Activation Functions

next

Activation Functions

Compare common activation functions and their derivatives. Drag the probe, toggle which functions are visible, and adjust parameters to explore saturation and gradient flow.

Visible
Ready

Concepts

  • Non‑linearity enables complex decision boundaries.
  • Saturation: flat derivative slows learning.
  • Gradient magnitude drives update size.
  • Range & symmetry influence training dynamics.

Probe

Drag the vertical probe to inspect f(x) and f'(x). Hover legend to highlight a function.


                

Legend

Guidance

Explore which functions keep gradients alive (derivative not ~0) across wider input ranges.