Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    PS5 emulation is suddenly making big strides on PC

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»What the ReLU Revolution Revealed About Biological Plausibility
    AI Tools

    What the ReLU Revolution Revealed About Biological Plausibility

    By No Comments11 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    What the ReLU Revolution Revealed About Biological Plausibility
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Introduction

    Few decisions in the design of a neural network have attracted as much attention, and as much revision, as the choice of activation function. On the surface it’s a narrow technical question, but the activation function is where the comparison with biology is most literal, since it decides what it means for an artificial neuron to fire. That makes it the natural place for an old question to resurface: how closely should an artificial neuron resemble a real one?

    The early answer was a great deal. Warren McCulloch and Walter Pitts proposed the first formal model of the neuron in 1943[1], a binary threshold device that fired whenever its inputs crossed a critical value. Frank Rosenblatt’s perceptron added adjustable weights in 1957[2], turning a static logical unit into something that could learn from data. These models were openly biological in their inspiration, since the ambition was not merely to compute but to think. That objective carried through to the sigmoid function, whose smooth, saturating curve was taken to mirror the graded firing of real neurons, and for decades biological plausibility served as both a design principle and a source of prestige.

    In the early 2010s, the Rectified Linear Unit (ReLU) displaced the sigmoid because it made deep networks trainable where the sigmoid had not, and it looked like the field had settled the matter in favour of pragmatism. The story is less tidy than that, however, and following it through calls into question how much weight biology ever carried.

    The sigmoid and the vanishing gradient

    From the 1980s into the early 2000s, the sigmoid was the activation function of choice for hidden layers. Its appeal was partly mathematical, since it’s smooth and differentiable everywhere, but its S-shaped curve also offered a natural neural interpretation, with the output read as the probability that a neuron fires given its input,

    f(x)=11+e−xf(x) = frac{1}{1 + e^{-x}}f(x)=1+e−x1​

    Weak inputs produce outputs near zero, while strong inputs drive the neuron toward saturation (see left panel in figure below). It was, in other words, a model of graded neural firing.

    The sigmoid activation function (left) and its derivative (right; image by author).

    The sigmoid’s limitations surfaced in how networks learn. Training adjusts each weight in proportion to the gradient of the loss with respect to that weight, and backpropagation computes these gradients by passing a signal backward from the output layer, one layer at a time. Because each step applies the chain rule, the gradient for a weight in layer l is a product of derivatives through every layer above it:

    ∂L∂W(l)=∂L∂a(L)⋅∂a(L)∂a(L−1)⋅⋯⋅∂a(l+1)∂a(l)⋅∂a(l)∂W(l)frac{partial L}{partial W^{(l)}} = frac{partial L}{partial a^{(L)}} cdot frac{partial a^{(L)}}{partial a^{(L-1)}} cdot cdots cdot frac{partial a^{(l+1)}}{partial a^{(l)}} cdot frac{partial a^{(l)}}{partial W^{(l)}}∂W(l)∂L​=∂a(L)∂L​⋅∂a(L−1)∂a(L)​⋅⋯⋅∂a(l)∂a(l+1)​⋅∂W(l)∂a(l)​

    where a(k)a^{(k)}a(k) is the activation of layer kkk.

    This multiplicative structure means the derivative of the activation function largely decides whether the gradient survives the journey. When the derivatives are consistently small, the product shrinks with every layer, and the gradient vanishes before it reaches the earliest layers, leaving them to learn slowly or not at all. When the derivatives are large, the opposite happens: gradients explode and training destabilises. Learning needs a regime between the two, in which gradients are large enough to carry meaningful information but not so large as to overwhelm it.

    The sigmoid falls firmly on the vanishing side. Its derivative takes a simple form:

    f′(x)=f(x)[1−f(x)]f'(x) = f(x)[1 – f(x)]f′(x)=f(x)[1−f(x)]

    and even at its peak, when x=0x = 0x=0, it reaches only f′(0)=0.25f'(0) = 0.25f′(0)=0.25 (see right panel in figure above). Each sigmoid layer therefore contributes a factor of at most 0.25 to the product, and far less once a unit saturates. In shallow networks this is manageable, but as depth increases, the signal decays rapidly and learning slows to a crawl. The hyperbolic tangent was often preferred in hidden layers because its outputs centre on zero, but it eased the problem without solving it. Its derivative peaks at one, yet it still falls away as the unit saturates. Careful weight initialisation and layer-wise pre-training could mitigate the decay, but saturating activation functions remained a significant barrier to building deep networks.

    The shift to ReLU

    In the early 2010s, a solution emerged that was so disarmingly simple it looked wholly unimpressive on paper (see, for example, Nair and Hinton[3]; Glorot, Bordes, and Bengio[4]). Rather than a smooth, S-shaped curve that compresses everything into a narrow range, ReLU returns its input unchanged when it is positive and outputs zero otherwise. Its definition fits on a single line:

    f(x)=max⁡[0,x]↦{0if x<0xif x≥0f(x) = max[0,x] mapsto begin{cases} 0 & quad text{if } x < 0 \ x & quad text{if } x geq 0 end{cases} f(x)=max[0,x]↦{0x​if x<0if x≥0​

    What made ReLU so effective is its derivative. For any positive input it is exactly one (if f(x)=xf(x) = xf(x)=x then f′(x)=1f'(x) = 1f′(x)=1; see right panel in figure below), so a gradient passing back through an active unit is neither diminished nor amplified by the activation itself. Each sigmoid layer contributed a factor of at most 0.25 to the chain of derivatives, whereas an active ReLU contributes a factor of one, which allows gradients to reach the earliest layers of a deep network largely intact. For negative inputs, the derivative is zero, a cost we will return to.

    The ReLU activation function (left) and its derivative (right; image by author).

    The kink at zero also changes what a network can represent. Each unit behaves as a switch, passing a linear response when active and nothing when not, so a layer of ReLUs carves its input space into regions, each with its own linear map. Stacking layers lets later ones subdivide the partitions made by earlier ones, and the number of regions can grow exponentially with depth[5]. That goes some way to explaining how networks built from such plain pieces can capture highly non-linear patterns.

    Those results came later, however. ReLU was adopted because deep networks built with it trained, often considerably faster than their sigmoid counterparts, and to most practitioners it looked like the moment the field traded biological fidelity for results. An unbounded, piecewise-linear function bore little obvious likeness to the graded, saturating response of a real neuron.

    But what of biology?

    There is an irony buried in this story, because the biological case for rectification arrived at the same moment as ReLU itself. In 2011, Xavier Glorot and colleagues presented a paper arguing that rectified units were a better model of cortical neurons than the sigmoid had ever been[4]. As they pointed out, real neurons are silent for much of the time and fire sparsely when they do, with a firing rate of zero below threshold, whereas a sigmoid unit receiving no input idles at half its maximum output. By that measure, the function the field had prized for its realism was the less realistic of the two.

    Nor was the idea new to neuroscience. A decade earlier, Richard Hahnloser and collaborators had built rectification into a cortex-inspired circuit, drawing on the threshold behaviour of excitatory neurons, and showed that networks of rectified units could both select among competing inputs and respond to them in a graded way[6].

    The case against ReLU as a biological analogue rests mainly on its unbounded output. That case is fair, since a real neuron cannot fire faster than its refractory period allows and its activity must eventually saturate, but it is narrower than ReLU’s reputation for artificiality suggests. It also isn’t what decided matters. ReLU spread because it materially improved the training of deep networks, and the biological argument, although made in print at the moment of its adoption, largely dropped out of the field’s narrative. A community that had once chosen its activation functions for their biology was handed a winner with better biological credentials than the incumbent, and it remembered that winner as a pragmatic hack.

    The cost of simplicity

    ReLU’s hard zero came at a price. Because the function outputs zero for any negative input, a unit can drift into a state where it outputs zero for every example in the training data, and its gradient is then zero too. The weights feeding into it stop updating. Changes elsewhere in the network can occasionally push its inputs back into the positive regime, but in practice such a unit rarely recovers. This became known as the dying ReLU problem, and in deep networks trained with high learning rates it could claim a significant fraction of a network’s capacity[7].

    The variants that followed (see figure below) share a common compromise, a negotiated balance between preserving what made ReLU effective and repairing the points of failure. Leaky ReLU keeps a small slope for negative inputs so that gradients never quite switch off[8]. ELU replaces the hard zero with a smooth exponential curve, which also pulls mean activations toward zero[9]. SELU fixes ELU’s constants so that activations tend to stay normalised as they pass through the network[10]. Softplus, Swish, and GELU[11] smooth out the kink itself. In each case, some of ReLU’s simplicity is traded for stability. For most feedforward and convolutional networks, plain ReLU remains a sensible default, with these variants as a straightforward remedy when its limitations surface.

    Activation functions (image by author).

    What these variants also share is the nature of their justification. Each was motivated by gradient behaviour, training dynamics, or benchmark performance, and none by a closer resemblance to real neurons. Swish makes the point most plainly, since it was found not by reasoning about neurons but by an automated search over candidate functions[12]. By this stage, the question of what a neuron should look like had shifted to a question of what trains best.

    After ReLU

    In one reading, ReLU won. It dethroned the sigmoid in hidden layers so completely that returning to it is now almost unthinkable. Yet at the frontier it has itself been displaced. The original transformer used ReLU[13], but its successors soon settled on GELU[14], and many recent large language models have since moved to gated variants such as SwiGLU[15,16]. Noam Shazeer offered no theory for why these variants work and, with some wit, attributed their success to “divine benevolence”[16]. It would be hard to find a more candid statement of where the field’s priorities now lie.

    The pattern is a familiar one. The sigmoid gave way to ReLU because depth exposed its limitations, and ReLU gave way to GELU and its successors because scale and new architectures exposed its own. There is no particular reason to think the process has reached an end.

    What this progression reveals is that the activation function has never been a fixed biological commitment but a working hypothesis revised under empirical pressure. The artificial neuron endures, but what it means for one to fire has been renegotiated with each generation of architectures. Biological plausibility, it turns out, was less a design principle than a loose inspiration, invoked when convenient and set aside when the numbers demanded it, and the story of ReLU suggests that was true even in the sigmoid era.

    That ReLU has been displaced at the frontier does not diminish what it made possible. The deep networks it enabled were the precondition for the architectures that eventually superseded it. In that sense the revolution was self-consuming: ReLU removed the barrier to progress, and progress, in turn, revealed its limits. That is not a demotion but precisely the mark of a foundational idea.

    ···

    References

    [1] Warren S. McCulloch and Walter Pitts, A logical calculus of the ideas immanent in nervous activity (1943), Bulletin of Mathematical Biophysics.

    [2] Frank Rosenblatt, The perceptron: a perceiving and recognizing automaton (1957), Cornell Aeronautical Laboratory, Inc. Report (No. 85-460-1).

    [3] Vinod Nair and Geoffrey Hinton, Rectified linear units improve restricted Boltzmann machines (2010), Proceedings of the 27th International Conference on Machine Learning.

    [4] Xavier Glorot, Antoine Bordes, and Yoshua Bengio, Deep sparse rectifier neural networks (2011), Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics.

    [5] Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio, On the number of linear regions of deep neural networks (2014), arXiv:1402.1869 [stat.ML].

    [6] Richard H. R. Hahnloser, Rahul Sarpeshkar, Misha A. Mahowald, Rodney J. Douglas, and H. Sebastian Seung, Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit (2000), Nature.

    [7] Lu Lu, Yeonjong Shin, Yanhui Su, and George Em Karniadakis, Dying ReLU and initialization: theory and numerical examples (2019), arXiv:1903.06733 [stat.ML].

    [8] Andrew Maas, Awni Hannun, and Andrew Ng, Rectifier nonlinearities improve neural network acoustic models (2013), ICML 2013 Workshop on Deep Learning for Audio, Speech and Language Processing.

    [9] Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter, Fast and accurate deep network learning by exponential linear units (ELUs) (2015), arXiv:1511.07289 [cs.LG].

    [10] Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter, Self-normalizing neural networks (2017), arXiv:1706.02515 [cs.LG].

    [11] Dan Hendrycks and Kevin Gimpel, Gaussian error linear units (GELUs) (2016), arXiv:1606.08415 [cs.LG].

    [12] Prajit Ramachandran, Barret Zoph, and Quoc V. Le, Searching for activation functions (2017), arXiv:1710.05941 [cs.NE].

    [13] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin, Attention is all you need (2017), arXiv:1706.03762 [cs.CL].

    [14] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, BERT: pre-training of deep bidirectional transformers for language understanding (2018), arXiv:1810.04805 [cs.CL].

    [15] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample, LLaMA: open and efficient foundation language models (2023), arXiv:2302.13971 [cs.CL].

    [16] Noam Shazeer, GLU variants improve transformer (2020), arXiv:2002.05202 [cs.LG].

    biological Plausibility ReLU revealed Revolution
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleSatlyt, founded by a former Google and SpaceX product manager, raises $8M to run AI on satellites
    Next Article PS5 emulation is suddenly making big strides on PC
    • Website

    Related Posts

    AI Tools

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    AI Tools

    How to Build a 10-Slide Deck in Decktopus AI: A Step-by-Step Walkthrough

    AI Tools

    Your AI Bill Is a Toll Booth. Stop Paying Twice.

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    0 Views

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    0 Views

    PS5 emulation is suddenly making big strides on PC

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Photon held a funeral for mobile apps. Now it has $4.5M to help replace them with agents

    0 Views

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    0 Views

    PS5 emulation is suddenly making big strides on PC

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.