Chapter I
Neurons as Logic
In 1943 the idea that the brain computes was put in precise form. Warren McCulloch, a neurophysiologist in Chicago, and Walter Pitts, a largely self-taught teenage logician, took the all-or-none impulse of electrophysiology and the excitatory and inhibitory synapses of Sherrington, and stripped them to essentials. A model neuron adds its inputs and fires if the sum reaches a threshold. They proved that networks of such units can compute any statement of logic, and, given a tape for memory, anything a Turing machine can. John von Neumann borrowed their notation when he described the design of the stored-program computer in 1945.
In 1949 the Canadian psychologist Donald Hebb proposed how such networks could learn. When one neuron repeatedly helps to fire another, the connection between them grows stronger. Groups of neurons that are often active together would bind into assemblies that stand for things and ideas. It was a guess, and a quarter of a century passed before long-term potentiation, found by the students of synaptic transmission, gave it strong support.
Chapter II
Machines That Learn
In 1958 Frank Rosenblatt described the perceptron, a network of threshold units with a rule for learning. After each wrong answer, the weights of the active inputs are nudged towards the right answer. He proved that if some setting of the weights solves a classification, the rule will find it, and built a machine, wired to a camera of 400 light sensors, that learned to tell simple shapes apart. Newspapers promised machines that would soon walk, talk and think.
In 1969 Marvin Minsky and Seymour Papert published Perceptrons, which proved that a single layer of trainable units cannot compute some simple functions. Multilayer networks could, but nobody knew how to train them. Rosenblatt died in 1971, and research on neural networks waned.
Meanwhile David Marr in Cambridge and then at MIT built theories tied to real circuits, first of the cerebellum and then of vision. His lasting contribution was a way of thinking. A process in the brain must be understood at three levels: what problem it solves, what representation and procedure solve it, and how the hardware carries that out. Knowing every neuron is not enough without the first level.
Chapter III
A Closer Look: Logic from Thresholds
A McCulloch–Pitts unit, in the simplified form usually taught today, takes inputs that are each 0 or 1, multiplies each by a weight , and fires, giving output 1, if
where is its threshold. With two inputs of weight 1, a threshold of 2 makes an AND gate, since both inputs must be active. A threshold of 1 makes an OR gate. A single input of weight with threshold 0 makes a NOT gate: it fires when the input is 0, because , and is silent when it is 1, because .
| AND () | OR () | XOR | ||
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 |
| 0 | 1 | 0 | 1 | 1 |
| 1 | 0 | 0 | 1 | 1 |
| 1 | 1 | 1 | 1 | 0 |
The last column, exclusive or, fires when exactly one input is active. No single unit can compute it. The unit must be silent for input , so . It must fire for and , so and . Adding these gives , which is greater than because is positive. So the unit would also fire for , which is wrong. A computer search over every integer weight and threshold from to finds none that works, as the argument says it must.
Two layers solve it. Feed and to an OR unit and an AND unit, then feed those to a third unit with weights from OR and from AND, and threshold 1. For the third unit receives and fires. For it receives and stays silent. This is the whole of Minsky and Papert's objection and its answer in miniature: one layer cannot, two layers can, and the difficulty was finding the weights of the hidden layer by learning rather than by hand.
Chapter IV
Energy, Errors and Wiring
In 1982 John Hopfield, a physicist, gave a network of symmetrically connected units an energy that can only fall as the units update. Memories stored by a Hebbian rule become valleys in this energy, and a network started from a fragment slides into the nearest whole memory, the same mathematics as a magnet in statistical mechanics. A network of 1,000 units can hold about 138 random memories before they blur together. In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams showed that backpropagation, passing errors backwards through the layers, trains the hidden units that Minsky and Papert had found lacking. Deep learning grew from this.
Models also needed real circuits. In 1986 John White and Sydney Brenner published the complete wiring of the worm C. elegans, and in 2024 the FlyWire consortium mapped all 140,000 neurons of a fruit fly's brain. Even with the worm's wiring known for decades, its behaviour cannot yet be predicted from it, because the diagram does not say how strong each synapse is or what signals it uses. How the brain itself solves the problem that backpropagation solves for machines, deciding which synapses to change after a mistake, is still unknown.