Skip to main content

Lesson 2 — Why depth changes everything

A single layer of neurons can, in theory, approximate almost any function — a result known as the universal approximation theorem. So why stack dozens of layers? The theorem says such a network exists; it says nothing about how wide it would have to be or whether you could ever train it. In practice depth is enormously more efficient than width, and the reason is worth understanding.

The hierarchy that builds itself​

The defining property of deep learning is that the network invents its own intermediate representations. This is not a metaphor; it is observable by inspecting what the layers respond to after training.

Take an image network trained only on labelled photographs, with no instruction about visual structure. Examine what each layer has learned to detect:

Layer depthWhat it responds to
First layersedges, corners, patches of colour, orientations
Early middlesimple textures, repeated patterns, curves
Middlerecognisable parts: an eye, a wheel, a leaf, a doorway
Deeperassemblies: a face, a car, a plant, a building
Final layerthe categories you asked for

Nobody programmed that progression. Nobody told the network that edges combine into textures, or that eyes and noses appear together in faces. It emerged because that hierarchy is genuinely present in the data, and building it was the most efficient way to reduce the training error.

Why this is the whole point

Lesson 1 of the Introduction to AI course described Polanyi's paradox: you cannot write down the rules you use to recognise a cat. Deep learning does not solve that problem by extracting the rules from you. It sidesteps it entirely by discovering its own — which is why it works precisely where feature engineering is impossible.

Why depth beats width​

Consider building a face detector.

Shallow and wide: one enormous layer that must recognise every possible face directly from raw pixels. Each unit has to handle every combination of position, lighting, angle and expression independently. The number of units required explodes, and there is no sharing of effort.

Deep and narrow: the first layers learn edges, which are useful for every face and indeed for every object. The next layers assemble edges into eyes and mouths, reusing the edge detectors. The next assemble those into faces, reusing the part detectors.

Depth allows reuse. Each layer builds on the abstractions below it instead of starting from pixels, so the same edge detector serves a face, a car and a tree. That compositional efficiency is why a deep network with a million parameters can outperform a shallow one with a hundred million.

The hierarchy an image network builds. Raw pixels enter, the first two layers detect edges, colours and orientations, layers 3 to 5 textures and curves, layers 6 to 10 object parts such as an eye or a wheel, layers 11 to 20 whole objects, and the final layer produces the prediction.Raw pixelsLayers 1 to 2edges, colours, orientationsLayers 3 to 5textures, curves, simple patternsLayers 6 to 10parts: an eye, a wheel, a leafLayers 11 to 20objects: a face, a car, a plantPrediction
Nobody programmed this progression. The edge detectors learned near the bottom are reused by everything above them, for every category — which is exactly the sharing of effort a wide shallow network cannot do.

What depth cost, and what fixed it​

Depth was understood to be desirable long before it was achievable. Networks past about twenty layers trained worse than shallower ones, which was baffling: a deeper network can always in principle imitate a shallower one by making the extra layers do nothing, so it should never be worse.

The culprit was the vanishing gradient, described in the maths course. Backpropagation multiplies rates of change along the chain from output back to input. Multiply thirty numbers smaller than one and you get something microscopic, so the earliest layers received essentially no signal and stopped learning.

Two changes resolved it:

ReLU activations. Their slope is exactly 1 for positive inputs, so gradients pass through undiminished rather than being repeatedly shrunk by a squashing function.

Residual connections. Introduced in 2015 with ResNet, and the more important of the two. A residual connection adds a shortcut that lets the input of a block skip past it and be added to its output. The gradient can then travel back along the shortcut without being attenuated by the layers.

The effect was immediate: networks jumped from around 20 usable layers to over 150, with better results. Residual connections are now in essentially every deep architecture, including every transformer, which makes them one of the most consequential small ideas in the field.

A residual block. The signal passes through two layers and reaches an addition. In parallel a shortcut leaves the block input, travels above both layers without crossing them, and joins the same addition.block inputlayer 1layer 2residual connection: a direct paththat crosses no layer at all+output
The shortcut adds the block input to the block output. The gradient can therefore travel back along a path that escapes the multiplications responsible for its vanishing — which is what took networks from twenty layers to over a hundred.

What depth does not fix​

Being clear about the limits, since depth is often proposed as a solution to problems it cannot touch:

Insufficient data. A deeper network has more parameters and therefore needs more data, not less. Depth on a small dataset produces confident memorisation.

Bad labels. The network will learn your labelling errors faithfully and with great precision.

Missing information. If the signal is not in the input, no architecture recovers it. Depth discovers features within the data; it does not invent data.

Tabular data. Adding layers does not make a neural network competitive with gradient boosting on a spreadsheet. The hierarchy that depth exploits exists in pixels and language; it is largely absent from a table of columns that were designed by a human in the first place.

The reflex to resist

When a model underperforms, "make it deeper" is rarely the answer. Check the data volume, the label quality, and whether the input actually contains the signal. Depth amplifies what is there, including the problems.

Two ideas you will see everywhere​

Two techniques accompany depth so consistently that they are worth naming.

Batch normalisation rescales the values flowing between layers so they stay in a well-behaved range. Deep networks otherwise suffer from values growing or shrinking as they propagate, which destabilises training. It made deep networks train faster and tolerate a wider range of learning rates.

Dropout randomly switches off a fraction of neurons during each training step. It sounds destructive and it is a regularisation technique: the network cannot rely on any single unit, so it is pushed towards redundant, robust representations rather than brittle memorisation. It is switched off at prediction time.

A four-layer network during training with dropout. Several neurons of the hidden layers are crossed out and their connections dashed, showing that they take no part in this particular training step.inputshidden 1hidden 2output
One training step with dropout: the crossed-out neurons are ignored, and the surviving ones must cope without them. Each step draws a different set, so no single unit ever becomes indispensable. Press the button to draw another step.

In three sentences​

Depth matters because each layer builds on the abstractions below it, so a network discovers a reusable hierarchy — edges, then textures, then parts, then objects — that nobody programmed and that makes deep networks far more efficient than wide ones. That hierarchy is exactly why deep learning succeeds where features cannot be hand-designed, and why it offers little on tabular data whose columns a human already designed. Depth was unusable until ReLU and residual connections let gradients travel back through many layers, which is what took networks from twenty layers to hundreds.


Next — Lesson 3: CNNs, how machines see →