What is a neural network?
A pile of multiplications and a 'how wrong was I?' signal — somehow, when you stack enough of them, the thing learns to read, see, and play chess.
On this page
The picture version
Six pictures for a reader who has never seen inside one. The prose below fills in the seams the pictures skip.
1 · The problem
Write down what makes a 7 a 7. Someone will break your rule.
2 · The simplest machine
Score the picture with one number. Now try to draw the line.
3 · Depth, and the piece that makes it real
Stacking flat layers changes nothing. The bend is the whole trick.
4 · How it learns
Guess, measure how wrong, turn every dial a hair. Repeat.
5 · Why it’s affordable
You can’t test ten million dials one at a time.
6 · Keep this card
The whole thing on one index card.
Why it exists
You photograph a cheque to deposit it in your banking app, and the app reads the amount out of your handwriting. Your 7 has a slash through it because you learned to write in Europe; the app doesn’t care. Before reading on: how would you write that program? What makes a 7 a 7 — a slanted top stroke, usually a horizontal bar, sometimes a crossbar through the middle, not always closed. By the time you’ve written the rules, someone’s grandmother writes a 7 you’ve never seen and your code says “1”.
That handwritten 7 is the running example for this whole post. A whole category of problems — recognizing faces, reading speech, telling spam from not-spam, predicting the next word — refused to yield to hand-coded rules. Not because the rules don’t exist. Because there are millions of them, they’re tangled, and no human can enumerate them faster than reality invents new edge cases.
Neural networks exist because you don’t have to know the rules. You only need examples. Show a flexible enough function enough labeled sevens, give it a way to measure how wrong it is, and let it nudge itself toward less wrong. The rules fall out as a side effect — encoded in millions of small numbers nobody has to interpret. Trade “I understand exactly what my program does” for “my program does things I couldn’t have written by hand.” On this class of problems, the field mostly took the second deal.
Why it matters now
The systems that made “AI” a consumer word are neural networks underneath: LLMs, image and video generators, speech recognition, recommendation feeds. They differ mostly in shape, size, and what they were trained on — not in kind. (Plenty of production machine learning is not a neural network; on tabular data — spreadsheets of rows and columns — gradient-boosted trees are still a standard first choice, and often the one that wins.)
You need a picture of one because the vocabulary of the field assumes you have it. Weights, layers, training, fine-tuning, gradients, parameters aren’t metaphors — they’re the literal mechanical parts. With the picture, almost everything else in modern ML becomes “okay, but bigger” or “okay, but with a clever twist.”
The short answer
neural network = stacked layers of (linear transform + non-linear activation), trained by gradient descent on a loss
Picture to keep: a wall of dials. The image of your 7 goes in one side, a guess comes out the other, and training turns every dial a hair in whichever direction made the guess less wrong, over and over. Except that nobody labels the dials: no single one means “has a crossbar,” and you can’t turn one to fix a specific mistake.
A neural net is a long pipeline of “multiply by some numbers, add some numbers, then bend the result.” Start with random numbers, feed in examples, measure how wrong the output is — that measurement is the loss — and adjust the numbers slightly toward less-wrong, which is what gradient descent means. The numbers that survive are the model.
How it works
Build it badly and let each failure force the next piece.
Naive attempt: score the pixels. Your 7 arrives as, say, 784 brightness numbers. Give each pixel a weight, add a bias, sum it up, and call a high score “7”:
score = w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b
This is one neuron, and it can genuinely learn “ink in the top row, none in the bottom left.” Why it breaks: it can only carve the input space with a straight cut. A 7 shifted two pixels right, or a slashed 7, or a 1 with a serif, all live on the wrong side of any single line you can draw.
Fix: stack layers — but bend them first. Put many neurons in parallel (a layer) and feed their outputs into another layer, so later units can combine “stroke here” and “no stroke there” into “corner.” Why that breaks on its own: a stack of linear layers is still linear — multiply the matrices together and a hundred layers collapse into one. Depth buys you nothing. The fix is a non-linear activation between layers, like ReLU, which clips negatives to zero. Without that kink, depth is arithmetic theatre. The simplest full shape is the MLP: input vector in, a few fully-connected hidden layers, ten outputs (one per digit). Stacks of simple bends are enough to approximate a very wide class of functions — that’s the universal approximation result (Cybenko 1989; Hornik, Stinchcombe & White 1989) — but it says nothing about whether you can find the right weights, which is the whole problem.
But now you have millions of dials and no idea where to set them. Random weights call your 7 a 3. You need a direction. So define a loss function that turns “predicted vs. actual” into one number. For your 7 that’s cross-entropy: the model outputs a probability for each of the ten digits, and the loss is large when it gave “7” a small one. (Other tasks swap the loss — squared error for predicting a number, next-token cross-entropy for an LLM — and the shape of the loop stays the same, even though plenty else about a real system doesn’t.) Now “less wrong” is measurable.
But measuring the effect of each dial separately is hopeless. The obvious way to ask “would nudging weight #418,203 lower the loss?” is to nudge it and re-run the network. With ten million weights, that’s ten million forward passes per example. The fix is the chain rule: the loss depends on the last layer’s weights, which depend on the previous layer’s, and so on, so you can push “how does the loss change if I nudge this?” backward through the whole network in a single sweep, getting a gradient for every weight at once. That bookkeeping is backpropagation, and it’s the difference between one backward sweep per batch and one forward pass per weight — which is why training deep networks is affordable at all.
Then take the step. Move each weight a little against its
gradient: w ← w − η · ∂L/∂w. Do that batch after batch. The dials
drift toward a setting where slashed sevens and grandmother sevens
both come out as 7.
That’s the entire algorithm: pick a shape, pick a loss, initialize randomly, forward / loss / backward / step. The pieces that made modern deep learning work — convolutions, attention, residual connections, Adam — are mostly refinements of one of those four slots rather than replacements for the loop.
Two seams worth flagging. First, the “knowledge” is just numbers — a big grid of floating-point values with no symbolic representation you can read off, which is why interpretability is its own active field; nothing in the trained network says “7”. Second, why this works as well as it does is an open research question. The loss surface of a big network is non-convex, so there’s no guarantee that rolling downhill lands anywhere good — let alone somewhere that also handles sevens it never saw. In practice it usually does. There are partial explanations (overparameterization, implicit regularization, high-dimensional geometry) and no consensus account.
You started with `neural network = stacked layers of (linear transform
- non-linear activation), trained by gradient descent on a loss
. What did building it badly add? —+ the non-linearity is load-bearing, not decoration, and+ backprop is why the whole thing is affordable`. Drop the first and depth is a lie; drop the second and the algorithm is correct but unrunnable.
Famous related terms
- Transformer —
transformer = stack of (attention + feed-forward) blocks. The architecture LLMs are built on — a specific shape, not a different kind of animal. - Backpropagation —
backprop = chain rule + caching intermediate values. Makes gradient computation linear in the number of weights instead of catastrophic. - Gradient descent —
gradient descent = repeatedly step against the gradient of the loss. Variants (SGD, Adam, AdamW) differ in how they scale the steps. - MLP / feed-forward network —
MLP = input → fully-connected hidden layers → output. The simplest shape; still the building block inside transformers. - Embedding —
embedding = learned vector representation of a discrete thing. How modern networks turn IDs into something continuous to multiply. - LLM —
LLM = transformer + next-token objective at scale. The most famous current consumer of all of the above.
Going deeper
- Rumelhart, Hinton & Williams, Learning representations by back-propagating errors (Nature, 1986) — the primary source: what did backprop actually claim to solve, in four pages, before any of this was fashionable?
- 3Blue1Brown’s Neural Networks series — answers “what is gradient descent doing, geometrically?” — and it works the same handwritten-digit example this post does.
- Andrej Karpathy’s Neural Networks: Zero to Hero — the rabbit hole, for “would I get the same answer if I derived it myself?”: builds backprop, then a language model, from an empty file.
Where the line sits: the mechanical story (neurons, layers, forward/loss/backward, gradient descent) is established and hasn’t changed in decades. Why deep nets generalize as well as they do, and what their internal representations actually mean, are still open research questions — so a confident, tidy story about either is a sign to check the source, not a sign of understanding.