Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What is a neural network?

A pile of multiplications and a 'how wrong was I?' signal — somehow, when you stack enough of them, the thing learns to read, see, and play chess.

AI & ML intro Apr 30, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Six pictures for a reader who has never seen inside one. The prose below fills in the seams the pictures skip.

1 · The problem

Write down what makes a 7 a 7. Someone will break your rule.

plain crossed tilted curled your rules has a flat top stroke has a bar through it never curls at the top each one broken above The rules exist. There are just far too many, and they interact. by the time you’ve enumerated them, someone’s grandmother writes a seven you’ve never seen
Recognising handwriting, faces, speech or spam all have this shape. The task isn’t unwritable because it’s mysterious — it’s unwritable because no person can list the rules faster than reality invents exceptions.

2 · The simplest machine

Score the picture with one number. Now try to draw the line.

pixels × a dial each add up score “seven?” 777777 111111 One straight cut. Some sevens always land on the wrong side. move the digit two pixels sideways and the score changes for reasons that have nothing to do with sevenness
A single scoring unit can learn something real — “ink up here, none down there” — but it can only ever split the world with one straight cut. The variety in scene 1 doesn’t fit on one side of any line.

3 · Depth, and the piece that makes it real

Stacking flat layers changes nothing. The bend is the whole trick.

layers with nothing between them straight straight straight equals one straight layer a hundred of them collapse into one. depth bought you nothing layers with a bend between them straight straight straight bend bend now the boundary can fold as many times as it needs to the bend is a tiny rule — e.g. “anything negative becomes zero” — applied after each layer
Two straight steps in a row are still one straight step, which is why a deep stack of plain layers is arithmetic theatre. Put a small kink after each layer and depth starts buying shapes a single cut can’t make.

4 · How it learns

Guess, measure how wrong, turn every dial a hair. Repeat.

one example millions of dials, none of them labelled its guess: “3” how wrong was that? one number turn every dial a hair in the direction that would have made that number smaller Do this a few million times and the crossed sevens come out as sevens.
Nobody sets the dials and nobody can say what any single one means. The only instruction the machine ever gets is “that was wrong by this much” — and the rules from scene 1 end up distributed across all of them.

5 · Why it’s affordable

You can’t test ten million dials one at a time.

the obvious way nudge dial #1 → run the whole thing nudge dial #2 → run the whole thing nudge dial #3 → run the whole thing … × 10,000,000 per example. hopeless. what’s actually done forward once then backward once and every dial has its direction each layer passes “how much did you affect the answer?” to the one before it One extra sweep instead of one re-run per dial. this is the difference between training being impossible and training being routine
The trick is that the layers are nested, so the effect of an early dial can be worked out from the effect of the later ones. That single backward pass is what made deep networks trainable at all — the idea, not the hardware, is the bottleneck it removed.

6 · Keep this card

The whole thing on one index card.

neural network = stacked layers of multiply, then bend + a number saying how wrong that was + nudge every dial, a few million times — the bend is load-bearing; the backward sweep is why it’s affordable
Picture to keep: a wall of dials. The image goes in one side, a guess comes out the other, and training turns every dial a hair in whichever direction made the guess less wrong — and no single dial means “has a crossbar”.

Why it exists

You photograph a cheque to deposit it in your banking app, and the app reads the amount out of your handwriting. Your 7 has a slash through it because you learned to write in Europe; the app doesn’t care. Before reading on: how would you write that program? What makes a 7 a 7 — a slanted top stroke, usually a horizontal bar, sometimes a crossbar through the middle, not always closed. By the time you’ve written the rules, someone’s grandmother writes a 7 you’ve never seen and your code says “1”.

That handwritten 7 is the running example for this whole post. A whole category of problems — recognizing faces, reading speech, telling spam from not-spam, predicting the next word — refused to yield to hand-coded rules. Not because the rules don’t exist. Because there are millions of them, they’re tangled, and no human can enumerate them faster than reality invents new edge cases.

Neural networks exist because you don’t have to know the rules. You only need examples. Show a flexible enough function enough labeled sevens, give it a way to measure how wrong it is, and let it nudge itself toward less wrong. The rules fall out as a side effect — encoded in millions of small numbers nobody has to interpret. Trade “I understand exactly what my program does” for “my program does things I couldn’t have written by hand.” On this class of problems, the field mostly took the second deal.

Why it matters now

The systems that made “AI” a consumer word are neural networks underneath: LLMs, image and video generators, speech recognition, recommendation feeds. They differ mostly in shape, size, and what they were trained on — not in kind. (Plenty of production machine learning is not a neural network; on tabular data — spreadsheets of rows and columns — gradient-boosted trees are still a standard first choice, and often the one that wins.)

You need a picture of one because the vocabulary of the field assumes you have it. Weights, layers, training, fine-tuning, gradients, parameters aren’t metaphors — they’re the literal mechanical parts. With the picture, almost everything else in modern ML becomes “okay, but bigger” or “okay, but with a clever twist.”

The short answer

neural network = stacked layers of (linear transform + non-linear activation), trained by gradient descent on a loss

Picture to keep: a wall of dials. The image of your 7 goes in one side, a guess comes out the other, and training turns every dial a hair in whichever direction made the guess less wrong, over and over. Except that nobody labels the dials: no single one means “has a crossbar,” and you can’t turn one to fix a specific mistake.

A neural net is a long pipeline of “multiply by some numbers, add some numbers, then bend the result.” Start with random numbers, feed in examples, measure how wrong the output is — that measurement is the loss — and adjust the numbers slightly toward less-wrong, which is what gradient descent means. The numbers that survive are the model.

How it works

Build it badly and let each failure force the next piece.

Naive attempt: score the pixels. Your 7 arrives as, say, 784 brightness numbers. Give each pixel a weight, add a bias, sum it up, and call a high score “7”:

score = w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b

This is one neuron, and it can genuinely learn “ink in the top row, none in the bottom left.” Why it breaks: it can only carve the input space with a straight cut. A 7 shifted two pixels right, or a slashed 7, or a 1 with a serif, all live on the wrong side of any single line you can draw.

Fix: stack layers — but bend them first. Put many neurons in parallel (a layer) and feed their outputs into another layer, so later units can combine “stroke here” and “no stroke there” into “corner.” Why that breaks on its own: a stack of linear layers is still linear — multiply the matrices together and a hundred layers collapse into one. Depth buys you nothing. The fix is a non-linear activation between layers, like ReLU, which clips negatives to zero. Without that kink, depth is arithmetic theatre. The simplest full shape is the MLP: input vector in, a few fully-connected hidden layers, ten outputs (one per digit). Stacks of simple bends are enough to approximate a very wide class of functions — that’s the universal approximation result (Cybenko 1989; Hornik, Stinchcombe & White 1989) — but it says nothing about whether you can find the right weights, which is the whole problem.

But now you have millions of dials and no idea where to set them. Random weights call your 7 a 3. You need a direction. So define a loss function that turns “predicted vs. actual” into one number. For your 7 that’s cross-entropy: the model outputs a probability for each of the ten digits, and the loss is large when it gave “7” a small one. (Other tasks swap the loss — squared error for predicting a number, next-token cross-entropy for an LLM — and the shape of the loop stays the same, even though plenty else about a real system doesn’t.) Now “less wrong” is measurable.

But measuring the effect of each dial separately is hopeless. The obvious way to ask “would nudging weight #418,203 lower the loss?” is to nudge it and re-run the network. With ten million weights, that’s ten million forward passes per example. The fix is the chain rule: the loss depends on the last layer’s weights, which depend on the previous layer’s, and so on, so you can push “how does the loss change if I nudge this?” backward through the whole network in a single sweep, getting a gradient for every weight at once. That bookkeeping is backpropagation, and it’s the difference between one backward sweep per batch and one forward pass per weight — which is why training deep networks is affordable at all.

Then take the step. Move each weight a little against its gradient: w ← w − η · ∂L/∂w. Do that batch after batch. The dials drift toward a setting where slashed sevens and grandmother sevens both come out as 7.

That’s the entire algorithm: pick a shape, pick a loss, initialize randomly, forward / loss / backward / step. The pieces that made modern deep learning work — convolutions, attention, residual connections, Adam — are mostly refinements of one of those four slots rather than replacements for the loop.

Two seams worth flagging. First, the “knowledge” is just numbers — a big grid of floating-point values with no symbolic representation you can read off, which is why interpretability is its own active field; nothing in the trained network says “7”. Second, why this works as well as it does is an open research question. The loss surface of a big network is non-convex, so there’s no guarantee that rolling downhill lands anywhere good — let alone somewhere that also handles sevens it never saw. In practice it usually does. There are partial explanations (overparameterization, implicit regularization, high-dimensional geometry) and no consensus account.

You started with `neural network = stacked layers of (linear transform

Going deeper

Where the line sits: the mechanical story (neurons, layers, forward/loss/backward, gradient descent) is established and hasn’t changed in decades. Why deep nets generalize as well as they do, and what their internal representations actually mean, are still open research questions — so a confident, tidy story about either is a sign to check the source, not a sign of understanding.