Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

A short history of AI, from Turing to today's LLMs

Seventy years of trying to make machines think — and how a single architecture from 2017 finally cashed the check that 1950s AI wrote.

AI & ML intermediate Apr 29, 2026 · updated Aug 24, 2026 · 13 min read

On this page

The picture version

Seventy years of AI in six pictures, told through one stubborn sentence. The prose below fills in the names, dates and seams the pictures skip.

1 · The problem

One sentence that grammar cannot translate.

The trophy didn’t fit in the suitcase because it was too big. French makes you pick a gender for “it” le trophée so “it” is il la valise so “it” is elle ? Nothing in the grammar decides it. You have to know that things which don’t fit are too big, and containers they don’t fit into are too small.
Picking il or elle needs knowledge about trophies and suitcases, not knowledge about French. Seventy years of AI is largely seventy years of failing, then finally succeeding, at sentences like this one.

2 · The first bet

Write the knowledge down by hand.

symbolic AI, 1960s–70s: intelligence is symbol manipulation, so type the rules in trophy doesn’t fit → too big suitcase it won’t fit in → too small piano → … doorway → … every other object pair anyone might mention no end to this list and a human types every single line narrow domains: it worked MYCIN · DENDRAL RULES RAN OUT OF RULES
Hand-written rules were genuinely good inside a fenced-off domain like medical diagnosis. Open language has no fence: the knowledge needed is unbounded, and humans have to type all of it.

3 · The second bet

Stop writing rules. Count what people already wrote.

1990s–2000s: piles of already-translated documents EN ↔ FR, by the million count which phrases go together? …because it was… → car il seen far more often …because it was… → car elle seen far less often pick the bigger pile il right, this time That is not working out what “it” refers to. And a human still had to decide what to count — the hand-written-rules problem, one level down.
Phrase statistics stitch together locally plausible fragments, so the pronoun is decided by whichever gendered form was commoner in similar surroundings. Sometimes that lands; it is never reasoning about trophies.

4 · The wall

Attention was already there. Recurrence was the problem.

attention, added to recurrent translation in 2014 — already working the trophy the suitcase because it step n cannot start until step n−1 has finished so the GPU cluster you paid for looks like this: busy waiting · waiting · waiting · waiting · waiting · waiting · waiting · waiting In an era where progress tracked compute spent, “cannot be parallelized” is close to fatal.
Letting the model look back at earlier words was solved in 2014 and it was not enough. The blocker was reading one step at a time, which no amount of attention on top can fix.

5 · The move

Keep the attention. Delete the recurrence.

2017 — the whole sentence trains in one shot, no step-by-step the trophy the suitcase because it “it” reads both candidates directly, and weighs them and the same cluster now looks like this: all of it busy — so you can afford to spend a data centre on the training run Attention was already there. Recurrence was the part you could drop.
Once training could saturate a cluster, the winning move was scale: predict the next token over enough text and the pronoun sorts itself out. Nobody in that lineage set out to build a translator — the translation arrived as a side effect.

6 · Keep this card

Three arrows, and each one is a wall someone hit.

AI history ≈ symbolic logic ↓ rules ran out of rules statistical learning ↓ statistics ran out of hand-designed features scaled neural nets each step down costs more compute than the last — a ratchet, not a pendulum ↓ recurrence ran out of parallelism
Picture to keep: AI history ≈ symbolic logic → statistical learning → scaled neural nets — where every arrow is a failure boundary, not a change of fashion. Read that way the sequence looks less like taste and more like a series of forced moves.

Why it exists

You land on a page in a language you don’t read, hit Translate, and get something you can actually follow. Try to remember what that button did in 2010: grammatically mangled word-salad you had to reverse-engineer. Nobody announced the fix. It just quietly got good.

That one button is a decent x-ray of the whole field, so let’s use a single sentence as our running example the whole way down:

The trophy didn’t fit in the suitcase because it was too big.

Translate that into French and you hit a wall that looks trivial and isn’t. French makes you choose a gendered pronoun: il if “it” means the trophy (le trophée), elle if it means the suitcase (la valise). No amount of grammar tells you which. You have to know that things that don’t fit are too big, and containers that things don’t fit into are too small. Seventy years of AI is, in large part, seventy years of failing and then finally succeeding at sentences like that one.

The history matters because almost every “new” idea you’ll read about — agents, reasoning, alignment, even the fear of superintelligence — was proposed, tried, abandoned, and revived at least once before. Knowing the arc is how you tell a genuinely new idea from a rebrand.

Why it matters now

The machinery you use this week is old machinery. Every LLM answering your questions is trained by backpropagation (1986) on a transformer (2017) — that combination is what every frontier lab currently ships — and the translation, autocomplete, dictation, and photo-search features on your phone are the same lineage wearing different product names. None of that is a fast-moving frontier; it’s the settled floor everything else stands on.

What that buys you is a working filter for the news. The field has a strong habit of removing one bottleneck and immediately hitting the next, and each generation’s failure is what names the next generation’s idea. Once you can see that chain, “we removed a bottleneck” and “we changed everything” stop sounding alike — which is most of what you need to read an AI announcement carefully.

The short answer

AI history ≈ symbolic logic → statistical learning → scaled neural nets

Picture to keep: a pendulum between two workshops — one where people write the rules down by hand, one where machines infer the rules from piles of examples — swinging toward the second workshop every time hardware gets cheap enough to feed it. Where the pendulum image breaks: it isn’t symmetric. Each swing back toward learning has come with more compute than the last, so it’s less a pendulum than a ratchet.

For seventy years, AI alternated between two bets: encode human reasoning as explicit rules (symbolic AI), or learn patterns from data (statistical or connectionist AI). The standard account of why the second bet kept losing is that the compute and the datasets weren’t there; that’s the reading Sutton’s Bitter Lesson argues for, and it’s the one this post follows. It isn’t the only reading — the training algorithms and the benchmarks changed too, and untangling which mattered most is genuinely contested.

How it works

Read the timeline as a chain of failures. Each era does something real, hits a wall, and the wall names the next idea. Our French translation problem is the test case at every step.

1950 — the question gets made answerable. Alan Turing publishes Computing Machinery and Intelligence, asks “can machines think?”, and immediately swaps it for something testable: the imitation game. No machine could play it for decades, but the move — replace the philosophy with a behavioral test — is why the field is engineering and not metaphysics.

1956 — Dartmouth names it. A summer workshop coins artificial intelligence and sets the tone. The 1955 proposal is worth reading purely for the confidence: it budgets two months and ten people to make “a significant advance.” Early programs proved geometry theorems and played checkers.

1960s–70s — symbolic AI, and where it breaks. The bet: intelligence is symbol manipulation, so write knowledge as rules and run logic over them. This produced expert systems that were genuinely impressive inside narrow domains — MYCIN for bacterial infections, DENDRAL for chemistry.

Why it breaks: to translate our sentence, you’d need a rule saying trophies that don’t fit are too big. Then one for suitcases. Then for pianos, doorways, and every other object pair anyone might mention. The knowledge required is unbounded, and humans have to type all of it. Narrow domains hid this; open language exposed it.

Meanwhile — the underdog, and its ceiling. Frank Rosenblatt’s perceptron (1958) showed a single-layer network could learn a pattern from examples rather than be told it. Minsky and Papert’s 1969 book Perceptrons showed that a single layer can’t represent even XOR.

Why it breaks: the fix — stack more layers — was known, but nobody had a practical way to train the stack. Neural-net funding drained. The broader first AI winter of the mid-1970s had wider causes than this one book — the ALPAC report on machine translation and the UK’s Lighthill report are the usual culprits cited — but Perceptrons is the reason the connectionist branch specifically went quiet.

1986 — backpropagation makes depth trainable. Rumelhart, Hinton, and Williams popularize backpropagation, which computes how to adjust every weight in a deep network by pushing errors backward through the layers. In principle, Minsky’s objection is now answered.

Why it breaks: in practice the networks were tiny and slow, and simpler statistical methods beat them on the benchmarks that existed. A second winter arrived in the early 90s as expert systems also failed to scale.

1990s–2000s — statistics wins the practical fights. Spam filtering, search ranking, speech recognition, and translation all get good via classical machine learning. Statistical machine translation learns which phrases tend to correspond across languages, from piles of already-translated documents. It’s how that Translate button worked in its early years — and why it produced word-salad. Phrase statistics stitch together locally plausible fragments; the pronoun choice gets decided by whatever gendered form was more common in similar surrounding phrases, which is not the same as working out what “it” refers to. Sometimes that lands. It isn’t reasoning about trophies.

Why it breaks: these methods learn from features humans design. Somebody has to decide what to measure. That’s the hand-written-rules problem again, moved one level down.

2012 — AlexNet, and features stop being hand-designed. A deep neural network wins the ImageNet contest by a large margin (a top-5 error of 15.3% against 26.2% for the runner-up), trained on GPUs rather than CPUs. Its features were learned, not specified. Within a few years deep learning dominates vision and speech.

Why it breaks (for language): images are fixed-size grids. Sentences aren’t.

2014–2016 — sequence models, and the serial wall. RNNs and especially LSTMs become standard for text. They read a sentence left to right, carrying a hidden state, and translate from that state. The first version squeezed the entire source sentence into one fixed-size vector, which is a bad place to keep a trophy you’ll need to refer back to. That specific problem was fixed before the transformer: Bahdanau et al. added an attention mechanism to RNN translation in 2014, letting the decoder look back at every source position. Attention is not what the transformer invented.

Why it breaks: recurrence itself. Step n can’t be computed until step n−1 finishes, so training can’t exploit a GPU’s parallelism no matter how good the attention on top is. In an era where progress was starting to track compute spent, “cannot be parallelized” is close to fatal.

2017 — attention removes the wall. A Google paper, Attention Is All You Need, introduces the transformer — and, tellingly, demonstrates it on machine translation. Its move is to keep attention and delete the recurrence: every position reads its context directly rather than one step at a time, so a whole sentence trains in one shot. “It” can attend straight back to “trophy” and “suitcase” and weigh both — and, crucially, the training run can now saturate a GPU cluster. The title is the argument: attention was already there; recurrence turned out to be the part you could drop.

Why it breaks: the architecture doesn’t tell you what to train it on, or how much of it to build.

2018–2020 — GPT and scaling laws. OpenAI’s GPT-1, GPT-2, and GPT-3 show something quietly radical: train a transformer on enough text to just predict the next token, and capability improves predictably with size, data, and compute. GPT-3 (175 billion parameters, 2020) wrote essays and code with no task-specific training. The scaling laws made spending on a training run something you could forecast a return on. (That reading of why the money arrived is mine, not a documented decision process at any lab.) Notice what happened to our sentence along the way: nobody in this lineage set out to build a translation system. Translating the trophy sentence became a side effect of predicting text well enough.

Why it breaks: a next-token predictor is not an assistant. Show GPT-3 our sentence and ask for French, and one statistically reasonable continuation is another exam question about trophies, because that’s what such text tends to look like on the internet. Making the model treat the prompt as an instruction rather than a passage to continue is a separate problem.

2022 — RLHF, and the product leap. The fix is RLHF: fine-tune on demonstrations of following instructions, then shape behavior using human preference ratings. The technique wasn’t new that November: OpenAI had already published it as InstructGPT earlier in 2022. What ChatGPT added was the conversational framing — a model tuned for back-and-forth dialogue, behind a text box anyone could type into. Ask for French now and you get French. Not primarily a capability jump — a usability one. It broke containment; ChatGPT was widely reported as the fastest-growing consumer application to that point, though the much-quoted “100 million users in two months” figure came from a third-party estimate, not from OpenAI.

Why it breaks: a helpful answer is still just an answer. It can’t look anything up or check itself.

2023–2025 — tools, multimodality, agents. Models learn to see images (GPT-4V, Claude 3) and to call external functions. Now the trophy sentence can be a photo of a page, and the model can consult a glossary before committing to il or elle. Put an LLM in a loop where it decides what to call next, and you get the agent pattern — see agent harness.

Why it breaks: answering in one shot has a ceiling on problems that need working-out.

2024–2026 — reasoning and long horizons. OpenAI’s o-series and DeepSeek-R1 popularize models trained with reinforcement learning to produce a long chain of reasoning tokens before the final answer. (How much of that chain reflects the computation actually driving the answer is an open research question — treat “the model is thinking” as a description of the output format, not of the internals.) Context windows in some frontier models reach a million tokens. The live question stops being “can the model translate the sentence?” and becomes “can it stay coherent across a day of work?” Open-weights models have narrowed the gap to closed frontier ones on public benchmarks, though how much depends heavily on which benchmark you pick — I wouldn’t trust a single summary number.

So: why did the Translate button quietly get good? Because the pronoun in our sentence stopped being a lookup problem and became a side effect of a system trained to predict text in general — and that system only became trainable at that scale once you could drop recurrence and spend a data center on it.

You started with AI history ≈ symbolic logic → statistical learning → scaled neural nets. What does that arrow-diagram leave out? — that each arrow is a failure boundary, not a change of fashion: rules ran out of rules, statistics ran out of hand-designed features, recurrence ran out of parallelism. Read that way, the sequence looks less like taste and more like a series of forced moves.

Check yourself

Before you go — the transformer paper was demonstrated on translation, and The Bitter Lesson argues that hand-built human knowledge keeps losing to raw scale. So why did symbolic AI dominate for thirty years if it was the losing bet?

Answer

Because “raw scale wins” is a claim about the limit, not about 1975. With the compute and data available then, a learned system couldn’t beat a rule-based one at anything useful — the perceptron’s ceiling was real and Minsky’s critique was correct on its own terms. Symbolic AI won for thirty years because it was genuinely the better engineering choice at that hardware budget. The lesson isn’t “they were fools”; it’s that the winner of an approach comparison can flip when an input like compute changes by many orders of magnitude.

And one more — suppose someone announces a new architecture that beats the transformer on quality per parameter, but has to process tokens strictly in order. Based on the pattern above, what’s your first question?

Answer

“How does it train?” That’s the wall recurrent models hit before 2017. A quality win per parameter is worth little if you can’t parallelize training across a GPU cluster, because you’ll be outscaled by a worse-per-parameter model that trains ten times faster on ten times the data. Architectures on this timeline mostly won on trainability at scale, not on elegance — which is also why “beats the transformer on a small benchmark” has been announced many times without sticking.

Going deeper