A short history of AI, from Turing to today's LLMs
Seventy years of trying to make machines think — and how a single architecture from 2017 finally cashed the check that 1950s AI wrote.
On this page
The picture version
Seventy years of AI in six pictures, told through one stubborn sentence. The prose below fills in the names, dates and seams the pictures skip.
1 · The problem
One sentence that grammar cannot translate.
2 · The first bet
Write the knowledge down by hand.
3 · The second bet
Stop writing rules. Count what people already wrote.
4 · The wall
Attention was already there. Recurrence was the problem.
5 · The move
Keep the attention. Delete the recurrence.
6 · Keep this card
Three arrows, and each one is a wall someone hit.
Why it exists
You land on a page in a language you don’t read, hit Translate, and get something you can actually follow. Try to remember what that button did in 2010: grammatically mangled word-salad you had to reverse-engineer. Nobody announced the fix. It just quietly got good.
That one button is a decent x-ray of the whole field, so let’s use a single sentence as our running example the whole way down:
The trophy didn’t fit in the suitcase because it was too big.
Translate that into French and you hit a wall that looks trivial and isn’t. French makes you choose a gendered pronoun: il if “it” means the trophy (le trophée), elle if it means the suitcase (la valise). No amount of grammar tells you which. You have to know that things that don’t fit are too big, and containers that things don’t fit into are too small. Seventy years of AI is, in large part, seventy years of failing and then finally succeeding at sentences like that one.
The history matters because almost every “new” idea you’ll read about — agents, reasoning, alignment, even the fear of superintelligence — was proposed, tried, abandoned, and revived at least once before. Knowing the arc is how you tell a genuinely new idea from a rebrand.
Why it matters now
The machinery you use this week is old machinery. Every LLM answering your questions is trained by backpropagation (1986) on a transformer (2017) — that combination is what every frontier lab currently ships — and the translation, autocomplete, dictation, and photo-search features on your phone are the same lineage wearing different product names. None of that is a fast-moving frontier; it’s the settled floor everything else stands on.
What that buys you is a working filter for the news. The field has a strong habit of removing one bottleneck and immediately hitting the next, and each generation’s failure is what names the next generation’s idea. Once you can see that chain, “we removed a bottleneck” and “we changed everything” stop sounding alike — which is most of what you need to read an AI announcement carefully.
The short answer
AI history ≈ symbolic logic → statistical learning → scaled neural nets
Picture to keep: a pendulum between two workshops — one where people write the rules down by hand, one where machines infer the rules from piles of examples — swinging toward the second workshop every time hardware gets cheap enough to feed it. Where the pendulum image breaks: it isn’t symmetric. Each swing back toward learning has come with more compute than the last, so it’s less a pendulum than a ratchet.
For seventy years, AI alternated between two bets: encode human reasoning as explicit rules (symbolic AI), or learn patterns from data (statistical or connectionist AI). The standard account of why the second bet kept losing is that the compute and the datasets weren’t there; that’s the reading Sutton’s Bitter Lesson argues for, and it’s the one this post follows. It isn’t the only reading — the training algorithms and the benchmarks changed too, and untangling which mattered most is genuinely contested.
How it works
Read the timeline as a chain of failures. Each era does something real, hits a wall, and the wall names the next idea. Our French translation problem is the test case at every step.
1950 — the question gets made answerable. Alan Turing publishes Computing Machinery and Intelligence, asks “can machines think?”, and immediately swaps it for something testable: the imitation game. No machine could play it for decades, but the move — replace the philosophy with a behavioral test — is why the field is engineering and not metaphysics.
1956 — Dartmouth names it. A summer workshop coins artificial intelligence and sets the tone. The 1955 proposal is worth reading purely for the confidence: it budgets two months and ten people to make “a significant advance.” Early programs proved geometry theorems and played checkers.
1960s–70s — symbolic AI, and where it breaks. The bet: intelligence is symbol manipulation, so write knowledge as rules and run logic over them. This produced expert systems that were genuinely impressive inside narrow domains — MYCIN for bacterial infections, DENDRAL for chemistry.
Why it breaks: to translate our sentence, you’d need a rule saying trophies that don’t fit are too big. Then one for suitcases. Then for pianos, doorways, and every other object pair anyone might mention. The knowledge required is unbounded, and humans have to type all of it. Narrow domains hid this; open language exposed it.
Meanwhile — the underdog, and its ceiling. Frank Rosenblatt’s perceptron (1958) showed a single-layer network could learn a pattern from examples rather than be told it. Minsky and Papert’s 1969 book Perceptrons showed that a single layer can’t represent even XOR.
Why it breaks: the fix — stack more layers — was known, but nobody had a practical way to train the stack. Neural-net funding drained. The broader first AI winter of the mid-1970s had wider causes than this one book — the ALPAC report on machine translation and the UK’s Lighthill report are the usual culprits cited — but Perceptrons is the reason the connectionist branch specifically went quiet.
1986 — backpropagation makes depth trainable. Rumelhart, Hinton, and Williams popularize backpropagation, which computes how to adjust every weight in a deep network by pushing errors backward through the layers. In principle, Minsky’s objection is now answered.
Why it breaks: in practice the networks were tiny and slow, and simpler statistical methods beat them on the benchmarks that existed. A second winter arrived in the early 90s as expert systems also failed to scale.
1990s–2000s — statistics wins the practical fights. Spam filtering, search ranking, speech recognition, and translation all get good via classical machine learning. Statistical machine translation learns which phrases tend to correspond across languages, from piles of already-translated documents. It’s how that Translate button worked in its early years — and why it produced word-salad. Phrase statistics stitch together locally plausible fragments; the pronoun choice gets decided by whatever gendered form was more common in similar surrounding phrases, which is not the same as working out what “it” refers to. Sometimes that lands. It isn’t reasoning about trophies.
Why it breaks: these methods learn from features humans design. Somebody has to decide what to measure. That’s the hand-written-rules problem again, moved one level down.
2012 — AlexNet, and features stop being hand-designed. A deep neural network wins the ImageNet contest by a large margin (a top-5 error of 15.3% against 26.2% for the runner-up), trained on GPUs rather than CPUs. Its features were learned, not specified. Within a few years deep learning dominates vision and speech.
Why it breaks (for language): images are fixed-size grids. Sentences aren’t.
2014–2016 — sequence models, and the serial wall. RNNs and especially LSTMs become standard for text. They read a sentence left to right, carrying a hidden state, and translate from that state. The first version squeezed the entire source sentence into one fixed-size vector, which is a bad place to keep a trophy you’ll need to refer back to. That specific problem was fixed before the transformer: Bahdanau et al. added an attention mechanism to RNN translation in 2014, letting the decoder look back at every source position. Attention is not what the transformer invented.
Why it breaks: recurrence itself. Step n can’t be computed until step n−1 finishes, so training can’t exploit a GPU’s parallelism no matter how good the attention on top is. In an era where progress was starting to track compute spent, “cannot be parallelized” is close to fatal.
2017 — attention removes the wall. A Google paper, Attention Is All You Need, introduces the transformer — and, tellingly, demonstrates it on machine translation. Its move is to keep attention and delete the recurrence: every position reads its context directly rather than one step at a time, so a whole sentence trains in one shot. “It” can attend straight back to “trophy” and “suitcase” and weigh both — and, crucially, the training run can now saturate a GPU cluster. The title is the argument: attention was already there; recurrence turned out to be the part you could drop.
Why it breaks: the architecture doesn’t tell you what to train it on, or how much of it to build.
2018–2020 — GPT and scaling laws. OpenAI’s GPT-1, GPT-2, and GPT-3 show something quietly radical: train a transformer on enough text to just predict the next token, and capability improves predictably with size, data, and compute. GPT-3 (175 billion parameters, 2020) wrote essays and code with no task-specific training. The scaling laws made spending on a training run something you could forecast a return on. (That reading of why the money arrived is mine, not a documented decision process at any lab.) Notice what happened to our sentence along the way: nobody in this lineage set out to build a translation system. Translating the trophy sentence became a side effect of predicting text well enough.
Why it breaks: a next-token predictor is not an assistant. Show GPT-3 our sentence and ask for French, and one statistically reasonable continuation is another exam question about trophies, because that’s what such text tends to look like on the internet. Making the model treat the prompt as an instruction rather than a passage to continue is a separate problem.
2022 — RLHF, and the product leap. The fix is RLHF: fine-tune on demonstrations of following instructions, then shape behavior using human preference ratings. The technique wasn’t new that November: OpenAI had already published it as InstructGPT earlier in 2022. What ChatGPT added was the conversational framing — a model tuned for back-and-forth dialogue, behind a text box anyone could type into. Ask for French now and you get French. Not primarily a capability jump — a usability one. It broke containment; ChatGPT was widely reported as the fastest-growing consumer application to that point, though the much-quoted “100 million users in two months” figure came from a third-party estimate, not from OpenAI.
Why it breaks: a helpful answer is still just an answer. It can’t look anything up or check itself.
2023–2025 — tools, multimodality, agents. Models learn to see images (GPT-4V, Claude 3) and to call external functions. Now the trophy sentence can be a photo of a page, and the model can consult a glossary before committing to il or elle. Put an LLM in a loop where it decides what to call next, and you get the agent pattern — see agent harness.
Why it breaks: answering in one shot has a ceiling on problems that need working-out.
2024–2026 — reasoning and long horizons. OpenAI’s o-series and DeepSeek-R1 popularize models trained with reinforcement learning to produce a long chain of reasoning tokens before the final answer. (How much of that chain reflects the computation actually driving the answer is an open research question — treat “the model is thinking” as a description of the output format, not of the internals.) Context windows in some frontier models reach a million tokens. The live question stops being “can the model translate the sentence?” and becomes “can it stay coherent across a day of work?” Open-weights models have narrowed the gap to closed frontier ones on public benchmarks, though how much depends heavily on which benchmark you pick — I wouldn’t trust a single summary number.
So: why did the Translate button quietly get good? Because the pronoun in our sentence stopped being a lookup problem and became a side effect of a system trained to predict text in general — and that system only became trainable at that scale once you could drop recurrence and spend a data center on it.
You started with AI history ≈ symbolic logic → statistical learning → scaled neural nets. What does that arrow-diagram leave out? — that each arrow is a
failure boundary, not a change of fashion: rules ran out of rules, statistics
ran out of hand-designed features, recurrence ran out of parallelism. Read that
way, the sequence looks less like taste and more like a series of forced moves.
Check yourself
Before you go — the transformer paper was demonstrated on translation, and The Bitter Lesson argues that hand-built human knowledge keeps losing to raw scale. So why did symbolic AI dominate for thirty years if it was the losing bet?
Answer
Because “raw scale wins” is a claim about the limit, not about 1975. With the compute and data available then, a learned system couldn’t beat a rule-based one at anything useful — the perceptron’s ceiling was real and Minsky’s critique was correct on its own terms. Symbolic AI won for thirty years because it was genuinely the better engineering choice at that hardware budget. The lesson isn’t “they were fools”; it’s that the winner of an approach comparison can flip when an input like compute changes by many orders of magnitude.
And one more — suppose someone announces a new architecture that beats the transformer on quality per parameter, but has to process tokens strictly in order. Based on the pattern above, what’s your first question?
Answer
“How does it train?” That’s the wall recurrent models hit before 2017. A quality win per parameter is worth little if you can’t parallelize training across a GPU cluster, because you’ll be outscaled by a worse-per-parameter model that trains ten times faster on ten times the data. Architectures on this timeline mostly won on trainability at scale, not on elegance — which is also why “beats the transformer on a small benchmark” has been announced many times without sticking.
Famous related terms
- Symbolic AI —
symbolic AI = rules + logic engine— the “write down what you know” school; dominant 1956–~1990, still alive in formal verification. - Connectionism —
connectionism ≈ many tiny units learning together— the neural-net family, which lost the early debates and won the war. - AI winter —
AI winter ≈ funding collapse after over-promising— at least two big ones, roughly the mid-1970s and around 1990. - Perceptron —
perceptron = single layer + threshold function— the original learnable neural unit (Rosenblatt, 1958). - Backpropagation —
backprop = chain rule + errors propagated backward through layers— the algorithm that made depth trainable; it’s still what the models you use are trained with. - AlexNet —
AlexNet = deep CNN + GPU training + ImageNet (2012)— the starting gun for the deep learning era. - Transformer —
transformer ≈ stack of (attention + feed-forward) layers— the architecture under today’s frontier models. - Scaling laws —
scaling laws = loss falls as a power law in compute, data, params— why labs are willing to spend billions on a single training run. - RLHF —
RLHF = supervised fine-tune + reward model + RL loop— the step that turned a next-token predictor into something that feels like an assistant.
Going deeper
- Turing, Computing Machinery and Intelligence (1950) — read this to see how the field decided that “can machines think?” should be settled by behavior rather than definition.
- Rich Sutton, The Bitter Lesson (2019) — one page, and the best answer to “is there a pattern behind all these swings, or is it just fashion?”
- Vaswani et al., Attention Is All You Need (2017) — for the reader who wants to see exactly what replaced recurrence, in the original translation setting.