Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why JSON beat XML

XML had a standards body, schemas, namespaces, transformations, and a decade head start. JSON had curly braces and a JavaScript parser. Curly braces won.

Computer Science intro Apr 29, 2026 · updated Aug 25, 2026 · 8 min read

On this page

The picture version

Four pictures for a reader who has never wondered why the braces won. The prose below fills in the seams the pictures skip.

1 · The problem

The same user record, in the two formats.

the blessed answer <user id="42"> <name>phi</name> <admin>true</admin> <tags> <tag>ops</tag> <tag>writer</tag> </tags> </user> what everything returns now { "id": 42, "name": "phi", "admin": true, "tags": ["ops", "writer"] } The side that lost had the standards body, the vendors and the textbooks. that record is the running example — watch what happens to it in each format
Length is the least of it. The XML version can’t tell you whether id is a number or a string, blurs attributes against child elements, and has no native list syntax — so a one-tag user and a two-tag user deserialise to different shapes.

2 · The mismatch

Markup is for prose with annotations. This is a record with fields.

what markup is good at <p>This is <em>important</em>.</p> prose, with annotations sprinkled through it what engineers actually had { id: 42, admin: true } a record with five typed fields XML was a markup language pressed into service as a data format. that framing is this post’s argument rather than a settled verdict — but it explains the specific awkwardness above
Use markup on a typed record and you are forced to carry the type information out-of-band, in a schema, and to express structure verbosely. Neither is a flaw in XML — they are what happens when you use a document tool on a data problem.

3 · Why the braces won

JSON wasn’t designed to beat XML. It was already sitting there.

the object in the server’s memory the bytes on the wire {"id": 42} the object in the browser’s memory Same shape at all three points. Nothing to translate. numbers are numbers. booleans are booleans. one element or ten, a list is still a list. there is exactly one way to write “an object with these fields”, and the whole spec is about ten pages Not cleverness. Provenance.
Crockford has been clear that he discovered JSON rather than invented it — the syntax was already implicit in JavaScript, and he wrote it down and named it. XML 1.0 plus Namespaces, Schema and XPath run to hundreds of pages.

4 · Keep this card

The whole thing on one index card — and what it cost.

JSON = JavaScript object literal syntax + “that’s the whole spec” what winning cost, pushed out of the format and into convention: no schema — so every API reinvents one no dates — so they are strings, and time zones are a footgun no comments — which is half of why YAML and JSON5 exist no binary — base64 it and pay 33% numbers past 2⁵³ lose precision in a browser — which is why big IDs ship as strings
Picture to keep: XML is a shipping crate with a packing list, customs forms, and a stencilled label describing the crate itself. JSON is the object handed over unwrapped, in the same shape it had on the shelf — because the receiving warehouse already stores things in exactly that shape. Winning didn’t make the problems above disappear; it moved them somewhere each team re-solves them badly.

Why it exists

Open the developer tools on any site you use, click the Network tab, click any request. The response pane is full of curly braces: {"id": 42, "name": "phi"}. Every site. Every time. That uniformity is so complete it reads like a law of nature — but it’s the outcome of a fight that the other side was overwhelmingly favoured to win.

That little user record is the running example for this post; watch what happens to it in each format.

Rewind. It’s the early 2000s. You want to send structured data between two programs over HTTP. The official answer, blessed by the W3C, is XML. There are entire books about it. There are degrees in it. Enterprise vendors have built tooling empires around it: XSD for schemas, XSLT for transformations, XPath for queries, SOAP for RPC. Microsoft, IBM, Oracle, Sun all agree XML is the way.

So why, twenty years later, does basically every web API you touch return {"id": 42, "name": "phi"} instead of <user><id>42</id><name>phi</name></user>?

The short version: XML was designed to mark up documents, but engineers needed to ship data structures. JSON happened to be a literal serialization of the data structure JavaScript already used, so the browser could parse it for free. Everything else followed from that one fact.

Why it matters now

Almost every API a software engineer touches today speaks JSON: REST endpoints, GraphQL responses, log lines, config files for tools that swore they’d stay YAML-only, the request and response bodies of every LLM API, the tool-call schemas inside those LLM requests. When you debug a production incident in 2026, you’re reading JSON.

Understanding why it won — and what it gave up to win — explains a lot of weird corners: why JSON has no comments, why dates are strings, why every API has its own ad-hoc convention for null vs. missing, and why tool-call payloads still wrestle with the same problems SOAP wrestled with in 2003.

The short answer

JSON = JavaScript object literal syntax + "that's the whole spec"

Picture to keep: XML is a shipping crate with a packing list, customs forms, and a stencilled label describing the crate itself. JSON is the object handed over unwrapped, in the same shape it had on the shelf — because the receiving warehouse already stores things in exactly that shape.

JSON is the subset of JavaScript syntax you’d use to write a nested object literal — strings, numbers, booleans, null, arrays, objects — frozen and called a data format. It won because it was small enough to fit in your head, parseable by eval() in any browser shipped after 1996, and shaped exactly like the data structures programmers already had in memory.

How it works (and how XML works differently)

The attempt: use the document format for the data. XML was the blessed answer, so take our user record and mark it up.

<user id="42">
  <name>phi</name>
  <admin>true</admin>
  <tags>
    <tag>ops</tag>
    <tag>writer</tag>
  </tags>
</user>
{
  "id": 42,
  "name": "phi",
  "admin": true,
  "tags": ["ops", "writer"]
}

Why it breaks. The XML version is longer, but length is the least of it. It doesn’t actually tell you whether id is a number or a string — everything in XML is text, so you need an external schema just to read your own payload. It blurs attributes (id="42") and child elements (<name>phi</name>), which are syntactically different but semantically often the same, so every team invents its own convention. And tags has no native list syntax: you wrap the repeating elements in a parent and hope the consumer knows a one-element <tags> is still a list, not a scalar. That last one bites in production constantly — a naive XML-to-object mapper turns a one-tag user into a string and a two-tag user into an array.

The fix: serialize the data structure you already have. JSON sidesteps every one of these, not through cleverness but through provenance. Numbers are numbers. Booleans are booleans. Lists have brackets — one element or ten, it’s still an array. There’s exactly one way to express “an object with these fields.” A parser written from the spec fits in maybe 200 lines of C.

This is the deep reason XML lost: XML was a markup language pretending to be a data format. Markup is the right tool when you have prose with annotations sprinkled in (<p>This is <em>important</em>.</p>). It’s the wrong tool when you have a record with five typed fields, because it makes you encode the type information out-of-band and the structure verbosely.

The seams JSON left exposed

The fix has its own failures. JSON winning didn’t make the underlying problems disappear — it pushed them out of the format and into convention, where each team re-solves them badly.

The standard account is that Douglas Crockford specified JSON in the early 2000s — he registered the application/json media type and ran json.org — but he’s been clear he discovered it rather than invented it. The syntax was already implicit in JavaScript. He just wrote it down and gave it a name. RFC 8259 is the current spec; it’s about ten pages of actual content. XML 1.0 plus the Namespaces, Schema, and XPath specs run into the hundreds.

Nobody has published a clean measurement of when JSON volume overtook XML volume on the public web, and any single number for it deserves suspicion — adoption was gradual, format-by-format, between roughly 2006 (when major sites started offering JSON endpoints alongside XML) and the early 2010s (when REST-plus-JSON became the default new-API choice). The shift from SOAP to REST, and from desktop apps to single-page JavaScript apps, dragged the data format with it.

You started with JSON = JavaScript object literal syntax + "that's the whole spec". What did the {"id": 42, ...} payload add? — + it won by matching the receiver's memory layout, not by being better designed. Notice what that predicts: JSON is weakest exactly where the JavaScript object model is weak — no dates, no binary, no integers past 2⁵³ — and those are precisely the corners every API still papers over by hand.

Going deeper