Why JSON beat XML
XML had a standards body, schemas, namespaces, transformations, and a decade head start. JSON had curly braces and a JavaScript parser. Curly braces won.
On this page
The picture version
Four pictures for a reader who has never wondered why the braces won. The prose below fills in the seams the pictures skip.
1 · The problem
The same user record, in the two formats.
2 · The mismatch
Markup is for prose with annotations. This is a record with fields.
3 · Why the braces won
JSON wasn’t designed to beat XML. It was already sitting there.
4 · Keep this card
The whole thing on one index card — and what it cost.
Why it exists
Open the developer tools on any site you use, click the Network tab, click any
request. The response pane is full of curly braces: {"id": 42, "name": "phi"}.
Every site. Every time. That uniformity is so complete it reads like a law of
nature — but it’s the outcome of a fight that the other side was overwhelmingly
favoured to win.
That little user record is the running example for this post; watch what happens to it in each format.
Rewind. It’s the early 2000s. You want to send structured data between two programs over HTTP. The official answer, blessed by the W3C, is XML. There are entire books about it. There are degrees in it. Enterprise vendors have built tooling empires around it: XSD for schemas, XSLT for transformations, XPath for queries, SOAP for RPC. Microsoft, IBM, Oracle, Sun all agree XML is the way.
So why, twenty years later, does basically every web API you touch return {"id": 42, "name": "phi"} instead of <user><id>42</id><name>phi</name></user>?
The short version: XML was designed to mark up documents, but engineers needed to ship data structures. JSON happened to be a literal serialization of the data structure JavaScript already used, so the browser could parse it for free. Everything else followed from that one fact.
Why it matters now
Almost every API a software engineer touches today speaks JSON: REST endpoints, GraphQL responses, log lines, config files for tools that swore they’d stay YAML-only, the request and response bodies of every LLM API, the tool-call schemas inside those LLM requests. When you debug a production incident in 2026, you’re reading JSON.
Understanding why it won — and what it gave up to win — explains a lot of weird corners: why JSON has no comments, why dates are strings, why every API has its own ad-hoc convention for null vs. missing, and why tool-call payloads still wrestle with the same problems SOAP wrestled with in 2003.
The short answer
JSON = JavaScript object literal syntax + "that's the whole spec"
Picture to keep: XML is a shipping crate with a packing list, customs forms, and a stencilled label describing the crate itself. JSON is the object handed over unwrapped, in the same shape it had on the shelf — because the receiving warehouse already stores things in exactly that shape.
JSON is the subset of JavaScript syntax you’d use to write a nested object literal — strings, numbers, booleans, null, arrays, objects — frozen and called a data format. It won because it was small enough to fit in your head, parseable by eval() in any browser shipped after 1996, and shaped exactly like the data structures programmers already had in memory.
How it works (and how XML works differently)
The attempt: use the document format for the data. XML was the blessed answer, so take our user record and mark it up.
<user id="42">
<name>phi</name>
<admin>true</admin>
<tags>
<tag>ops</tag>
<tag>writer</tag>
</tags>
</user>
{
"id": 42,
"name": "phi",
"admin": true,
"tags": ["ops", "writer"]
}
Why it breaks. The XML version is longer, but length is the least of it. It doesn’t actually tell you whether id is a number or a string — everything in XML is text, so you need an external schema just to read your own payload. It blurs attributes (id="42") and child elements (<name>phi</name>), which are syntactically different but semantically often the same, so every team invents its own convention. And tags has no native list syntax: you wrap the repeating elements in a parent and hope the consumer knows a one-element <tags> is still a list, not a scalar. That last one bites in production constantly — a naive XML-to-object mapper turns a one-tag user into a string and a two-tag user into an array.
The fix: serialize the data structure you already have. JSON sidesteps every one of these, not through cleverness but through provenance. Numbers are numbers. Booleans are booleans. Lists have brackets — one element or ten, it’s still an array. There’s exactly one way to express “an object with these fields.” A parser written from the spec fits in maybe 200 lines of C.
This is the deep reason XML lost: XML was a markup language pretending to be a data format. Markup is the right tool when you have prose with annotations sprinkled in (<p>This is <em>important</em>.</p>). It’s the wrong tool when you have a record with five typed fields, because it makes you encode the type information out-of-band and the structure verbosely.
The seams JSON left exposed
The fix has its own failures. JSON winning didn’t make the underlying problems disappear — it pushed them out of the format and into convention, where each team re-solves them badly.
- No schema. JSON itself has no types beyond the primitives. Every API reinvents the wheel: JSON Schema, OpenAPI, Protobuf-as-JSON, ad-hoc TypeScript types. XML at least converged on one blessed answer — though not immediately: XML 1.0 became a Recommendation in 1998 and XML Schema only in 2001. We arguably have more schema fragmentation now than we did then.
- No dates. JSON has no date type, so dates are strings — usually ISO 8601, but not always. Time zones are a footgun.
- No comments. JSON has no comment syntax. The reason usually given — that Crockford removed them so nobody could smuggle parsing directives into them — is repeated everywhere but hard to pin to a primary source, so take the explanation loosely and the absence literally. Config files have suffered for it ever since, which is half of why YAML and JSON5 exist.
- Number weirdness. JSON’s grammar puts no length limit on a number, but the spec doesn’t promise you arbitrary precision either: RFC 8259 says implementations may set their own range and precision limits, and that you get good interoperability by expecting no more than an IEEE 754 double gives you. JavaScript numbers are doubles, so a 64-bit integer ID round-tripped through a browser silently loses precision past 2^53. Every modern API that uses big IDs ships them as strings to dodge this.
- No binary. Want to send bytes? Base64 them into a string and pay the 33% overhead, or use a different format entirely.
The standard account is that Douglas Crockford specified JSON in the early 2000s — he registered the application/json media type and ran json.org — but he’s been clear he discovered it rather than invented it. The syntax was already implicit in JavaScript. He just wrote it down and gave it a name. RFC 8259 is the current spec; it’s about ten pages of actual content. XML 1.0 plus the Namespaces, Schema, and XPath specs run into the hundreds.
Nobody has published a clean measurement of when JSON volume overtook XML volume on the public web, and any single number for it deserves suspicion — adoption was gradual, format-by-format, between roughly 2006 (when major sites started offering JSON endpoints alongside XML) and the early 2010s (when REST-plus-JSON became the default new-API choice). The shift from SOAP to REST, and from desktop apps to single-page JavaScript apps, dragged the data format with it.
You started with JSON = JavaScript object literal syntax + "that's the whole spec". What did the {"id": 42, ...} payload add? — + it won by matching the receiver's memory layout, not by being better designed. Notice what that predicts: JSON is weakest exactly where the JavaScript object model is weak — no dates, no binary, no integers past 2⁵³ — and those are precisely the corners every API still papers over by hand.
Famous related terms
- XML —
XML = SGML simplified for the web + namespaces + schema layer— still dominant in document-shaped domains: SVG, Office files, RSS, legacy enterprise integration. Not dead, just specialized. - YAML —
YAML = JSON superset + significant whitespace + comments + anchors— what JSON should have been for config, with all the indentation pain that implies. - Protobuf —
protobuf = schema-first binary format + generated code per language— what you reach for when JSON’s text overhead and lack of typing finally hurts enough. - JSON Schema —
JSON Schema = JSON document that validates other JSON documents— the schema layer JSON didn’t ship with, retrofitted later. The thing OpenAPI and LLM tool-call specs both lean on.
Going deeper
- RFC 8259 — the primary source, and the fastest way to answer “is that really the entire specification?” Yes: about ten pages of actual content.
- Douglas Crockford’s JSON: The Fat-Free Alternative to XML (2006) — read this for the argument as it was made at the time, before the outcome was obvious.
- Tim Bray’s blog posts from the mid-2000s — the rabbit hole for the other side: Bray co-edited the XML 1.0 spec and later wrote candidly about where XML got misused. Worth reading in his own words rather than paraphrased.