<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Why-How</title><description>Notes on famous terms in technology and science — why first, then how.</description><link>https://til.phipham141.dev/</link><language>en</language><item><title>Why database isolation levels exist</title><link>https://til.phipham141.dev/posts/data/isolation-levels/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/isolation-levels/</guid><description>Only the top level is simply correct. The others exist because correctness costs money — and the standard&apos;s list of what can go wrong turned out to be one entry short.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>How does an AI &apos;see&apos; a video?</title><link>https://til.phipham141.dev/posts/ai-ml/how-ai-sees-video/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/how-ai-sees-video/</guid><description>Upload an hour-long video to Gemini and ask what happens at minute 40 — it answers. But no model &apos;watches&apos; anything. It reads a flipbook, and the flipbook is missing most of the pages.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Hardening a cloud server</title><link>https://til.phipham141.dev/posts/security/hardening-a-cloud-server/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/hardening-a-cloud-server/</guid><description>Spin up a fresh VPS, wait an hour, and the auth log already has thousands of brute-force attempts from across the internet. Every server-hardening guide says roughly the same things — here&apos;s what each one actually stops, and where the rules are theater.</description><pubDate>Mon, 25 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>How end-to-end encryption works</title><link>https://til.phipham141.dev/posts/security/end-to-end-encryption/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/end-to-end-encryption/</guid><description>Open WhatsApp and a banner tells you Meta can&apos;t read your messages. That claim sits on a specific protocol — Diffie–Hellman key agreement plus a &apos;double ratchet&apos; that changes the key on every message. Here&apos;s the shape of it.</description><pubDate>Sun, 24 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>How does an LLM &apos;see&apos; an image?</title><link>https://til.phipham141.dev/posts/ai-ml/how-llms-process-images/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/how-llms-process-images/</guid><description>You paste a screenshot into ChatGPT and it reads the text, describes the scene, answers questions. But the model only ever predicts text tokens — so how does a picture get into it at all?</description><pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What are &apos;weights&apos; in an LLM?</title><link>https://til.phipham141.dev/posts/ai-ml/weights/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/weights/</guid><description>When Meta releases &apos;open weights&apos; for Llama, what&apos;s actually in that file? A giant table of numbers and nothing else — so how does a pile of numbers know things?</description><pubDate>Sat, 16 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why compression works at all</title><link>https://til.phipham141.dev/posts/computer-science/why-compression-works/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/why-compression-works/</guid><description>Zip a photo and it shrinks; zip the zip and it doesn&apos;t. Compression isn&apos;t magic — it only ever exploits the patterns that were already there.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why CPUs have three levels of cache</title><link>https://til.phipham141.dev/posts/science/why-cpus-have-cache-levels/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/why-cpus-have-cache-levels/</guid><description>Look at a CPU die shot and you&apos;ll find more area spent on memory than on math — and that memory is split into L1, L2, and L3. The split exists because no single memory technology is both big and fast, so the chip builds a ladder out of several instead.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why deadlocks need four conditions</title><link>https://til.phipham141.dev/posts/systems/why-deadlocks-need-four-conditions/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/why-deadlocks-need-four-conditions/</guid><description>A deadlock feels like bad luck, but it can only happen when four specific conditions all hold at once — and breaking any one makes it impossible.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why garbage collectors pause your program</title><link>https://til.phipham141.dev/posts/systems/why-gc-pauses-happen/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/why-gc-pauses-happen/</guid><description>A tracing collector can&apos;t safely move or free an object while your code is mid-read. Freezing every application thread is the obviously correct answer — and generations, write barriers and concurrent marking are all ways of making that freeze shorter without giving up what it buys.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>What is tool use (a.k.a. function calling)?</title><link>https://til.phipham141.dev/posts/ai-ml/tool-use/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/tool-use/</guid><description>A model that only emits text somehow ends up booking your flight. The trick isn&apos;t in the weights — it&apos;s in the contract between model, harness, and your code.</description><pubDate>Thu, 07 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What does &apos;X parameters&apos; mean in an LLM?</title><link>https://til.phipham141.dev/posts/ai-ml/llm-parameters/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/llm-parameters/</guid><description>Llama 3.1 70B, DeepSeek-V3 671B, Phi-4 14B — what is that number actually counting, and why is it the headline figure on every model release?</description><pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why model merging works at all</title><link>https://til.phipham141.dev/posts/ai-ml/why-model-merging-works/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-model-merging-works/</guid><description>Take two fine-tunes of the same model, average their weights element-wise, and you often get a model better than either parent. Naively, this shouldn&apos;t work — neural net loss surfaces are wildly non-convex. The reason it works tells you something deep about where fine-tuning actually lives.</description><pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why EUV lithography blasts tin droplets with lasers</title><link>https://til.phipham141.dev/posts/science/euv-lithography/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/euv-lithography/</guid><description>The most demanding layers of leading-edge chips are patterned by a machine that, tens of thousands of times per second, vaporizes a falling droplet of molten tin with a high-power laser. The setup is absurd — and nothing else makes 13.5nm light in a production tool.</description><pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why Spectre still isn&apos;t fully patched</title><link>https://til.phipham141.dev/posts/security/why-spectre-isnt-fully-patched/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/why-spectre-isnt-fully-patched/</guid><description>Eight years after disclosure, new Spectre-class vulnerabilities keep landing. The reason isn&apos;t sloppy patching — the speculation being exploited is what makes modern CPUs fast, and the list of channels it leaks through has no end.</description><pubDate>Mon, 04 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>How can I tell when an LLM is making the answer up?</title><link>https://til.phipham141.dev/posts/ai-ml/how-to-spot-hallucinations/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/how-to-spot-hallucinations/</guid><description>True answers and fabricated ones come out of the same pipe, in the same tone. There&apos;s no red light. But there are seams — places hallucinations cluster, shapes they tend to take, tells you can learn to read.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why dropout disappeared from modern LLMs</title><link>https://til.phipham141.dev/posts/ai-ml/why-dropout-disappeared-from-llms/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-dropout-disappeared-from-llms/</guid><description>Dropout was the regularization workhorse of the deep-learning era. Frontier LLM pretraining quietly stopped using it. The reason isn&apos;t that dropout broke — it&apos;s that the problem dropout solved stopped being the problem.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why LLMs can&apos;t count the r&apos;s in &apos;strawberry&apos;</title><link>https://til.phipham141.dev/posts/ai-ml/why-llms-cant-count-letters/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-llms-cant-count-letters/</guid><description>A model that can write a sonnet stumbles on a question a five-year-old gets right. The reason isn&apos;t intelligence — it&apos;s that the model never sees the letters.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why reward hacking is RLHF&apos;s hardest problem</title><link>https://til.phipham141.dev/posts/ai-ml/why-reward-hacking-happens/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-reward-hacking-happens/</guid><description>You can&apos;t write down a loss function for &apos;be helpful,&apos; so you train a model to predict it — and then a much bigger model spends all its optimization pressure looking for holes in that prediction. That gap is reward hacking, and it doesn&apos;t go away with scale.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why synthetic data works for modern LLM training</title><link>https://til.phipham141.dev/posts/ai-ml/why-synthetic-data-works/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-synthetic-data-works/</guid><description>The open web ran out of high-quality text years before frontier models stopped getting better. The new training signal didn&apos;t come from a fresh internet — it came from models writing for models, with filters in front.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why floating-point addition isn&apos;t associative</title><link>https://til.phipham141.dev/posts/computer-science/floating-point-not-associative/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/floating-point-not-associative/</guid><description>Schoolroom math says (a + b) + c equals a + (b + c). On a real computer it doesn&apos;t, and that one fact ripples out into nondeterministic GPU reductions, irreproducible training runs, and LLM outputs that aren&apos;t bit-stable across hardware.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why ECDSA nonce reuse leaks the private key</title><link>https://til.phipham141.dev/posts/security/ecdsa-nonce-reuse/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/ecdsa-nonce-reuse/</guid><description>ECDSA needs a fresh random number for every signature. Use the same one twice and anyone watching can recover the private key with two lines of algebra — which is exactly how the PS3&apos;s master key fell out.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why &apos;harvest now, decrypt later&apos; is driving post-quantum crypto adoption</title><link>https://til.phipham141.dev/posts/security/why-post-quantum-crypto-matters/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/why-post-quantum-crypto-matters/</guid><description>A sufficiently large quantum computer doesn&apos;t exist yet. Encrypted traffic from 2018 might already be sitting on a tape, waiting for one. That asymmetry — encrypt now, decrypt later — means the damage starts when the recording happens, not when the machine arrives.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why supply-chain attacks dominate the JavaScript ecosystem</title><link>https://til.phipham141.dev/posts/security/why-supply-chain-attacks-dominate-npm/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/why-supply-chain-attacks-dominate-npm/</guid><description>A small npm install pulls in a thousand-odd packages from hundreds of strangers, and some of that code runs before you type anything. JavaScript&apos;s deep, trusting dependency graph is the attack surface, and every step of the attack is a feature someone shipped on purpose.</description><pubDate>Sat, 02 May 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>How does League of Legends keep ten players in sync at low latency?</title><link>https://til.phipham141.dev/posts/networking/online-game-netcode/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/online-game-netcode/</guid><description>Ten strangers on ten different ISPs share one world that has to feel instantaneous. The trick is that nobody&apos;s screen shows exactly the same thing — and that&apos;s the feature, not the bug.</description><pubDate>Fri, 01 May 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>10 famous AI-ML terms</title><link>https://til.phipham141.dev/posts/ai-ml/famous-ai-ml-terms/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/famous-ai-ml-terms/</guid><description>The vocabulary you keep hearing on every podcast — neural network, transformer, RLHF, RAG — compressed to one line each, then unpacked.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is attention (in transformers)?</title><link>https://til.phipham141.dev/posts/ai-ml/attention/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/attention/</guid><description>Every token in a sequence gets to peek at every other token and decide which ones matter. That trick is the engine inside every modern LLM.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is harness engineering?</title><link>https://til.phipham141.dev/posts/ai-ml/harness-engineering/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/harness-engineering/</guid><description>Most of the work that turns a frontier model into a reliable product happens around the model, not inside it. Harness engineering is the name for that work.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>How does an AI model decide what to say?</title><link>https://til.phipham141.dev/posts/ai-ml/how-models-decide-answers/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/how-models-decide-answers/</guid><description>It looks like one big choice — you type a question, you get an answer. Underneath it&apos;s thousands of tiny choices, made one token at a time, with no plan and no rewind.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is a neural network?</title><link>https://til.phipham141.dev/posts/ai-ml/neural-network/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/neural-network/</guid><description>A pile of multiplications and a &apos;how wrong was I?&apos; signal — somehow, when you stack enough of them, the thing learns to read, see, and play chess.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is a transformer?</title><link>https://til.phipham141.dev/posts/ai-ml/transformer/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/transformer/</guid><description>The neural network architecture behind essentially every modern LLM — and the one big idea that made it work: drop recurrence, let every token look at every other token directly.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do attention sinks exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-attention-sinks-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-attention-sinks-exist/</guid><description>Trained transformers funnel a startling fraction of their attention onto the very first token — a token that&apos;s usually semantically meaningless. The pattern looks like a bug, behaves like a feature, and falls out cleanly from one constraint in the softmax.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why FlashAttention was a breakthrough</title><link>https://til.phipham141.dev/posts/ai-ml/why-flash-attention-was-a-breakthrough/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-flash-attention-was-a-breakthrough/</guid><description>Same math, same exact outputs, same asymptotic compute — and yet it made attention several times faster and unlocked long context. The trick was noticing attention was a memory problem, not a compute problem.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why FP8 training is stable</title><link>https://til.phipham141.dev/posts/ai-ml/why-fp8-training-is-stable/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-fp8-training-is-stable/</guid><description>FP8 has only 256 representable values. Training a frontier model in it sounds insane — and it almost is. Here&apos;s the trick that makes it work.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why grouped-query attention exists</title><link>https://til.phipham141.dev/posts/ai-ml/why-grouped-query-attention-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-grouped-query-attention-exists/</guid><description>Multi-head attention is a memory-bandwidth disaster at decode time. GQA keeps most of the quality and throws away most of the bandwidth bill.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why MLA replaced MHA</title><link>https://til.phipham141.dev/posts/ai-ml/why-mla-replaced-mha/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-mla-replaced-mha/</guid><description>DeepSeek-V2 cut its KV cache by 93% by attacking the bottleneck differently than GQA — and in their own matched ablation it scored higher, not lower.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does PagedAttention exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-paged-attention-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-paged-attention-exists/</guid><description>Naive KV-cache allocation reserves a contiguous slab for the worst-case sequence length, then watches 60–80% of it sit unused. PagedAttention asks: what if we treated GPU memory the way an operating system treats RAM?</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why RoPE replaced sinusoidal positional encoding</title><link>https://til.phipham141.dev/posts/ai-ml/why-rope-replaced-sinusoidal/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-rope-replaced-sinusoidal/</guid><description>The original transformer added a fixed sine/cosine vector to each token. Almost no frontier model does that anymore. RoPE rotates queries and keys instead — and that one structural change is what made long context tractable.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why AI runs away in verifiable domains</title><link>https://til.phipham141.dev/posts/ai-ml/why-verifiable-domains-run-away/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-verifiable-domains-run-away/</guid><description>AI is getting superhuman fastest at things a computer can grade — math, code, formal proofs — and dragging behind on things it can&apos;t. The reason isn&apos;t that those domains are &apos;easier.&apos; It&apos;s that training has a feedback step, and feedback needs a verifier.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is a hash function?</title><link>https://til.phipham141.dev/posts/computer-science/hash-function/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/hash-function/</guid><description>A deterministic shrinker that turns any blob of bytes into a fixed-size fingerprint — the same primitive that powers hash tables, Git commits, and password storage.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>What is TCP?</title><link>https://til.phipham141.dev/posts/networking/tcp/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/tcp/</guid><description>IP delivers packets best-effort — they can vanish, duplicate, or arrive out of order. Almost every program wants a clean stream of bytes instead. TCP is the layer that turns one into the other.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>What is TLS?</title><link>https://til.phipham141.dev/posts/networking/tls/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/tls/</guid><description>TCP gives you a reliable byte stream that any router along the path can read and modify. TLS is the layer that wraps that stream so you get confidentiality, integrity, and proof of who&apos;s on the other end.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>What is public-key cryptography?</title><link>https://til.phipham141.dev/posts/security/public-key-crypto/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/public-key-crypto/</guid><description>Until 1976, published cryptography required both sides to already share a secret. Public-key crypto broke that chicken-and-egg problem and quietly became the substrate of the modern internet.</description><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>A short history of AI, from Turing to today&apos;s LLMs</title><link>https://til.phipham141.dev/posts/ai-ml/ai-history/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/ai-history/</guid><description>Seventy years of trying to make machines think — and how a single architecture from 2017 finally cashed the check that 1950s AI wrote.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does chain-of-thought prompting work?</title><link>https://til.phipham141.dev/posts/ai-ml/chain-of-thought/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/chain-of-thought/</guid><description>Adding &apos;let&apos;s think step by step&apos; to a prompt makes models measurably better at hard problems. Nobody fully agrees on why, and the wrong story will mislead you about how to use it.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do LLMs hallucinate confidently instead of saying &apos;I don&apos;t know&apos;?</title><link>https://til.phipham141.dev/posts/ai-ml/hallucination/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/hallucination/</guid><description>The model isn&apos;t lying. It was never trained to know when to stop talking.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do embeddings exist?</title><link>https://til.phipham141.dev/posts/ai-ml/embeddings/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/embeddings/</guid><description>Computers want numbers, but you also want &apos;cat&apos; and &apos;kitten&apos; to live next to each other. Embeddings are the trick that makes both true at once.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is an agent harness?</title><link>https://til.phipham141.dev/posts/ai-ml/agent-harness/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/agent-harness/</guid><description>The loop and scaffolding around a language model that turns &apos;a thing that emits tokens&apos; into &apos;a thing that does work in the world.&apos;</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why is the KV cache a thing?</title><link>https://til.phipham141.dev/posts/ai-ml/kv-cache/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/kv-cache/</guid><description>The model has to read your whole prompt every time it picks a token. Why doesn&apos;t it choke? Because of a quiet trick almost nobody mentions in the docs.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does in-context learning work?</title><link>https://til.phipham141.dev/posts/ai-ml/in-context-learning/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/in-context-learning/</guid><description>You paste three examples into a prompt and the model suddenly does the task. Nothing got trained. So what just happened?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>What is an LLM?</title><link>https://til.phipham141.dev/posts/ai-ml/llm/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/llm/</guid><description>A neural network trained to predict the next token of text — and why that simple goal scaled into something that feels like reasoning.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does MCP exist?</title><link>https://til.phipham141.dev/posts/ai-ml/mcp/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/mcp/</guid><description>Every AI app was reinventing the same plumbing to talk to the same tools. MCP is the standard that turns an M×N integration mess into M+N.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does GPU memory bandwidth matter more than FLOPS for LLM inference?</title><link>https://til.phipham141.dev/posts/ai-ml/memory-bandwidth/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/memory-bandwidth/</guid><description>You bought the GPU for the teraflops. At inference time, almost none of them are doing anything. The bottleneck is moving the weights, not multiplying them.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>RAG: why retrieval didn&apos;t die when context windows got huge</title><link>https://til.phipham141.dev/posts/ai-ml/rag/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/rag/</guid><description>Long context windows were supposed to kill retrieval-augmented generation. They didn&apos;t. Here&apos;s why the bottleneck moved instead of disappearing.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does temperature exist as a knob?</title><link>https://til.phipham141.dev/posts/ai-ml/temperature/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/temperature/</guid><description>If the model knows the right answer, why is there a dial that asks it to be wrong on purpose?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do small models exist?</title><link>https://til.phipham141.dev/posts/ai-ml/small-models/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/small-models/</guid><description>If bigger models always benchmark better, why does anyone ship a 3B model? The answer is mostly about latency, cost, and the place the model has to live.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does tokenization exist?</title><link>https://til.phipham141.dev/posts/ai-ml/tokenization/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/tokenization/</guid><description>Computers can already read bytes. So why do language models insist on chopping text into these weird half-words first?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why Adam beat plain SGD for LLMs</title><link>https://til.phipham141.dev/posts/ai-ml/why-adam-beat-sgd-for-llms/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-adam-beat-sgd-for-llms/</guid><description>Vision models are mostly trained with SGD + momentum. Transformers are almost always trained with Adam or AdamW. Why did one optimizer win one regime and lose the other?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why vector search is approximate on purpose</title><link>https://til.phipham141.dev/posts/ai-ml/why-ann-not-exact/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-ann-not-exact/</guid><description>Exact nearest-neighbor search exists, works, and is correct. At scale, the AI-era retrieval stack quietly walks away from it. The reason is more interesting than &apos;it&apos;s faster.&apos;</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why is attention quadratic?</title><link>https://til.phipham141.dev/posts/ai-ml/why-attention-is-quadratic/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-attention-is-quadratic/</guid><description>Doubling the context length makes attention 4× more expensive, not 2×. That single fact shapes every trade-off in modern LLM serving — and explains what FlashAttention actually changed (it&apos;s not what most people think).</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why agents fall apart over long horizons</title><link>https://til.phipham141.dev/posts/ai-ml/why-agents-fail-at-long-horizons/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-agents-fail-at-long-horizons/</guid><description>Your agent solves any single step beautifully. Run it for fifty steps and it falls off a cliff. The math behind that cliff is older than LLMs, but a newer twist makes it worse.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why beam search died for LLMs</title><link>https://til.phipham141.dev/posts/ai-ml/why-beam-search-died-for-llms/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-beam-search-died-for-llms/</guid><description>Beam search was the default way to decode neural sequence models for years. Then chatbots arrived and quietly stopped using it. The reason is stranger than &apos;sampling is more creative.&apos;</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does continuous batching exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-continuous-batching-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-continuous-batching-exists/</guid><description>Static batching works fine for image classifiers and breaks immediately for LLMs. The problem isn&apos;t the batch — it&apos;s that generation lengths vary, and the slowest sequence holds the GPU hostage.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why image generation went diffusion, not autoregressive</title><link>https://til.phipham141.dev/posts/ai-ml/why-diffusion-models-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-diffusion-models-exist/</guid><description>LLMs are autoregressive: predict the next token. Image models could have been the same — predict the next pixel. Almost none of the dominant ones are. Here&apos;s why the field walked away from that approach.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why model distillation exists</title><link>https://til.phipham141.dev/posts/ai-ml/why-distillation-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-distillation-exists/</guid><description>A small model trained on a big model&apos;s outputs often beats the same small model trained on the original labels. That shouldn&apos;t be obvious — and the reason it works is the actually interesting part.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why is fine-tuning so cheap compared to pretraining?</title><link>https://til.phipham141.dev/posts/ai-ml/why-fine-tuning-is-cheap/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-fine-tuning-is-cheap/</guid><description>Pretraining a frontier model costs tens of millions of dollars. Fine-tuning the same model on your data can cost less than a pizza. Why the four-orders-of-magnitude gap?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why GPU kernels are still hand-tuned</title><link>https://til.phipham141.dev/posts/ai-ml/why-gpu-kernels-are-hand-tuned/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-gpu-kernels-are-hand-tuned/</guid><description>A modern GPU can do tens of teraflops of matrix math. A naive, correct implementation of the same math leaves most of that on the floor. Here&apos;s why moving the bytes — not doing the FLOPs — is the actual job.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why LayerNorm (and RMSNorm) exist</title><link>https://til.phipham141.dev/posts/ai-ml/why-layernorm-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-layernorm-exists/</guid><description>Every transformer block has a normalization step. Pull it out and training falls apart in the first thousand steps. Why is this tiny operation load-bearing?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why is evaluating an LLM so much harder than testing normal software?</title><link>https://til.phipham141.dev/posts/ai-ml/why-llm-eval-is-hard/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-llm-eval-is-hard/</guid><description>Unit tests pass or fail. LLM outputs don&apos;t. The hard part isn&apos;t running the eval — it&apos;s deciding what &apos;correct&apos; even means when there are a million right answers.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why long-context models still get lost in the middle</title><link>https://til.phipham141.dev/posts/ai-ml/why-lost-in-the-middle/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-lost-in-the-middle/</guid><description>Your model has a 1M token context window. It can recall the first paragraph perfectly. It can recall the last paragraph perfectly. The thing in the middle? Coin flip. This is not a bug — it&apos;s what happens when you ask a model trained one way to behave a different way.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why LoRA exists</title><link>https://til.phipham141.dev/posts/ai-ml/why-lora-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-lora-exists/</guid><description>Full fine-tuning a 70B model means storing optimizer state for 70 billion weights. LoRA trains under 1% of the parameters and, on the tasks people have tested, often matches the result. The trick is a hypothesis about the shape of the update.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does mixture-of-experts exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-mixture-of-experts-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-mixture-of-experts-exists/</guid><description>A 671B-parameter model whose per-token compute is closer to a 37B one. The trick isn&apos;t compression — it&apos;s that most of the weights sit out most of the time.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does predicting the next token end up doing reasoning?</title><link>https://til.phipham141.dev/posts/ai-ml/why-next-token-prediction-generalizes/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-next-token-prediction-generalizes/</guid><description>An LLM is trained on one objective: guess the next token. From that one task, you get translation, code, arithmetic, and arguments. Why is autocomplete this powerful?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does prompt caching exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-prompt-caching-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-prompt-caching-exists/</guid><description>Your agent sends the same 50,000-token system prompt on every turn. Providers charge a fraction of the usual rate when they recognize it — not out of generosity, but because they stopped doing the work.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do positional encodings exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-positional-encodings-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-positional-encodings-exist/</guid><description>A transformer cannot tell &apos;dog bites man&apos; from &apos;man bites dog&apos; on its own. The attention math is symmetric in token order — until you bolt on a position signal. Every modern LLM does, and the choice of how shapes long-context behavior more than people realize.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why quantization works</title><link>https://til.phipham141.dev/posts/ai-ml/why-quantization-works/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-quantization-works/</guid><description>Stuffing a 70-billion-parameter model into 4-bit weights sounds like it should ruin it. It mostly doesn&apos;t — and the reason is more about how the model gets used at inference than about the math of rounding.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why reasoning models exist</title><link>https://til.phipham141.dev/posts/ai-ml/why-reasoning-models-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-reasoning-models-exist/</guid><description>Why we suddenly have a separate class of LLMs that &apos;think before answering&apos; — and what changed to make spending compute at inference, not training, the new lever.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why RLHF exists</title><link>https://til.phipham141.dev/posts/ai-ml/why-rlhf-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-rlhf-exists/</guid><description>A pretrained language model knows everything and answers nothing. RLHF exists because the gap between &apos;predict the next token&apos; and &apos;do what the user asked&apos; is wider than prompt engineering can paper over.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why do scaling laws exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-scaling-laws-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-scaling-laws-exist/</guid><description>Bigger model, more data, more compute — and the loss falls along a straight line on a log-log plot for seven orders of magnitude. Nobody fully knows why that line is so straight.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why does speculative decoding exist?</title><link>https://til.phipham141.dev/posts/ai-ml/why-speculative-decoding-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-speculative-decoding-exists/</guid><description>A small fast model guesses, a big slow model checks. Somehow you get the big model&apos;s exact output, faster. The trick isn&apos;t cleverness — it&apos;s that your GPU was already sitting idle.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why SwiGLU replaced ReLU in transformers</title><link>https://til.phipham141.dev/posts/ai-ml/why-swiglu-replaced-relu/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-swiglu-replaced-relu/</guid><description>Modern LLMs ditched the simplest activation function in deep learning for a multiplicative gate nobody can fully explain. Here&apos;s why.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why is structured output so hard?</title><link>https://til.phipham141.dev/posts/ai-ml/why-structured-output-is-hard/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-structured-output-is-hard/</guid><description>You ask the model for JSON. Sometimes it gives you a trailing comma. Sometimes a markdown fence. Sometimes prose. Why is this still a problem?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why VRAM is the bottleneck for LLM serving</title><link>https://til.phipham141.dev/posts/ai-ml/why-vram-is-the-bottleneck/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-vram-is-the-bottleneck/</guid><description>It&apos;s not FLOPS, it&apos;s not network, it&apos;s not the CPU. The thing that decides whether your model fits and how many users you can serve is a number printed on the GPU&apos;s spec sheet — and three things fight to consume it.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why bf16 won the training format wars</title><link>https://til.phipham141.dev/posts/computer-science/bf16-vs-fp16/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/bf16-vs-fp16/</guid><description>Half-precision floats came in two flavors: fp16, which had been around for years, and bf16, which kept fp32&apos;s exponent and threw away mantissa bits. The less-precise format won. Here&apos;s why that&apos;s not a typo.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why isn&apos;t temperature 0 actually deterministic?</title><link>https://til.phipham141.dev/posts/ai-ml/why-temperature-zero-isnt-deterministic/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/ai-ml/why-temperature-zero-isnt-deterministic/</guid><description>You set temperature to 0, send the same prompt twice, get two different answers. The math says argmax is a function. The hardware disagrees.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>AI &amp; ML</category></item><item><title>Why networks are big-endian but your CPU is little-endian</title><link>https://til.phipham141.dev/posts/computer-science/endianness/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/endianness/</guid><description>Two halves of the same machine disagree on which end of a number comes first. The split is older than you, and it&apos;s never going away.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why Git stores snapshots, not diffs</title><link>https://til.phipham141.dev/posts/computer-science/git-snapshots/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/git-snapshots/</guid><description>Git&apos;s reputation says &apos;version control = diffs.&apos; Git&apos;s actual model says &apos;version control = snapshots, hashed.&apos; That swap is the whole reason Git feels different.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why UTF-8 won</title><link>https://til.phipham141.dev/posts/computer-science/utf-8/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/utf-8/</guid><description>Unicode could have been a fixed 4-byte-per-character encoding. Instead, the web runs on a variable-width hack — and that hack is why everything still works.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>What is a hash table?</title><link>https://til.phipham141.dev/posts/computer-science/hash-table/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/hash-table/</guid><description>A data structure that lets you find, insert, and delete by key in roughly constant time — the workhorse behind dictionaries, sets, and most fast lookup in modern code.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why WebAssembly exists</title><link>https://til.phipham141.dev/posts/computer-science/why-webassembly-exists/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/why-webassembly-exists/</guid><description>JavaScript already runs everywhere. So why did browser vendors agree to ship a second, lower-level execution target — and why is it now showing up in CDNs, plugin systems, and AI runtimes?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why B-trees still dominate database indexes</title><link>https://til.phipham141.dev/posts/data/b-tree-indexes/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/b-tree-indexes/</guid><description>Disks read in blocks, not bytes. B-trees were designed around that one fact — and decades later, even on SSDs, no one has dethroned them.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why Bloom filters exist</title><link>https://til.phipham141.dev/posts/data/bloom-filters/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/bloom-filters/</guid><description>A data structure that answers &quot;have I seen this?&quot; with &quot;definitely no&quot; or &quot;maybe&quot; — and saves enormous amounts of work by being wrong on purpose.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why columnar storage won for analytics</title><link>https://til.phipham141.dev/posts/data/columnar-storage/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/columnar-storage/</guid><description>Row stores read every column to answer a one-column question. Columnar stores refuse — and that refusal is what makes Parquet, ClickHouse, DuckDB, and every modern data warehouse fast.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why CRDTs exist</title><link>https://til.phipham141.dev/posts/data/crdts/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/crdts/</guid><description>Two people typing into the same document at the same time, possibly offline, possibly across the world. The merge has to come out the same on both screens with no central referee. CRDTs are the data structures that make that arithmetic instead of a fight.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why &apos;eventually consistent&apos; became acceptable</title><link>https://til.phipham141.dev/posts/data/eventual-consistency/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/eventual-consistency/</guid><description>For decades, anything weaker than strict consistency was a bug. Then the internet got big enough that strict consistency stopped being affordable — and a generation of engineers learned to live with the gap.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why LSM trees exist</title><link>https://til.phipham141.dev/posts/data/lsm-trees/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/lsm-trees/</guid><description>B-trees write where the key lives. LSM trees refuse to do that — and that refusal is the whole point.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why Merkle trees are everywhere</title><link>https://til.phipham141.dev/posts/data/merkle-trees/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/merkle-trees/</guid><description>A hash tree that lets you prove one tiny piece of a giant dataset is correct without re-downloading the whole thing — the trick that quietly underpins Git, Bitcoin, IPFS, and certificate transparency.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why Postgres uses MVCC</title><link>https://til.phipham141.dev/posts/data/postgres-mvcc/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/postgres-mvcc/</guid><description>Readers shouldn&apos;t have to wait for writers, and writers shouldn&apos;t have to wait for readers — so Postgres keeps multiple versions of every row.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why UUIDv7 is quietly replacing autoincrement IDs</title><link>https://til.phipham141.dev/posts/data/uuidv7-vs-autoincrement/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/uuidv7-vs-autoincrement/</guid><description>Autoincrement IDs can&apos;t be minted client-side and leak how many rows you have. Random UUIDs trash your index. UUIDv7 is the boring fix almost nobody noticed shipping.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Write-ahead logging</title><link>https://til.phipham141.dev/posts/data/write-ahead-log/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/data/write-ahead-log/</guid><description>Why databases write your change to a log before they write it to the actual table — and why crash recovery is impossible without it.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Data</category></item><item><title>Why is the central limit theorem load-bearing?</title><link>https://til.phipham141.dev/posts/math/central-limit-theorem/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/math/central-limit-theorem/</guid><description>Almost every confidence interval, A/B test, and gradient-noise argument quietly leans on one fact: averages of independent things look Gaussian, even when the things themselves don&apos;t.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Math</category></item><item><title>Why does cosine similarity dominate over Euclidean distance in embeddings?</title><link>https://til.phipham141.dev/posts/math/cosine-similarity/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/math/cosine-similarity/</guid><description>Two vectors can be far apart and still mean the same thing. Cosine similarity asks the only question that turns out to matter: are they pointing the same way?</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Math</category></item><item><title>Why matrix multiplication is the bottleneck of modern ML</title><link>https://til.phipham141.dev/posts/math/matmul-bottleneck/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/math/matmul-bottleneck/</guid><description>Modern ML is mostly one operation in a trench coat. Understanding why matmul dominates explains hardware, software, and why GPUs eat the world.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Math</category></item><item><title>Why does information entropy use log base 2?</title><link>https://til.phipham141.dev/posts/math/entropy-log-base-2/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/math/entropy-log-base-2/</guid><description>Shannon could have picked any base for the logarithm in his entropy formula. He picked 2 — and the choice quietly fixes the unit you measure information in.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Math</category></item><item><title>Why does softmax look like that?</title><link>https://til.phipham141.dev/posts/math/softmax/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/math/softmax/</guid><description>Softmax is the function that turns a vector of arbitrary numbers into probabilities. The exponential in the middle isn&apos;t decorative — it&apos;s what makes the whole machine differentiable, well-behaved, and historically inevitable.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Math</category></item><item><title>Why JSON beat XML</title><link>https://til.phipham141.dev/posts/computer-science/json-beat-xml/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/computer-science/json-beat-xml/</guid><description>XML had a standards body, schemas, namespaces, transformations, and a decade head start. JSON had curly braces and a JavaScript parser. Curly braces won.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Computer Science</category></item><item><title>Why fiber-optic beat copper for long distances</title><link>https://til.phipham141.dev/posts/science/fiber-vs-copper/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/fiber-vs-copper/</guid><description>Copper carries electrons; fiber carries photons. The reasons one wins over kilometers come down to physics — how fast the signal fades, how much room there is around the carrier to put signal in, and the fact that light doesn&apos;t care about your neighbor&apos;s microwave.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why AI accelerators are wrapped in stacks of HBM</title><link>https://til.phipham141.dev/posts/science/hbm-stacked-memory/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/hbm-stacked-memory/</guid><description>Open any photo of a modern AI GPU and you&apos;ll see the giant compute die in the middle, ringed by short, fat towers of memory soldered millimeters away. Those towers are HBM, and they exist because regular DRAM physically cannot feed a matrix engine fast enough.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why GPUs ended up running AI even though they were built for graphics</title><link>https://til.phipham141.dev/posts/science/gpus-for-ai/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/gpus-for-ai/</guid><description>GPUs were designed to shade pixels. Then the same hardware turned out to be exactly what neural networks needed. That isn&apos;t luck — graphics and deep learning make the same demand of silicon: identical arithmetic, millions of times over, with nothing to branch on.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why leap seconds exist (and why they&apos;re being abolished)</title><link>https://til.phipham141.dev/posts/science/leap-seconds/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/leap-seconds/</guid><description>The Earth doesn&apos;t spin on a schedule, but our clocks do. Leap seconds tried to bridge the two — and broke the internet doing it.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why CPU clock speeds stopped climbing</title><link>https://til.phipham141.dev/posts/science/power-wall/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/science/power-wall/</guid><description>Around the mid-2000s, GHz numbers on CPUs flatlined while core counts started growing. The reason isn&apos;t engineering laziness — it&apos;s that switching a transistor costs energy, and energy turns into heat you can&apos;t get rid of fast enough.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Science</category></item><item><title>Why does CORS exist?</title><link>https://til.phipham141.dev/posts/networking/cors/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/cors/</guid><description>CORS isn&apos;t there to keep you out of an API — it&apos;s there to stop a webpage you&apos;re visiting from quietly using your logged-in cookies on a different site. The whole design only makes sense once you see that.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why retry with exponential backoff — and why jitter?</title><link>https://til.phipham141.dev/posts/networking/exponential-backoff-and-jitter/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/exponential-backoff-and-jitter/</guid><description>Retrying on failure sounds simple until you ship it at scale. Hammer the server and you make outages worse; back off but synchronize, and you accidentally rebuild the herd. Backoff is the timing rule; jitter is the part that keeps it from biting itself.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why does HTTPS need certificates if encryption already works?</title><link>https://til.phipham141.dev/posts/networking/https-certificates/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/https-certificates/</guid><description>Encryption alone gets you a private channel — to whoever&apos;s on the other end. Certificates are how the browser decides that &apos;whoever&apos; is the bank you meant to reach, not someone sitting in the middle pretending to be.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why is DNS hierarchical?</title><link>https://til.phipham141.dev/posts/networking/dns-hierarchy/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/dns-hierarchy/</guid><description>DNS could have been a giant flat lookup table — one machine somewhere mapping every name in the world to an IP. It isn&apos;t, and the reason is less about technology than about who gets to be in charge of what.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why does QUIC exist when TCP already works?</title><link>https://til.phipham141.dev/posts/networking/quic/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/quic/</guid><description>TCP works fine — until you&apos;re on a flaky phone connection, juggling a dozen multiplexed streams, and one lost packet stalls all of them. QUIC is the protocol designed around that specific frustration.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why does TCP have congestion control?</title><link>https://til.phipham141.dev/posts/networking/tcp-congestion-control/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/tcp-congestion-control/</guid><description>The internet didn&apos;t always have it. Once, in 1986, it nearly fell over. The fix wasn&apos;t a protocol change — it was endpoints learning to back off.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why do CDNs exist when we already have fast servers?</title><link>https://til.phipham141.dev/posts/networking/why-cdns-exist/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/why-cdns-exist/</guid><description>Your origin server can be the fastest box on Earth and your users in São Paulo will still hate it. CDNs exist because the speed of light, not your CPU, is the bottleneck.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why GPU clusters need NVLink and InfiniBand</title><link>https://til.phipham141.dev/posts/networking/why-gpu-clusters-need-nvlink/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/why-gpu-clusters-need-nvlink/</guid><description>Training a frontier model means thousands of GPUs taking the same step at the same time. Ethernet wasn&apos;t built for that, and PCIe gave up a long time ago.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>ASLR: why we shuffle memory before every run</title><link>https://til.phipham141.dev/posts/security/aslr/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/aslr/</guid><description>Attackers used to know exactly where your code lived in memory. ASLR reshuffles it every run, so an exploit has to learn the layout before it can use it — which is why modern exploit chains start by leaking a pointer.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why JWTs are controversial</title><link>https://til.phipham141.dev/posts/security/jwts-controversial/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/jwts-controversial/</guid><description>JWTs solve a real problem — stateless auth across services — and then keep solving it past the point where the cure is worse than the disease. Here&apos;s where the seams are.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Passkeys: why the password is finally being replaced</title><link>https://til.phipham141.dev/posts/security/passkeys/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/passkeys/</guid><description>Passwords are a shared secret you keep retyping into whatever site asked. Passkeys replace it with a key pair whose private half never reaches the site at all.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why password hashing is deliberately slow</title><link>https://til.phipham141.dev/posts/security/password-hashing/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/password-hashing/</guid><description>SHA-256 is fast and that&apos;s exactly why you must not use it for passwords. Password storage is the rare corner of computing where being slow — and greedy with memory — is the feature.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why prompt injection isn&apos;t a bug to be patched</title><link>https://til.phipham141.dev/posts/security/prompt-injection/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/prompt-injection/</guid><description>SQL, XSS, and command injection are all fought the same way: separate the code channel from the data channel. An LLM has labels for that boundary and nothing that enforces them, so the move that works everywhere else has nowhere to land. The vulnerability is the architecture.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why constant-time comparison is a thing</title><link>https://til.phipham141.dev/posts/security/constant-time-comparison/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/constant-time-comparison/</guid><description>An ordinary equality check leaks the secret it&apos;s supposed to protect — one byte at a time, through the clock. Constant-time comparison exists because == is faster than it should be.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why public-key signatures are not just &apos;encryption in reverse&apos;</title><link>https://til.phipham141.dev/posts/security/signatures-vs-encryption/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/security/signatures-vs-encryption/</guid><description>They look symmetric — encrypt with one key, decrypt with the other — but signatures and encryption answer different questions, and conflating them is how real cryptosystems get broken.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Security</category></item><item><title>Why do LLM responses stream?</title><link>https://til.phipham141.dev/posts/networking/why-llm-responses-stream/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/networking/why-llm-responses-stream/</guid><description>It&apos;s not for show. The model literally generates one token at a time, and forcing it to buffer the full answer before sending would make every chat app feel broken. Streaming is the network shape of an autoregressive process.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Networking</category></item><item><title>Why circuit breakers exist</title><link>https://til.phipham141.dev/posts/systems/circuit-breakers/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/circuit-breakers/</guid><description>Backoff makes a single retry polite. But when a downstream is plainly down, every caller in your fleet generously retrying it is the actual problem. A circuit breaker is the small piece that says: stop calling for a while — the answer isn&apos;t going to change in the next 50 ms.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why monotonic time is different from wall-clock time</title><link>https://til.phipham141.dev/posts/systems/monotonic-vs-wall-clock-time/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/monotonic-vs-wall-clock-time/</guid><description>Wall-clock time tells you what to put on a calendar. Monotonic time tells you how long something took. Confusing them is how you get bugs that look like physics violations.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why memory-mapped files exist</title><link>https://til.phipham141.dev/posts/systems/memory-mapped-files/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/memory-mapped-files/</guid><description>Why hand a file to the OS as memory instead of reading it byte by byte? Because the OS was already caching it that way — and pretending otherwise costs you a copy you don&apos;t need. What you pay for deleting the copy is knowing when the I/O happens.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why syscalls are expensive</title><link>https://til.phipham141.dev/posts/systems/syscalls-are-expensive/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/syscalls-are-expensive/</guid><description>A function call costs a few cycles. A system call costs hundreds — sometimes thousands. The gap isn&apos;t sloppy engineering; it&apos;s the price of the user/kernel boundary.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why Linux has an OOM killer</title><link>https://til.phipham141.dev/posts/systems/oom-killer/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/oom-killer/</guid><description>Linux promises memory it doesn&apos;t have, then has to break the promise — the OOM killer is the reaper that decides who dies so the system can keep running.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why virtual memory exists</title><link>https://til.phipham141.dev/posts/systems/virtual-memory/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/virtual-memory/</guid><description>Every process thinks it owns the whole machine. That lie is the foundation almost every modern OS feature is built on.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why containers won over VMs</title><link>https://til.phipham141.dev/posts/systems/why-containers-beat-vms/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/why-containers-beat-vms/</guid><description>Both promise isolated, reproducible environments. One boots in milliseconds and ships in megabytes; the other boots in seconds and ships in gigabytes. The reason isn&apos;t &apos;containers are lighter VMs&apos; — they&apos;re a different kind of thing entirely.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why fork() is such a weird API</title><link>https://til.phipham141.dev/posts/systems/why-fork-is-weird/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/why-fork-is-weird/</guid><description>Other systems take a program and arguments. Unix takes your whole process and clones it. The reasons are half historical accident, half deep insight — and the seams still show.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item><item><title>Why idempotency keys exist</title><link>https://til.phipham141.dev/posts/systems/idempotency-keys/</link><guid isPermaLink="true">https://til.phipham141.dev/posts/systems/idempotency-keys/</guid><description>The network can drop your response after the work is done. Now you have to retry — and you have no idea whether you&apos;d be doing it for the first time or the second. Idempotency keys are the small protocol the client and server agree on so the retry is safe.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Systems</category></item></channel></rss>