Blog
The Export That Got Slower Every Page
A CSV export paged through a million rows with LIMIT/OFFSET and slowed down the deeper it went. Deep OFFSET re-walks every row it skips, so the last page reads almost the whole table. Measured on Postgres 16.
The Commits That Didn't Survive the Failover
A MySQL primary told a thousand clients their writes were committed, then crashed. The replica came up in seconds, the dashboard went green, and every one of those thousand rows was gone. Async replication will do that quietly; I went to measure exactly how much semi-sync buys you back, and what it costs.
The Read That Couldn't See Its Own Write
You point your reads at a replica pool to take load off the primary, and then a user saves a setting, the page reloads, and the old value comes back. The write was fine. The read went to a replica that hadn't heard yet. Here's the LSN gate that gives you read-your-writes without giving up the replica.
The Events That Wouldn't Compress
I batched analytics events and watched zstd hit a wall at 4.7x. The floor turned out to be the events themselves, the identifiers that make each one unique are also the bytes that won't compress.
Catching the Bots Without Remembering the Clicks
A click firehose has bots hammering it from a handful of sources, and the obvious way to flag them is a counter keyed by source. That counter grows forever. Count-Min Sketch, Top-K and a Bloom filter do the same job in a fixed few megabytes. Measured on Redis 7.4.7.
Cranking Up the Confirm Window Made RabbitMQ Slower
Turn on publisher confirms, throughput drops, so you widen the in-flight window to win it back. Past a moderate sweet spot, widening it made throughput go down, not up. And fire-and-forget's higher number turned out to be a backlog, not throughput. Measured with PerfTest.
The Latency Numbers Nobody Reruns
Everyone quotes "L1 1ns, main memory 100ns, SSD 16µs" from a table that's fifteen years old. I reran the whole thing on my laptop with a pointer-chase harness. Some rows moved 20x, one hasn't budged since 2012, and one got slower than the number everybody memorized.
The Objects That Never Died
A merged service ran hotter for the same traffic. p50 and p99 were fine, but the GC threads were pinned. G1 was spending its afternoons re-marking a heap that was half long-lived cache, proving over and over that objects which were never going to die were still alive. Moving them off-heap cut total GC time 97.8% and lifted throughput 30%, and the latency cost I braced for never showed up.
The Slowest Request Was a Garbage Collection
p50 and p99 flat, and then one request in the trace takes 136 milliseconds. Nobody sent a slow request. The JVM stopped every thread to collect garbage, and the fix was one flag that traded half the throughput for a tail that never freezes.
One Byte Over the Slab
memcached rounds every item up to the next slab class. Land one byte over a boundary and you hand a fifth of your RAM to the allocator for nothing. I measured it.
The Messages RabbitMQ Confirmed and Lost Anyway
A publisher confirm is supposed to mean the broker has your message. On a classic mirrored queue I confirmed 5,000 messages and a single node failure lost all 5,000. Here's the reproduction on RabbitMQ 3.13, and what quorum queues do differently on 4.0.
The Query That Asked Every Shard
One Postgres table outgrows one machine, so you split it across eight. Then you find out that a query carrying the shard key touches one shard in 2ms, and the same query without it has to ask all eight and comes back in 12. The shard key wasn't a detail. It was the whole design.
Want to get blog posts over email?
Enter your email address and get notified when there's a new post!