Blog
The Debug Line You Disabled Is Still Running
A DEBUG line that never prints still cost me 5,260 ns a call, because the logger checks the level after you've already built the message. That, sync writes on the request thread, and logging every event instead of sampling: three ways logging quietly eats a hot path.
The Priority Queue That Didn't Cut the Line
One queue carried the marketing blast and the OTPs. I added a priority queue and the OTP still waited 3.8 seconds behind the backlog. The culprit was prefetch, not priority. Here's the reproduction.
The Machine That Rehashed Everything
You shard a table across four Postgres machines with hash % 4 and everything gets faster. Then you add the fifth machine and find out that changing the divisor moves four-fifths of your rows, because the row count was baked into the address. The fix is to never hash to machines at all, and the whole design turns on the number you pick before any of that happens.
The Export That Got Slower Every Page
A CSV export paged through a million rows with LIMIT/OFFSET and slowed down the deeper it went. Deep OFFSET re-walks every row it skips, so the last page reads almost the whole table. Measured on Postgres 16.
The Commits That Didn't Survive the Failover
A MySQL primary told a thousand clients their writes were committed, then crashed. The replica came up in seconds, the dashboard went green, and every one of those thousand rows was gone. Async replication will do that quietly; I went to measure exactly how much semi-sync buys you back, and what it costs.
The Read That Couldn't See Its Own Write
You point your reads at a replica pool to take load off the primary, and then a user saves a setting, the page reloads, and the old value comes back. The write was fine. The read went to a replica that hadn't heard yet. Here's the LSN gate that gives you read-your-writes without giving up the replica.
The Events That Wouldn't Compress
I batched analytics events and watched zstd hit a wall at 4.7x. The floor turned out to be the events themselves, the identifiers that make each one unique are also the bytes that won't compress.
Catching the Bots Without Remembering the Clicks
A click firehose has bots hammering it from a handful of sources, and the obvious way to flag them is a counter keyed by source. That counter grows forever. Count-Min Sketch, Top-K and a Bloom filter do the same job in a fixed few megabytes. Measured on Redis 7.4.7.
Cranking Up the Confirm Window Made RabbitMQ Slower
Turn on publisher confirms, throughput drops, so you widen the in-flight window to win it back. Past a moderate sweet spot, widening it made throughput go down, not up. And fire-and-forget's higher number turned out to be a backlog, not throughput. Measured with PerfTest.
The Latency Numbers Nobody Reruns
Everyone quotes "L1 1ns, main memory 100ns, SSD 16µs" from a table that's fifteen years old. I reran the whole thing on my laptop with a pointer-chase harness. Some rows moved 20x, one hasn't budged since 2012, and one got slower than the number everybody memorized.
The Objects That Never Died
A merged service ran hotter for the same traffic. p50 and p99 were fine, but the GC threads were pinned. G1 was spending its afternoons re-marking a heap that was half long-lived cache, proving over and over that objects which were never going to die were still alive. Moving them off-heap cut total GC time 97.8% and lifted throughput 30%, and the latency cost I braced for never showed up.
The Slowest Request Was a Garbage Collection
p50 and p99 flat, and then one request in the trace takes 136 milliseconds. Nobody sent a slow request. The JVM stopped every thread to collect garbage, and the fix was one flag that traded half the throughput for a tail that never freezes.
Want to get blog posts over email?
Enter your email address and get notified when there's a new post!