That Broken Performance Comparison

Last year I published an internal write-up — it never left the company, which is why I can be honest here — claiming a 4.3× speedup for a batch job after a rewrite. The number was real. The conclusion was wrong. Three months later, in production, the new version was maybe 15% faster, on a good day, with a tailwind.

This post is the post-mortem of my own benchmark. Not because the mistake was exotic — every part of it is a classic — but because I made it with full knowledge of the classics, and the checklist I use now is the actual useful artifact.

What went wrong, in order of embarrassment

1. The old version was cold, the new one was warm

The comparison ran the incumbent first and the rewrite second, in the same process, on the same data. By the time the rewrite ran, the page cache held the entire input, buffers were sized right, and the JIT-equivalent (Go's runtime; insert your ecosystem's equivalent) had profiled everything. I compared a warm engine against a cold start and called the difference an improvement.

2. One machine, one run, one number

I ran it once. On the machine I had. Which was newer than the fleet median and had, we later found, a NIC that the office Wi-Fi loved more than the datacenter switches did. A single run has no variance, so it produces no distrust, which is exactly what makes it dangerous.

3. The dataset was the demo dataset

It was small, uniform, and fit in memory end to end. The production input was 60× larger and had a long tail of pathological records that the rewrite handled through a slower but safer path. The benchmark measured the rewrite's happy path; production lived in the tail.

4. I measured the wrong endpoint

The benchmark timed “processing complete.” The user-visible requirement was “results available.” Those differ by the flush and index-build phase, which the rewrite deferred and which, in production, happened during the hour people actually read the results. The old version front-loaded that cost invisibly. The new version won the race that did not matter.

The checklist I use now

  1. Interleave and randomize run order. A/B/A/B, not A/A then B/B. If warmup exists, warm up both or neither.
  2. Three machines minimum, fleet-typical hardware. Report per-machine, not averaged; the spread is the finding.
  3. Report the distribution. Median and p95 of N runs, plus the input variance. One number is a vibe.
  4. Two datasets: production-shaped and adversarial. If the tail changes the ranking, the tail is the result.
  5. End-to-end user timing. Measure from the boundary the user experiences, including flushes, fsyncs, and index builds.
  6. Re-run the whole thing after a week. On the real workload, in staging, with production's cron pattern. The benchmark is a hypothesis; this is the experiment.
  7. Write the falsification line first. Before running: “if the gain is under X%, we do not ship the rewrite.” Deciding the line after seeing the numbers is how 4.3× gets published.

The uncomfortable part

The rewrite did ship, and it was fine — cleaner, testable, and 15% faster, which over a year was worth it. But it would have shipped on those grounds anyway. The 4.3× did not change the decision; it changed the story, which is worse, because the story gets repeated and the caveat does not. The people who heard “4.3×” quote it still. The people who heard “15%, worth it” have, so far, all remembered it correctly.

That is the real lesson: precision in benchmarks is not about the number, it is about what the number will be allowed to justify later. A benchmark you cannot defend is a rumor with a graph.