I almost shipped a benchmark that said parsing a NASDAQ message takes 0.98 nanoseconds. That’s about three CPU cycles. For decoding sixteen bytes across three fields.
I was proud of it for about ten minutes, which is exactly how long it took a friend to ask “why is it faster than a memory access?” Then I got suspicious, and then I got annoyed, because the compiler had been lying to me with a perfectly straight face and I’d been the one holding the pen.
Google Benchmark works like Go’s testing.B. You write a loop, it decides how many iterations to run until it has enough samples. Mine ran ~163 million iterations and reported 0.98 ns.
The number is wrong for a boring reason: the compiler deleted my parser.
I had written DoNotOptimize(parse_A(d, out)). That pins the return value — a bool. It says nothing about the struct the function filled in. So the compiler looked at the writes to out, saw nobody ever read them, and quietly removed the whole decode. I was benchmarking an empty loop.
The fix is to pin the thing that matters — the output:
parsers::nasdaq::parse_A(d, out);
benchmark::DoNotOptimize(out);
Now the fields have to actually be computed. The number went up to a bizarrely large… 1.77 ns. Still fast. Still real.
Even with the output pinned, d was the same bytes every iteration, so the compiler computed the result once and reused it. Loop-invariant hoisting — the benchmark measured a register read.
So I made the input move:
d[19] ^= 1; // flip the side byte each iteration
Now it can’t cache the answer. Honest number: 1.77 ns. I trust it now, and honestly it’s more impressive than the fake one, because it’s true.
Here’s the shape of the output once you turn on repetitions:
BM_BookReduce_median 3.42 ns
BM_BookReduce_stddev 0.037 ns
BM_BookReduce_cv 1.07 %
I had a beautiful table of micro-benchmarks:
| op | ns |
|---|---|
| parse A/F/X | 1.7–2.0 |
| best() | 0.32 |
| reduce | 3.3 |
| add / remove | 14 / 18 |
| depth(5) | 42 |
And none of it was the point. When I profiled the whole replay, it ran at 11 million messages per second — about 90 ns per message — and ~75% of that was gzip decompression. The order book I’d been micro-optimising was a rounding error in the real workload.
That’s the lesson I keep failing to internalise: micro-benchmarks make the thing you measured look important. Only the macro run and a profiler tell you whether it actually is. The micro numbers are accurate. Accurate and irrelevant is the most expensive kind of result, because it’s so easy to act on.