This guy has had a few talks at CppCon over the years, and covers the basics of how to attempt benchmarking with
actual computer science instead of myth and voodoo rituals:
https://www.youtube.com/watch?v=koTf7u0v41oIn practice, it's not critical to go that far. In my experience, huge performance wins are trivial to benchmark with just a modicum of effort, and tiny performance wins are typically not worth any effort at all. At FAANG scale they might be, but I don't work at a FAANG, so not my problem! :)
Simple methods I use and can recommend are:
- Ideally, collect data from production, where the change behaves exactly as it would in production because it is production!
- Interleave tests on the same machine. E.g.: A, B, A, B, A, B, etc...
- Run tests on every machine you can get your hands on over an extended period and draw a histogram or a scatter chart of some sort. You don't have to do this at 100% load, it's actually better to mix this in with normal workloads. This gives you an idea of the noise versus signal, platform variance, variation due to time of day, etc...
- Run tests after a cold boot to get a sense of worst-case performance instead of idealised and often unrealistic "hot loop" performance.
- Run tests with different input sizes and draw a graph. Try find the explanation for every feature of the graph.
- Be aware of every cache layer in the system. Don't forget about the potentially huge RAM or SSD cache of the storage subsystem. Typical SAN arrays now have 256 GB or more. The previous graph above can help, but only if your input size goes waaaay past the biggest cache sizes in the system.
As a real-world example: I recently reduced the size of a database platform by a few vCPUs to save on licensing costs, but increased the per-core performance and simultaneously did some network tuning. Some benchmarks showed the same performance, others showed an improvement. Both were extremely noisy before and after the change.
I would like to take credit and say that I improved performance while saving money. However, the scientifically honest result is that any performance change is statistically insignificant.
The cost saving is significant, because it has very little variance. That's good enough, and I don't have to make a statement about performance that the numbers can't really back up.