Your advice is good in general. However, note the OP is reporting user time (probably as shown by the "time" command). This time is the total CPU time used by the process. It's not measuring how much wall clock time it took to complete or the time used to complete involved IO operations. I think trusting that number should be fine in this case.
It can cause problems with competing for cache space, which could have large effects on CPU time.
You're right and I stand corrected.
He easily could have discovered that during debugging or when he was actually trying to make progress on the project, not benchmarking. But I agree, he should be explicit about that variable.
you're just increasing the uncertanity, but not drastically so. If you use your laptop during benchmarks that show 10x improvement, then your thesis still stands
Not if you were playing Crysis 3 during part of one benchmark and looking at Facebook the rest of the time.
No. To be honest PyPy is very sensitive to cache usage, so running any other program trashing the cache might be a serious problem (can lead to 30% performance degradation, depending on the load even if the core is unoccuppied at all)
I'm confused. You started with "No" and then continued with something that is either in agreement or orthogonal to my point.
That was "I agree with you" kind of no. English is hard, sorry.
Ah, I follow now. I knew I was missing something, and there it is. As somebody who writes English professionally, I agree wholeheartedly.
For the end benchmark I've used "runlevel 3" (multiuser without window manager) to perform this tasks, to maximize cache and RAM usage.
That's the spirit! You should also stop services known for spikes in CPU usage like the cron daemon and run the benchmark multiple times.
A good way to measure CPU performance in a CPU-agnositic way (not sure if thats the right phrase to use) is instruction-count. you would have to disregard things like cache when looking at instruction count, which is probably bad, but it serves as a good CPU measure.
No, it's not good at all. Cache stalls can easily account for 1/2 of your processing time. On top of that you have CPU pipelining and multi-issue CPUs. It was a good idea a while ago, now it's really not that great.
You do if that's similar to the expected deployment environment.
Only if it's something you can replicate exactly for each run of the benchmark.
There's no need for perfect repeatability when statistical analysis is good enough. After all, even if you have control down to the iron, the randomness in external interrupts and the effect of temperature on the hardware will cause some unpredictability.