For what it's worth, Python's standard library `timeit` command-line tool does this by default. (By importing the module you can programmatically assume more direct control over the runs.) And that's admittedly primitive (PyPy warns against using it).
A few years back I did some refactoring and minor enhancements on that code and wrote a blog post about it (https://zahlman.github.io/posts/timing/). The whole thing could probably use more work, and I'm sure I didn't contribute anything novel to the science of benchmarking; but it was fun.