992 karma · joined February 1, 2015
No, we never thought about HTML output. However, there are multiple other export options and we also ship Python scripts that can be used to plot the benchmark results. The script is not very large, so far, but we are happy to add new scripts if the need for one should arise. What kind of diagrams would you like to see?
Also, I have never thought about using criterion.rs. My feeling was that it is suited for benchmarks with thousands of iterations, while we typically only have tens of iterations in hyperfine (as we typically benchmark programs with execution times > 10 ms). Do you have anything specific criterion feature in mind that we could benefit from?
> Most -- nearly all -- benchmarking tools like this work from a normality assumption
I don't think that hyperfine makes any assumption about normality. Sure, we do report sample mean and sample standard deviation by default, but we also report sample minimum and the maximum. You can also easily export all the benchmark results and inspect in more detail with the supplied Python scripts.
> In fact, performance numbers (latencies) often follow a heavy-tailed distribution
So when is this really the case? In my understanding, if I am measuring the runtime of a deterministic program with the same input, the runtime should only be influenced by external factors that are out of my control (other programs being scheduled, caching effects, hardware-specific influences, ..). These are exactly the things that I want to "average out" by running the benchmark multiple times.
> What's worse is when these tools start to remove "outliers".
Hyperfine never removes outliers. What we do is to try and detect outliers. We do this by computing robust statistical estimates that specifically DO NOT assume a normal distribution (see https://github.com/sharkdp/hyperfine/blob/master/src/hyperfi... for details).
We perform this outlier detection to warn users about potentially interfering processes or caching effects.
Take a look at these results, for example: https://i.imgur.com/XRvE6Ys.png
I benchmarked a file-searching program. The underlying distribution, while probably not normal, seems to be "well behaved" and I think that the sample mean and the sample standard deviation could be quantities with a reasonably predictive power.
What you do NOT see in the histogram is a single outlier at 1.15 seconds, far outside the plot to the right. This was the first benchmark run where the disk caches were still cold. In such a case, hyperfine warns the user:
Warning: The first benchmarking run for this command was significantly slower than the rest (1.152 s). This could be caused by (filesystem) caches that were not filled until after the first run. You should consider using the '--warmup' option to fill those caches before the actual benchmark. Alternatively, use the '--prepare' option to clear the caches before each timing run.
In conclusion, I am not quite sure how your critisism applies to hyperfine, but I'd be happy to get further feedback.
Old discussion: https://news.ycombinator.com/item?id=16193225
Looking forward to your feedback!
> The dbg! macro works exactly the same in release builds. This is useful when debugging issues that only occur in release builds or when debugging in release mode is significantly faster.
Note, however, that the C++ dbg(..) macro can be easily disabled to a no-op (identity-op, to be precise) with the DBG_MACRO_DISABLE flag.
Exactly. Drop-in compatibility with 'cat' is one of the goals of bat (see https://github.com/sharkdp/bat#project-goals-and-alternative... and a list of alternatives here: https://github.com/sharkdp/bat/blob/master/doc/alternatives....).
Another thing that I use frequently is previewing a whole set of files in a single (pager) output. Something like
bat src/*.cpp
This also allows you to easily search across a whole set of open files.Web version: https://insect.sh/
Your calculation: https://insect.sh/?q=10m%20*%201000kg%20*%209.8m%2Fs%5E2%20t...
If you feel that there is anything we could to to improve fd's UX, it would be great if you could share it on GitHub: https://github.com/sharkdp/fd
> Why separating the extension from the filename? Is this a default
It does not need to be! You can just also just use "fd README.md". Admittedly, the "." will actually search for any character (regex), so if you want to be precise, you need to escape it.
I removed my fork...
Benchmark #1: hexyl $(which hexyl)
Time (mean ± σ): 169.8 ms ± 8.2 ms [User: 152.5 ms, System: 17.1 ms]
Range (min … max): 162.2 ms … 189.1 ms 16 runs
Benchmark #2: hexdump -C $(which hexyl)
Time (mean ± σ): 188.5 ms ± 4.4 ms [User: 186.2 ms, System: 2.2 ms]
Range (min … max): 184.1 ms … 198.2 ms 14 runs
Benchmark #3: xxd $(which hexyl)
Time (mean ± σ): 72.8 ms ± 2.7 ms [User: 71.9 ms, System: 1.1 ms]
Range (min … max): 71.0 ms … 87.8 ms 40 runsYes, it's a shame. But I don't think there is too much we can do about it. We have to print much more to the console due to the ANSI escape codes and we also have to do some conditional checks ON EACH BYTE in order to colorize them correctly. Surely there are some ways to speed everything up a little bit, but in the end I don't think its a real issue. Nobody is going to look at 1MB dumps in a console hex viewer (that's 60,000 lines of output!) without restricting it to some region. And if somebody really wants to, he can probably spare 1.5 seconds to wait for the output :-)
.. which is apparently also called "Hexyl". Granted, I was just looking for a short word that starts with "Hex" and I was never good at organic chemistry :-)
fd 'any.*thing'Hyperfine currently tracks real time (= wall-clock time), user time (= time spent in user mode) and system time (= time spent in kernel mode).
Unfortunately, I have never heard of dtrace. What kind of other metrics would you be interested in?
I believe I would like hyperfine to focus on timing-aspects.
I personally use command-line benchmarking to compare different tools. You might want to compare grep, ack, ag and ripgrep. I currently use it to profile my find-alternative fd and to compare it with find itself (https://github.com/sharkdp/fd-benchmarks).
You could also use it to find an optimal parameter setting for a command-line tool (make -j2 vs. make -j8).
Yes, having a "cold" or a "warm" disk cache makes a massive difference for I/O-heavy programs. For one of my other programs, I differentiate between "cold-cache" and "warm-cache" benchmarks: https://github.com/sharkdp/fd-benchmarks
Not yet, but I've just created a ticket here: https://github.com/sharkdp/hyperfine/issues/20
Should be easy to implement.