How the GNU coreutils are tested (2017)
pixelbeat.org
pixelbeat.org
That's pretty respectable, given that coreutils include 98 programs (some are simple like yes(1) and true(1), but most of them are used millions of times a day to do real work: ls(1), kill(1), cat(1), wc(1).
In fact, I used wc(1) to count the number of separate programs inside coreutils.
Not that simple: https://github.com/coreutils/coreutils/blob/master/src/yes.c
It also, frankly, feels like the wrong layer for such an optimization. I would have hoped there was a c "write to stdout" method that does all the buffering and performance tricks this thing does.
> If you have a vague recollection of the internals of a Unix program, this does not absolutely mean you can’t write an imitation of it, but do try to organize the imitation internally along different lines, because this is likely to make the details of the Unix version irrelevant and dissimilar to your results.
> For example, Unix utilities were generally optimized to minimize memory use; if you go for speed instead, your program will be very different.
So I think a lot of coreutils etc were written for extreme speed so there would obviously be no crossover with existing UNIX source.
I love using it as an example of how implementing something so simple can lead to learning about so many seemingly unrelated things like context switches and memory alignment.
That inefficiency is a far bigger sin than the "slowness" of yes needing a write() call for every line it emits, given the intended purpose of the command (which is not, as the code suggests, to saturate a unix pipe as fast as you can.)
It prints them as long as the receiving process reads them. That’s how pipes work. When piping to a process that only reads from stdin once (you know, like the original purpose of the command, to respond “yes” to prompts from processes that are asking for confirmation), it will only write once (because stdout is line-buffered, so printing a string with a newline will block until something reads it.) Filling a buffer with 4,000 instances of “y\n” on the off chance that the receiving process will actually read all of those, is doing extra work that may not be needed.
“Maddening research for performance” in this case is coming at the expense of energy expenditure: you’re always allocating that memory and always filling a buffer with thousands of “y”s even when it’s just going to get thrown out. That should not be applauded. People should be just as concerned about energy usage as they are about wall clock time.
self-hosting achievement unlocked
How the GNU coreutils are tested - https://news.ycombinator.com/item?id=29459185 - Dec 2021 (30 comments)
How the GNU coreutils are tested (2017) - https://news.ycombinator.com/item?id=18034265 - Sept 2018 (9 comments)