A little story about the `yes` Unix command
endler.dev
endler.dev
The /home directory was mounted to an autoexpanding EFS on AWS.
23.4tbs and 2 months later we noticed the bill :)
I assume a smaller place would just beg AWS for forgiveness and probably get it.
[0] https://docs.aws.amazon.com/lambda/latest/operatorguide/recu...
But the typical worst case-practices, are the small companies who think avoiding vendor lock-in is a thing that matters at their scale. Look, you're never going to change cloud providers - if you do, that will be a problem you can solve then. You're never going to go multi-cloud. If you do that's a problem you can solve then. Preemptively DIY'ing everything from database copies to security to encryption is going to break you - you're now engineered on a brittle substrate of hacks with no support, all so that one day you could maybe consider saving 10% on your cloud bill by moving to a different provider. The day that migration will save you more than one engineer per month, consider it. Until then, you're just making your own costs worse.
/rant
Vendor lock-in is only tolerable when the service is so easily swapped out that it's not actually vendor lock-in.
There is no valid argument for not worrying about that before it happens, and bending pretty far to avoid it. No matter how hard you work to stay as portable as possible early and at each daily step along the way, it's 10x or 1000x less than dealing with it later.
If you're just talking about a pluggable service, well then by definition that's not really lock-in.
Portability is a huge myth that eats engineers hours like a snack every single day and rarely pays off.
This. If AWS ever were to even consider dramatically increasing their prices, there are players whose AWS bills are two or three or more orders of magnitude more expensive than yours who will howl and gnash their teeth and whose potential departure from AWS does far more to protect you than you could ever do to protect yourself.
Similar stupidity includes spending engineering time on such issues like how to deal with an S3 outage. If S3 is down, your competitors are down too. Nobody cares.
It's nice that it's manageable and also a learning experience, but over here that would be like 2 months' take home salary for a software engineer.
Kind of why development/test environments shouldn't have autoexpanding or scalable anything, in my experience.
First, to retain an air of "vaguely knows what they're doing" about me, even though everyone makes mistakes and that should be treated as something that's okay - especially if you can limit the impact of mistakes, like with automated spending limits.
Secondly, because I wouldn't want to risk doing something like that in a personal project, given that my wallet is likely to be much thinner than those of organizations.
I'd expect that you'd reach the maximum once your card is rejected. :)
But truthfully many platforms out there will let you set up spending alerts, but not outright set limits because then you get into a bunch of difficult questions - should further data just be redirected and piped to /dev/null? Should you as the service provider instead limit IOPS in some way, or allow slower network connectivity if allowed egress amount of data is exceeded? What about managed databases, slow it down or throw it out altogether?
I've talked about those in detail with some of the people here ages ago and there are actually companies that take the "graceful degradation" approach, like Time4VPS who host almost all of my cloud VPSes for now: https://www.time4vps.com/?affid=5294 (affiliate link, feel free to remove affid if you'd prefer not to have it)
What happens if I run out of bandwidth?
We reduce your VPS server’s port speed 10 times until the new month starts. No worries, we won’t charge any extra fees or suspend your services.
Honestly, that's a really cool idea for handling resources, though one has to also understand that storage is a bit different and if you've built your entire platform around the concept of scalability and dynamically allotting more resources, you might make the choice of having the occasional story of large bills (some of which you'll probably forgive for good PR), as opposed to more frequent stories about things going down because the people forgot to pay you, as well as many enraged individuals complaining about their data being deleted, though it's supposedly your fault.So in a way it's also a business choice to be made, though one can feasibly imagine hard spend limits being feasible to implement.
> Not using an expandable storage in test risks deviations from prod (and then you get the prid only bugs that are difficult to keep fixed).
I concede that this is an excellent point, you also should be able to test automatic scaling when necessary etc.
Though the difference is probably in being able to test it but not leave it without (more conservative) limits when you're not looking at it.
If you are going to use ANY cloud provider, learn about alarms, or you will get screwed.
Writing portable, robust code is nontrivial. This is a great example.
For those who didn't RTFM:
1979 version:
main(argc, argv)
char **argv;
{
for (;;)
printf("%s\n", argc>1? argv[1]: "y");
}
Current 128-line GNU version:https://github.com/coreutils/coreutils/blob/master/src/yes.c
What would be way more interesting is fixing the standard library/OS interfaces to make the slow version fast, because that would likely benefit other simple filter commands as well. Might require using non-POSIX interfaces to do well, of course.
edit: POSIX note
What is your argument that it is silly? Have you done a survey of usage, or researched its history. Don't be so quick to criticize code, especially from a huge, old, project (that might even be older than you).
> Don’t in any circumstances refer to Unix source code for or during your work on GNU! (Or to any other proprietary programs.)
> If you have a vague recollection of the internals of a Unix program, this does not absolutely mean you can’t write an imitation of it, but do try to organize the imitation internally along different lines, because this is likely to make the details of the Unix version irrelevant and dissimilar to your results.
> For example, Unix utilities were generally optimized to minimize memory use; if you go for speed instead, your program will be very different. You could keep the entire input file in memory and scan it there instead of using stdio. Use a smarter algorithm discovered more recently than the Unix program. Eliminate use of temporary files. Do it in one pass instead of two (we did this in the assembler).
> Or, on the contrary, emphasize simplicity instead of speed. For some applications, the speed of today’s computers makes simpler algorithms adequate.
> Or go for generality. For example, Unix programs often have static tables or fixed-size strings, which make for arbitrary limits; use dynamic allocation instead. Make sure your program handles NULs and other funny characters in the input files. Add a programming language for extensibility and write part of the program in that language.
> Or turn some parts of the program into independently usable libraries. Or use a simple garbage collector instead of tracking precisely when to free memory, or use a new GNU facility such as obstacks.
[1]: https://cvsweb.openbsd.org/src/usr.bin/yes/yes.c?rev=1.9&con...
[2]: https://git.savannah.gnu.org/cgit/coreutils.git/tree/src/yes...
Similarly, there is the NSFW comparison between Plan 9’s cat and GNU coreutils’ cat [3].
[3]: http://9front.org/img/longcat.png
Just to reiterate in the end though. This is not an argument against optimisation and learning how to make something blazingly fast. But, there is such a thing as optimising the wrong thing and using speed as the only justification for merging a patch is probably not the right thing to do for a bigger project.
The reason it is useful is because yes can output anything, and so is useful to produce any repeated data for test files etc.
You can see the justification detailed in the original optimization commit: https://github.com/coreutils/coreutils/commit/35217221
I guess there's no ideal solution. But I think the "new" program does what the first one did better, and does not do anything worse.
We are talking about a program that still under 1000 lines of code and that's not getting new features every month, or at all anyway, so maintainability does not seem to be a big issue?
I see the reasoning but I don't see any actual practical drawback to have improved the original program directly in this specific case. I don't see any advantages of keeping "yes" dead simple neither. The new version is still pretty much readable and the extra time it takes to read it and modifying without doing mistakes seems worth the advantages.
My point is, I'd never thought of using yes for this purpose. So in this case, you could make a command called 'outputsomethingfast' and you could make a command called yes, that internally calls 'outputsomethingfast --output=yes' or something like that.
To me this is way more logical, and more in line with the Linux philosophy, right?
However, this is independent from this optimization, "yes" already had this feature of outputting anything, I think?
But I expect this kind of accident to happen in any working system that has been long enough. This seems unavoidable. So we'd better put up with this kind of mess probably.
also, if there's a new program, say, "fastrepeat" wouldn't that be a duplication of functionality between "yes" which just outputs "y" and "fastrepeat 'y'" which is, like, even more bloat since now you need both ?
well, I definitely don't. I don't want to encumber my mind with a name for every single of the 25000 "one thing" things I have to do.
Did this advice from GNU inspire you to optimise `yes`, did your optimised `yes` inspire GNU to write this, or is there no historical connection between your optimised `yes` and this advice?
You compare two things and see one is longer and more complex than the other. How much cost is that really? 10x more code sounds bad, but 100 more lines might put it in perspective. And how complex is the code really? And what is the benefit? A common complaint seen on OpenBSD lists is that performance is behind competitors so you take a bunch of those complaints and make an equally sound argument the other way.
I will say that a lot of the tools and libraries and functionality I have seen the hell optimized out of and functionality added to, allows solutions to be put together which would be infeasible or impossible with simple / naive implementations. More layers or custom code or more complexity can be avoided. Let's say a database layer could be avoided if filesystem operations are fast enough. Or a shell script + cmd line tools can be used instead of writing a new program if fork+exec+exit+context switching+pipe IO, and these kind of tools (yes and cat) are fast. If malloc+free are fast then you don't need to write your own caching allocator in front of it. Etc etc. So you might end up with an end-to-end solution that meets your requirements and actually has less code, or at least less bespoke complexity and more that is long maintained and used by many.
[0]: https://www.gnu.org/prep/standards/html_node/Reading-Non_002...
[1]: http://trillian.mit.edu/~jc/humor/ATT_Copyright_true.html
> However, I do think that one should not arrive at the belief that this kind of
> optimisation is warranted everywhere and that code simplicity can also be a goal
Definitively not warranted everywhere, especial if premature and at such detail, but core utils like `cat` and `yes` are IMO prime examples of where such optimizations are warranted:- They're in use daily on a huge amount of setups, even small benefits add up much more than in some niche tool.
- They got a clear and small feature set that won't change anytime soon, so there won't be much code churn and thus maintenance effort will be relatively low.
> Similarly, there is the NSFW comparison between Plan 9’s cat and GNU coreutils’
> cat [3].
>
> [3]: http://9front.org/img/longcat.png
IMO that isn't an exactly fair comparison though, as the difference is not only in optimizations but a lot in boilerplate license/copyright comments and in features like line numbering, modes for showing (non-printable) special characters or whitespace and option parsing for said features.Strip all that out, and you got (eye balled) 1/3 of that, and 200 lines for an extremely fast core util is really not much nor hard to maintain, as it won't get any new features soon anyway.
Is the posted link supposed to be a reference? That is your own tweet aka self-reference, boasting your own contribution.
In the tweet itself you say "estimate"! How do arrive at such a grandiose estimate?
How do you attribute a saved carbon footprint to an optimization in a command line tool? You cannot even approximate that. I would argue that such tiny optimizations make 0, nil difference in overall energy consumption on my local machine, all my laptops and all the servers in this building.
I'm not saying that we shouldn't run optimized code, but everyone can scream around random numbers.
I remember a page somewhere saying plan9 people where very much against making cat be a generic tool for all those use cases, and it should do what it says in the name: concatenate files.
While in general efficient programs can save energy, the "saved" MB/s do not necessarily correlate with saved J. There is no direct cost for a instruction, there is a severe overhead from the machine plainly being switched on. And it's not like you will always be able to "use" that saved MB/s for something else.
And you entirely neglect the "cost" of optimizing. Alone the time spent looking at inefficient code probably cost more energy than all the actually saved energy by a single change.
Consider the time someone could have spent on something else, with significantly more impact.
About the second part of your comment, it's true that optimizing has a cost, and if I were creating a "yes" or "cat" program for personal use it would obviously be pointless to optimize. But if it's a program ran by millions of people often (probably more true of "cat" than "yes"), it's not that obvious to me that millions of little savings cannot offset the time of one person optimizing.
Also very simple FreeBSD implementation [1] is not too slow on my 10-years old notebook:
> time yes test_string | dd of=/dev/null bs=1M count=65536
...
2646329200 bytes transferred in 1.709196 secs (1548288918 bytes/sec)
0.023u 0.850s 0:01.71 50.8% 5+166k 0+0io 0pf+0w
Firefox probably have used more CPU time while I was composing this comment - thanks to JS (in other tabs, HN is a rare example of a site which doesn't abuse my CPU). FF is almost always on the 1st line in top.If you'll check top/powertop and a typical desktop or a server you'll likely find better targets than 'yes' to reduce energy use.
[1] https://github.com/freebsd/freebsd-src/blob/main/usr.bin/yes...
The maintenance argument isn't a good enough one. It's a factor, just not a strong one.
There is no reason for every tiny bit of something as foundational as the os to just get better and better forever.
The only reason we write things down in the first place, is so that we can do the work once and then refer to it many times without having to re-create it each time. So there is very little argument for keeping a program small and simple like the first version.
Some, but just not much. Because black-boxifying that complexity is what writing (be it a legal document or a program) is for in the first place. Making a more sophisticated better performing version 2 of something is simply using the tools of writing that exist for no other purpose ultimately.
It does go the other way too. Version 3 could be to invest yet more brainpower into figuring out how to get the same performance in fewer operations. And on that day soneone will wonder if it's worth optimizing that when it already works fine and compute resources are infinite. The answer then as ow will be the same "Yes. Of course."
> using speed as the only justification for merging a patch is probably not the right thing to do for a bigger project
I don't get this. Speed is important. Energy efficience is important. Have we gotten so used to the bloat that performance and energy saving must be disregarded?
You know, if there was a SFW version, it would be nicely illuminating about the complexity for getting similar stuff done and handling edge cases and whatnot.
That said, I think the Plan 9 version could use a few comments to decrease the cognitive load, since individual bits of code felt more approachable to me in the GNU version.
Though with the code itself being shorter, one liners or even just a few lines at the top of the function definition could be sufficient.
Similarly, I tried writing a naive version in node.JS a few years ago and it was several orders of magnitude slower than coreutils…, which is usually the case whenever I try and clone one of the “simple” utilities from coreutils.
"[urandom] is designed for security, not speed, and is poorly suited to generating large amounts of random data"
`cat /dev/zero` is pretty much as fast as `yes` on my system however (2.1GB/s)
[1] https://adamdrake.com/command-line-tools-can-be-235x-faster-...
I can't see that any of this optimisation is actually useful in the real world. Does yes actually need such a high throughput? Would anything suffer if it didn't? Probably not. Still fun though.
> *By far the best score so far is by @ais523 - generating FizzBuzz at a throughput that seems to average somewhere around 54-56GiB/s.
The best 'yes' in the linked article gets about 3GiB/s. To be fair: the optimized fizzbuzz is insane.
My entry outputs 28 TB/s. (I…may have done some creative interpretation of the rules.)
Impressive out of the box thinking.
Isn’t filling a buffer going to be a huge waste of CPU for the common case where the command only needs to provide one or two “yes”es? (Common case as in, its intended use of working around scripts that interactively prompt you to continue, etc…)
2. Filling a buffer has much less CPU (power) cost as it doesn't involve context switching like I/O.
However, the author had the opposite results. I wonder if the architecture makes a difference.
And let them rip.
This is how I thermally stress CPUs but I usually do `yes speed | head -n 8 | xargs -n 1 -P 0 openssl` instead.
https://github.com/tud-zih-energy/FIRESTARTER
With my system hooked up to a watts up (but including the monitor and a couple of other things) yes > /dev/null on each core gets a bit above 42w, openssl speed on each core occasionally gets above 45w and running FIRESTARTER for a bit gets above 57w (on an i5-6260U (NUC) with hyperthreading disabled, 15W TDP with some attempted power restraint in the BIOS, firefox on decent websites like HN tend to use about 31w and I think that is something like 10-15w (I think 12w but it has been a bit) measuring the computer only).
FIRESTARTER has some evolutionary algorithms as well, but at least on my CPU after hours they were still doing much worse than the default. I was wondering why there wasn't much discussion of BIOS power settings that I could find and after some testing found out that they are not effective at restraining max power use (I forget if they had any effect on typical power use either but I don't think it was much if they did). Also, the integrated GPU can use more power than the CPU and can't be limited. For that GpuTest is handy (but not open source):
https://www.geeks3d.com/gputest/
For me on the internal GPU, Pixmark Piano and Furmark use the most power, either can get 60-61w and adding FIRESTARTER in the background only adds another w or two.
Similarly, checking temps with turbostat (PkgTmp), 2x yes seems to max out about 70, testing one of the higher power openssl tests on each core reaches 75, FIRESTARTER alone quickly reaches 80 and slowly ramps up to 90, and adding gputest got up to 96. Interstingly, the max temperature is reached when the CPU dethrottles too quickly after the gputest is done. It takes a few gputests alone in a row to get into the 80s (I got bored after 2x each alternating Furmark and Piano that hit 83). Similarly, looking at power useage with turbostat the max PkgWatt with yes is 8.26, with openssl speed ecdsa 9.88, with FIRSTARTER alone 17.62, + gputest 19.97 (or gputest alone, seems unreliable).
Anyway, this is a long diversion to say that even fancy yes or openssl speed tests are not that great as CPU stress tests :).
It takes the useful insights from this article and packages them back into a Python program. But there's so little actual Python code running that the efficiency relative to a compiled language is swamped by all the data copying to the kernel.
vmsplice()ing is probably the next step in speed, however, vmsplice apparently doesn't fit well into the semantics of python (or probably rust) programs. trying to reduce the number of python bytecodes by calling `os.writev(1, many_bufs)` actually harmed performance, not sure why.
Once we found one logged in and decided to add 'yes' as the last line of his bash login.
That day we learned, via an angry phone call, that you can't Ctrl^C out of 'yes' via a remote shell. Today after reading this I'm assuming it's because of the way it floods the buffer.
Poor guy had to call the one dude with root access to fix it for him.
How is GNU `yes` so fast?
When I returned to work in the morning, 2 of the nodes were unhealthy because their hard drives were full.
A couple years later I was load testing Datadog logging from a cluster and also used yes (no logs on disk this time).
I used the daily log quota for the entire Datadog account in 30 minutes.
#include <libc.h>
#define BUFSIZE (512*1024)
static char buf[BUFSIZE];
int main(void){
memset(buf,'y',BUFSIZE);
while (1) {
write(1,buf,BUFSIZE);
}
return 0;
}https://github.com/coreutils/coreutils/blob/master/src/true....
Busybox `true` is 38 lines
https://github.com/mirror/busybox/blob/master/coreutils/true...
The simplest implementation is 0 bytes