An update on rust/coreutils
sylvestre.ledru.info
sylvestre.ledru.info
[0] https://www.reddit.com/r/unix/comments/6gxduc/how_is_gnu_yes...
[1] https://github.com/coreutils/coreutils/blob/master/src/yes.c
[2] https://github.com/openbsd/src/blob/master/usr.bin/yes/yes.c
On my system it manages 20GB/s vs 6 GB/s (piping through pv into /dev/null).
These OSs distribute their own coreutils.
> yes(1) turns out to be really slow without that weirdness[0]
I'm not sure there is a workflow where yes writing slower than 10.2GB/s is the bottleneck. Being able to go so fast is a nifty engineering feat, but hardly a necessity IMHO.
No reason except the complexity of maintaining huge amounts of code to support dead processors and operating systems, or ones that are so rare that they might as well be dead.
brew install coreutils
etc., which I suspect is sort of what Spivak was referring to, and not that MacOS ships with them.I think the GP was getting at the fact that Rust/rustc handles many of these edge cases at the language/compiler/standard-library level, instead of at the application level. Which would mean, at least in theory, the new code would actually not have to be as ugly as the original in many cases.
Since then, I’ve discovered (thanks to Hacker News) Decoded: GNU coreutils¹, a “resource for novice programmers exploring the design of command-line utilities” but unfortunately, I no longer have the capacity or free time to spend on coding.
¹ https://www.maizure.org/projects/decoded-gnu-coreutils/index...
Outside of the test failures and missing features, the only reason one might not want the coreutils is the size.
Right now, cat on my system is 44K large, while the default output size for the cat executable in release mode is 4.4 megabytes in release mode. If you enable a bunch of options in Cargo.toml to reduce the default bloaty settings (lto = true, codegen-units = 1, strip = true, debug = false), you get that down to 876K. Which is still 20 times larger than the native cat. For true, it's a similar story. 40K vs 812K.
To give an answer to your other question: a lot of the additional code size is a one time cost. That means that the relative difference is the largest in these cases, for binaries with little logic of their own. The more complex the program gets, the less the relative difference is. However, even if you don't account for these one time costs, Rust creates larger programs than C, due to heavy use of generics compared to C.
It's a lovely library. Generated help text with colors and line wrapping based on your terminal width, informative error messages that suggest corrections for typos, and a declarative model that lets you express subtle relationships between options. You don't get that from getopt.
But it's also pretty large. https://github.com/rust-cli/argparse-benchmarks-rs has some numbers.
(Disclaimer: I wrote one of the competitors on that page, lexopt. But I also use clap, depending on the project, and I'm happy with it.)
One of the most attractive features of Linux is how it suits a large array of hardware. Replacing C solutions with some bloated bandwagon alternative is a bad idea.
I'm sure it's possible to cut it more. It'll never be tiny, but it can be smaller than it is now.
That doesn't seem so bad, the GNU coreutils Debian package has an Installed-Size of 17.9MB. Though there might be hidden drawbacks I don't know about.
That's a definitely acceptable size, thanks for pointing that out. Yeah a bunch of symlinks should solve the issue. If the binary is kept cached in RAM at least, there should be no overhead in starting it. I think it's pretty rare to have environments that are so memory constrained that they can't hold onto 8 MB of RAM.
Note that not all targets are that small though. Musl has indeed 8.3 MiB, but in coreutils-0.0.12-x86_64-unknown-linux-gnu.tar.gz, the size of the coreutils executable is 12.6 MiB. On the bright side, on the same target, a local build with the cargo features to reduce bloat enabled gives only 7.1M. All these numbers are acceptable.
For example, if ls had a --shell option that would write out file information as quoted eval'able variables, or even JSON, or anything that was easily and reliably parsable it would remove a huge portion of scripting headaches and errors.
This will help making sure that we aren't regressing (more) ;)
Instead of 30 to 60 patches per month, we jumped to 400 to 472
patches every month. Similarly, we saw an increase in the
number of contributors (20 to 50 per month from 3 to 8).
More contributors = more hands to fix issues and tune performance. I believe this is an absolute win.Probably not. In general, functionalities are enabling or disabling a behavior when doing an operation. In the code, it translates most of the time by a simple if else. For example, adding new options usually looks like this PR: https://github.com/uutils/coreutils/pull/2880/files
The performance wins are usually produced by using some fancy Rust features.
Also performance is not always the only factor to consider: you can optimize a program for speed, but also for RAM usage, or the size of the binary itself.
In program like Coreutils to me is more important that the programs are small than the rest. Typically you use a lot of commands in a script, to do some trivial operations (the input of the program is usually small), thus simpler programs (that have less startup time) are usually better.
Then use BusyBox/Matchbox....because..well that's the whole point of those projects.
Also ‘if’ causes branching which costs a small amount of CPU overhead. So it’s not a free operation and can quickly add up if you needed inside a hot path.
This kind of thing is where Rust really shines. The ecosystem was built post-unicode, so things tend to support it by default. Ripgrep for example has been unicode aware from the beginning, and you have to opt-out if you don't want that.
For example suppose you're implementing case-insensitive sort, you write a test and tweak it slightly so that it passes as you expected. I come along and write a slightly faster case-insensitive sort, and mine fails. Upon examining the test I discover it thinks I ought to sort (rat, doG, cat, DOG, dog, DOg) into (cat, doG, dog, DOG, DOg, rat) but I get (cat, doG, DOG, dog, DOg, rat) my answer seems, if anything better and certainly not wrong but it fails your test.
So while I agree that without historical context and compatibility for human consumption your way of sorting is probably fine, you could wreak major havoc when trying your sort as a drop in replacement. If you only have a few users change management for something like this is relatively easy. The user base of coreutils? Think twice if you want to try that change management.
It's the other way around, though.
Notice that the "failed" example exhibits stability which you claimed was desirable, while neither exhibits sorting by case, this is after all a case-insensitive sort, the "successful" example is just swapping some of the list items for whatever reason, maybe it was how their chosen algorithm worked, maybe it's a bug, they wrote the test so they get to fill out a "correct" answer that matches their behaviour.
Now, striving to pass such tests gets you bug-for-bug compatibility which is what you want if you're an emulator, but the GNU project started out deliberately not doing bug-for-bug because it means people accuse you of copying, and so I don't see why this project should be different.
Unfortunately I couldn't find a list of failed/successful tests, if that's available I'd be happy if someone linked it
The "Run GNU tests" is probably what you are looking for.
So first get it to 100% compatibility, then, and only then concentrate on performance. Because if you don't do it that way you will end up foregoing compatibility because you'd have to say goodbye to your beautifully tuned code that unfortunately can't be 100% compatible. As long as it does not pass all the tests: it does not work. Even if it works for some selected cases, the devil is in the details, and those details can eat up performance like there is no tomorrow.
- Make it once.
- Make it right.
- Make it fast.
"Make it right" as the first step can trick an unseasoned developer into never finishing a prototype. I don't mind the first iteration of something being sloppy and then pursuing correctness with an existing solution to the problem in hand.I'm curious if you've got a sense of the interplay between prototyping and correctness similar to your sense of the interplay between performance tuning and correctness? Any thoughts?
While I do generally agree, it can easily become an excuse for not thinking things through from the start.
If you have the wrong architecture it might be very difficult or even impossible to optimize performance later on.
Also: If working on a project for a client, performance can be difficult to sell when the feature already works. But that feature might break when it's put under load.
My advice would be to always have performance in mind. But otherwise to stay away from "optimizations" until they are needed.
God level programmers (of which I've met exactly one and know about one other over my whole career) can do this the first time around.
What I've seen many times is "best practices" like this being thoughtlessly applied.
A react component that rerenders 20 times, but passes the test. A database with no indices, no problem (until the number of records grow).
Things like that can easily be defended with "no need to optimise prematurely" or "I wrote the least amount of code to make the test pass".
What helps in those cases: to have a good idea of how long you think something should take and then to verify that it indeed is within an order of magnitude of that first estimate. If it is much slower you are probably in trouble, if it is much faster than you will have to check if you are really doing all the work that you should be doing.
Colin Wright (also on HN) wrote: "You can't make computers faster, but you can make them do less work". The corollary is that if your program is faster than expected that may be because it is doing less work!
Is this one of the intentions of this team? It sounds like it could make "scripting" in rust very nice if all of the CLI functions you're used to exist as language libraries.
- Refactoring out buggy, convoluted, highly-backwards-compliant code for cleaner more practical code? - Reimplementing code in a manner which simplifies previously complicated bottlenecks that were there in response to bug reports? (and whose simplification potentially risks reintroducing said bugs again)
In all honesty, I would expect that "reimplementing coreutils" as above would have resulted in a speedup even if it was written in c again.
Am I wrong? Is there something about rust that inherently leads to an increase in speed which one could not ever hope to obtain with clean, performant c code?
Many libc functions are also much slower than they could be because of POSIX requirements and being a shared library. For example, libc’s fwrite, fread, etc. functions are threadsafe by default and acquire a globally shared lock even when you aren’t using threads (you can opt out, but it’s quite annoying and non-standard) which makes them horribly slow if you’re doing lots of small reads and writes. Because libc is a shared library, calls to its functions won’t get inlined, which can be a major performance issue as well. By comparison Rust’s read and write primitives don’t need to acquire a lock and can be inlined, so a small read or write (for example) could be just a couple of instructions instead of what a C program will do, which is a function call (maybe even an indirect one through the PLT) and then a lock acquisition, only to write a few bytes into a buffer. That’s a lot of overhead!
And finally, Rust’s promise of safe multithreading no doubt encourages programmers to write code that utilizes threads in situations where only the truly courageous would attempt it in C.
Depending on how much a rewrite it is, GNU's intellectual property may be enforced at some point. But that's a tricky question for a lawyer. If it were me, I'd have asked GNU first.
In the real world, corporate engineering middle managers do not understand the GPL and do not understand how it does and doesn’t restrict them. They avoid it rather than taking the time to learn it.
It supports lots of different legacy encoding.
It has two advantages on the GNU version apart from just memory safety with Rust.
1) Ease of portably building. Just type ‘cargo build —release’ and it will just work, even on Windows.
2) MIT License. A company can take these and distribute them as part of a commercial offering without having to worry about GPL compliance.
It's also going to be _really_ hard to be more portable than GNU coreutils, when it comes to platforms it's available on.
I do embedded linux for work and while OpenEmbedded makes it pretty easy to manage license obligations, it’s always a pain in the ass to deal with GPL code in a larger team where people will find any excuse to bikeshed about their misunderstandings of the GPL. For questionable historical reasons, my current job manually whitelists every GPL package we ship, so it being MIT makes things less abrasive to teams like mine.
That’s an advantage for everyone, because it makes me more likely to use it and it makes me more likely to submit bug reports and fixes.