Carefully but Purposefully Oxidising Ubuntu
discourse.ubuntu.com
discourse.ubuntu.com
I'm not any more a fan of POSIX locales than the next person[1], but AIUI, that seems a likely requirement for uutils to be used in a distro like Ubuntu.
I'd be curious how they plan to address this. At least from my perspective, unless uutils has already been designed to account for locales from the start (I don't know if it has), it seems likely that a significant investment of time will be required to add support for it.
[1]: https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
Is performance a frequent rationale for rewriting C applications in Rust?
But that's at least partially a maintainability argument, not just a performance one. Rust can make achieving higher levels of performance easier and less risky than doing so in C or C++ would have, but you do still have to work for it a little, it's not going to be magically faster.
$ curl -LO 'https://burntsushi.net/stuff/subtitles2016-sample.en.gz'
$ gzip -d subtitles2016-sample.en.gz
$ time rg -c 'Sherlock Holmes' subtitles2016-sample.en
629
real 0.099
user 0.063
sys 0.035
maxmem 923 MB
faults 0
$ time LC_ALL=C grep -c 'Sherlock Holmes' subtitles2016-sample.en
629
real 0.368
user 0.285
sys 0.082
maxmem 25 MB
faults 0
$ time rg -c '^\w{42}$' subtitles2016-sample.en
1
real 1.195
user 1.162
sys 0.031
maxmem 928 MB
faults 0
$ time LC_ALL=en_US.UTF-8 grep -c -E '^\w{42}$' subtitles2016-sample.en
1
real 21.261
user 21.151
sys 0.088
maxmem 25 MB
faults 0
(Yes, ripgrep is matching a Unicode-aware `\w` above, which is why I turned on GNU grep's locale feature. To make it apples-to-apples.)Now to be fair, you did say "usually." But actually, sometimes, even when functional parity[1] has been achieved (and then some), perf can still be wildly improved.
[1]: ripgrep is not compatible with GNU grep, but there shouldn't be much you can do with grep that you can't do with ripgrep. The main thing would be stuff related to the marriage of locales and regexes, e.g., ripgrep can't do `echo 'pokémon' | LC_ALL=en_US.UTF-8 grep 'pok[[=e=]]mon'`. Conversely, there's oodles that ripgrep can do that GNU grep can't. For example, transparently searching UTF-16. POSIX forbids such wildly popular use cases (e.g., on Windows).
ripgrep gains comes from a ton of brilliant optimizations and strategies done by it's author. They wrote articles about such tricks.
I didn't say it was, and this isn't even remotely close to my point. The comment I was replying to wasn't even talking about Rust versus C. Just new tools versus old tools.
> ripgrep gains comes from a ton of brilliant optimizations and strategies done by it's author. They wrote articles about such tricks.
I know. I'm its author!
I was replying to this:
> Is performance a frequent rationale for rewriting C applications in Rust?
But I now realize your message was much more specific so I stand corrected with regards to context. Your point was indeed different.
In the blink of an eye to search a couple of gigabytes.
I just checked and did a full search across 100 gigabytes of files in only 21 seconds.
The software is fantastic, and moreover it goes to show what our modern hardware is capable of. In these days of unbelievable software waste and bloat, stuff like ripgrep, dua, and fd reminds me there is hope for a better world.
ls, chgrp, chown, etc are applications that the current user will use to interactively work with them. Or maybe use in a shell script. They are not to be exposed to malicious inputs, so any bugs are usually not security problems, because you can't really do anything that the user couldn't do anyways. There is no security boundary, so nothing to secure there.
However, if you e.g. write your webserver as a shell script or using shell commands you are doing it wrong and you deserve the evil things an attacker does to you.
Also, very often, the insecurity is codified in the relevant standards such as POSIX, so you cannot really fix anything without breaking everything. E.g. newlines and various kinds of whitespace characters in filenames are a huge pain to handle safely in shell commands and especially shell scripts, if possible at all. Because shells do split things at any whitespace (unless quoted properly, which is hard to impossible to do) and shell commands usually separate their stdin/stdout at newlines (unless you know about -print0 and everything in your pipe does as well...). None of this is fixed by rust, everything could be fixed by throwing away existing standards first, but then the language doesn't matter.
I dunno, I'd expect that you could
ls -l untrusted.exe
mv untrusted.exe untrusted.exe.donotrun
chmod 600 untrusted.exe.donotrun
sha256sum untrusted.exe.donotrun
without anything bad happening. That said, I'd like some evidence that GNU's coreutils aren't already safe; I suspect that their limited scope means there's minimal exposure even when operating on untrusted inputs (note that in my example there, only hashing the file actually involves its contents).That mv and chmod might operate on a whole different file due to a symlink or hardlink being in place, and due to the inherent race condition between all those operations that just use the file name. Also, mv might overwrite an existing file.
sha256sum will run escaping on the filename output, thereby making automated comparisons fail if you are unaware of that feature. On the other hand, when turning off escaping with -z, you need to know how to handle the \0 separator.
Not that GNU coreutils must be vulnerable to this. But I suppose they are explicitly hardened against such inputs, despite the C language not providing any guard rails.
> Long-term, my concern would be that this may somewhat muddy the picture for which packages need substantive fixes. If it is extremely easy to just revert, what is the benefit to switching?
??? Is this not the ideal situation? It provides low friction both for moving to the new cool thing and moving back to the existing tools when you have software that hard depends on GNU coreutils. You want the changeover to be high risk because it's hard to undo? I guess that's one way to force yourself to commit but the real users on the ground won't be happy when going to the LTS is substantially more work.
This would be 3 different symlink managers in Ubuntu all used for different sets of software. The alternatives system at least has the benefit of integrating tightly with apt.
How this will work on CPU architectures other than x86 and Arm? Ubuntu also supports ppc64le and IBM s390. Is LLVM usefully able to built binaries from Rust code for those architectures now?
Most likely none.
The most likely bugs you'll encounter are due to Fil-C using musl, not glibc (that always leads to some incompatibilities for GNU code).
> Are you subsetting the language
No.
> or do those libraries/applications run into some "benign" UB under normal operation that is caught by Fil-C?
Sometimes, but rarely.
source https://github.com/pizlonator/llvm-project-deluge/blob/delug...
$ hyperfine "LC_ALL=C sort < subtitles2018.en" "LC_ALL=C gsort < subtitles2018.en"
Benchmark 1: LC_ALL=C sort < subtitles2018.en
Time (mean ± σ): 2.007 s ± 0.005 s [User: 1.940 s, System: 0.064 s]
Range (min … max): 1.997 s … 2.015 s 10 runs
Benchmark 2: LC_ALL=C gsort < subtitles2018.en
Time (mean ± σ): 898.5 ms ± 9.7 ms [User: 2795.8 ms, System: 93.6 ms]
Range (min … max): 875.0 ms … 906.9 ms 10 runs
Summary
LC_ALL=C gsort < subtitles2018.en ran
2.23 ± 0.02 times faster than LC_ALL=C sort < subtitles2018.en
Info about the tools: $ which sort
/usr/bin/sort
$ which gsort
/opt/homebrew/bin/gsort
$ brew info coreutils | head -n3
==> coreutils: stable 9.6 (bottled), HEAD
GNU File, Shell, and Text utilities
https://www.gnu.org/software/coreutils/
There are other tools in coreutils for which this applies as well.The GNU tools have overall been pretty heavily optimized. Why do you think they did that if it literally didn't matter? Just for shits & giggles?
You should try that benchmark with sort compiled with Fil-C.
But what you said is (emphasis mine):
> For these tools, you won’t notice.
> None of these tools that they’re oxidizing is compute bound
just blatantly wrong and a >2x difference in perf is absolutely relevant and something I would notice personally. Maybe I'm the only one who likes sorting to be as fast as possible, but I'd guess not.
> You should try that benchmark with sort compiled with Fil-C.
Feel free to post a complete MRE for doing this and I'd be happy to run it.
The interesting question - going back to my original post - is what the Fil-C slow down would be. Just because a program takes 100% CPU doesn't mean it'll experience bad overheads when compiled with Fil-C.
I don't know what you mean by "complete MRE". You can download the Fil-C binaries and point coreutils' configure script at the compiler and see what happens. It's not hard.
I showed you in my original comment. Both are C. One is the `sort` that comes with macOS, as shown at `/usr/bin/sort`, and the other is from GNU coreutils.
> Also, it's just one benchmark of sort, so for all we know the two sort implementations really have the same perf if you test a broader set of cases.
Sure, you're welcome to come up with a more comprehensive benchmark suite to support YOUR claim that tools like `sort` are not "compute bound" and you won't "notice" a difference in speed. I agree that 1 benchmark is insufficient to make generalized claims about performance. But it is very obviously better than 0 benchmarks (which is the number you have provided to support your claim).
See also: https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)
In the interest of over-communicating, please note the qualification I gave originally: for files cached in RAM.
> I don't know what you mean by "complete MRE". You can download the Fil-C binaries and point coreutils' configure script at the compiler and see what happens. It's not hard.
Cool, then you should be able to show me a transcript of the precise commands necessary to do this very easily!
There's enough data from the Android ecosystem that it's much better to focus oxidisation on new software instead of old.
Whereas there's just one thing to target with X11. Far more consistent, less buggy overall. At least up to now. But with megacorps like IBM/Red Hat having their employees abandon X11 entirely for Gtk5, etc, it's gonna be rough going everywhere.
Canonical is just so gross at every level. I saw a job posting in my area recently and it was the slimiest corpo-speak job post I've seen in.. well, a couple of weeks at least.
I'm all for Rusting more-or-less all the things, but Rust isn't actually magic or anything. Things aren't automatically better just because Rust. And as someone who keeps at least modest track of CVEs and such, these utilities aren't exactly throwing CVEs out left right front and center. I don't like the amount of C in the world but being run billions of times a day in every conceivable environment is itself a really, really good test. We well know from history, still not always perfect, but honestly these aren't the things that need a priority Rusting; it's the code that doesn't fit that description that really needs it.
By all means, rewrite it in Rust, but don't go slamming it in to Ubuntu. Debian unstable. Gentoo. Something other than what I put on my boomer-generation, stereotypical "knows nothing about computers" father-in-law's laptop so I don't have to support him every week.
No
But it's still rolling-only code. And having to change it for every ubuntu release (even if per-release everything is targeted to the static version of rustc in Ubuntu) is silly since cargo editions won't save you. And since anyone actually compiling rust programs will have had to update their rustc constantly they'll have to do some containerization or something to compile the $version-written tool for $ubuntu-release-number.
Rust 1.82 was released in October last year. It looks like this gets bumped periodically, but it isn't pulled to latest and presumably a bugfix release, if they were doing those, wouldn't change MSRV so that seems basically fine?
Remember when the switched they shell to dash? Not to mention, Upstart, Mir, Unity, Snap.
But I agree in principle
> Debian Almquist shell
So yes:)... Although apparently Ubuntu was first to adopt it as /bin/sh
upstart was used by RedHat must have been good enough, before they NIH it Unity was actually better then GNOME at that time, but in the end RH NIH everything except Qt/KDE
+1. Still is.
What Unity really demonstrated to me is that despite the constant "Vim gives me editing at the speed of light" mantra, most Linux users do not in fact know how to drive a Windows-like or Windows-compatible desktop GUI with the keyboard.
Knowing all the keyboard controls for one editor does not help someone with anything else. If anyone is a keen keyboard user and has the time and inclination to learn a single app's keyboard UI, then it's a good use of their time to learn the keyboard UI of the underlying OS.
There is a standard. It works on Windows and has done since I first met Windows in 1988 with Windows 2.01, and it still does. Decades ago most Linux WMs honoured it, and it still works today in Xfce, Metacity, much of LXDE and LXQt, and some of MATE. And in Unity.
Not just simple stuff like Alt+F4 closing a window, but richer stuff like Super+1--9 opens the nth app pinned to the taskbar or panel.
If you use the keyboard, Unity works great and the cosmetic resemblance to Mac OS X can be largely ignored.
If you don't... "OMG it's so different aaaaargh where am I noooooo I hate this!"
GNOME >3 threw keyboard UI in the bin along with 30Y of HCI R&D.
KDE... never bothered to learn to properly replicate Win98 so it invents its own. Or, more typically of KDE, half a dozen slightly different versions.
On the other hand, I'd prefer for basic system tools to be provided by an established and well-funded project, and not just switch to something written in Rust because Rust is the hype these days.
Why?
The GNU project has an extremely utopian and unrealistic vision of open-source.
Indeed, the Chief GNUisance himself is very adamant about that[0]. Of course, the GNU project is a lot bigger than just RMS, and we shouldn't condemn the project and its struggle against non-free software just because RMS is problematic.
[0]: <https://www.gnu.org/philosophy/open-source-misses-the-point....>
For the details on the author of the Stallman Report -- Drew Devault -- who tried to pretend he didn't write it, see https://dmpwn.info/
Also, please consider not using the term "open-source". It's a term designed to pander to businessmen, who are afraid of the Free Software movement's goals, and to muddy the waters as to what rights you should have with such software.