Debian Running on Rust Coreutils
sylvestre.ledru.info
sylvestre.ledru.info
Replacements like rg or fd are often not even compatible to the unix originals let alone the GNU extensions. Yes, some of the GNU tools are badly in need of UI improvements, but you'd have to abandon compatibility with a great deal of scripts.
I think both a coreutils rewrite as well as some end user facing software like rg has its place on modern unix systems. I'm a very happy user of rg! But I'd like some more respect for tradition by some of those tools. For example, "fd" in a unix setting refers to file descriptors. They should rename IMO.
Of course, ripgrep is still ultimately a lot more similar to grep than it is different. Where possible, I used the same flag names and prescribed similar behavior---if it made sense. Because it's good to lean on existing experiences. It makes it easier for folks to migrate.
So it's often a balancing act. But yes, POSIX compatibility is certainly a non-goal of ripgrep.
https://github.com/BurntSushi/ripgrep/blob/master/GUIDE.md#c...
(The other killer reason is the blitzing speed.)
EDIT: So I've quickly looked into it and it seems nobody did an extensive comparison to the grep feature set or the POSIX specification. If I have some time later this week I might do this and check whether something like this would be viable.
There are lots of incompatibilities. The regex engine itself is probably one of the hardest to fix. The incompatibilities range from surface level syntactic differences all the way down to how match regions themselves are determined, or even the feature sets themselves (BREs for example allow the use of backreferences).
Then of course there's locale support. ripgrep takes a more "modern" approach: it ignores locale support and instead just provides what level 1 of UTS#18 specifies. (Unicode aware case insensitive matches, Unicode aware character classes, lots of Unicode properties available via \p{..}, and so on.)
Thanks anyway, I enjoy rg quite a lot :)
To give backreference support, ripgrep can optionally use PCRE. But PCRE already comes with its own drop in grep replacement...
But still, that only solves the incompatibilities with the regex engine. There are many others. The extent to which ripgrep is compatible with POSIX grep is that I used flag names similar to GNU grep where I could. I have never taken a fine toothed comb over the POSIX grep spec and tried to emulate the parts that I thought were reasonable. Since POSIX wasn't and won't be a part of ripgrep's core design, it's likely there are many other things that are incompatible.
A POSIX grep can theoretically be built with a pretty small amount of code. Check out busybox's grep implementation, for example.
While building a POSIX grep in Rust sounds like fun, I do think you'd have a difficult time with adoption. GNU grep isn't a great source of critical CVEs, it works pretty well as-is and is actively maintained. So there just isn't a lot of reason to. Driving adoption is much easier when you can offer something new to users, and in order to do that, you either need to break with POSIX or make it completely opt-in. (I do think a good reason to build a POSIX grep in Rust is if you want to provide a complete user-land for an OS in Rust, perhaps if only for development purposes.)
1. Distributions could adopt rg as default and ship with it only, adding features at nearly no cost
2. The performance advantage over "traditional" grep
Number 1 is basically how bash became the default; since it is a superset of sh (or close enough at least), distributions could offer the feature set at no disadvantage. Shipping it by default would allow scripts on that distribution to take advantage of rg and, arguably, improve the situation for most users at no cost.
If one builds two programs in one with a switch, you're effectively shipping optional software, but in a single binary, which makes point 1 pretty moot. If you then also fall back on another engine, point 2 is moot as well - so the only point where this would actually be useful is if rg could become a good enough superset of grep that it would provide sufficient advatages (most greps _already_ provide a larger superset of POSIX, though). Everything else would just add unnecessary complexity, in my opinion.
But it would have been nice :)
I have been playing around with regex parsing through building parsers through parser combinators at runtime recently, no clue how it will perform in practice yet (structuring parser generators at runtime is challenging in general in low-level languages) but maybe that could pan out and lead to an interesting way to support broader sets of regex syntaxes like POSIX in a relatively straightforward and performant way.
The regex crate does a lot of literal optimizations to speed up searches. More than most regex engines in my experience.
ed : within 30s of learning about ripgrep I learned that VSCode's search routine calls ripgrep and this has just gifted me leverage to introduce Rust in my company... we're a vms shop.. searching for rust and vms gets you the delightful statement that since Rust isn't a international standard there's no guarantee it'll be around in.. oh about right now...the real reason why any language doesn't last long on vms is if the ffi implementation is shoddy. but that's another story that turns into a book the moment I try explaining how much value vms ffi brings... my brain suddenly just now thought "slingshot the critical systems devs outta vms on a Rust hook and things will happen ".
(longer explicative comment in profile - please do check it out if you'd like to know why this really made my day and some)
I do have plans to add an indexing mode to ripgrep[1], but it will probably be years before that lands. (I have a newborn at home, and my free-time coding has almost completely stopped.)
There is a long history of newer, better tools supporting backward compatibility to be a replacement for their predecessors:
- bash and zsh support backward compatibility to be replacements for sh.
- vim has compatibility mode to be a replacement for vi.
If the Rust implementation is not backward compatible, it should not be called "Coreutils".
Having tried to reproducibly package various software, the C layer (and especially the coreutils layer) are the absolute worst, and I wouldn't shed a tear if we started afresh with something more holistically designed.
If you have something better in mind, please implement it. It’d get used.
> If you have something better in mind, please implement it. It’d get used.
Nonsense. It's not a dearth of better options that causes C folks to clutch their autotools and bash scripts and punting-altogether-on-dependency-management; we've had better options for decades. Change will happen, but it will take decades as individual packages become less popular in favor of newer alternatives with better features--build hygiene will improve because these newer projects are much more likely to be built in Rust or "disciplined C" by developers who are increasingly of a younger generation, less wed to the Old Ways of building software.
I build and package golang/Python/rust/C/C++ binaries using GNU make, bash, and Debian's tooling. I have dependency management, parallel builds, and things are reproducible. I do it that way, because I need a glue that will deploy to VMs, containers, bare metal, whatever. I don't use it because I'm scared of other tooling. I use it because I haven't seen a better environment for getting software I work on, into heterogeneous environments.
I'm not attached to the old way; I'm attached to ways that I can be productive with.
To be clear, I'm not arguing that there are better tools for working around C/C++'s build ecosystem; I'm arguing that our lives will be better when we minimize our dependencies on those ecosystems.
For example, for the overwhelming majority of Rust and Go projects, there are explicit dependency trees in which all nodes are built roughly the same way such that a builder can do "cargo build" and get a binary. No need to understand each node's bespoke, cobbled-together build system or to figure out which undocumented dependencies are causing your build to fail (or where those dependencies need to be installed on the system or how to resolve conflicts with other parts of the dependency tree).
To be clear, that means I have devoted zero effort to it, and certainly it didn't require discipline, I have absolutely no expertise in the process, and my knowledge of the dependency tree is "gcc -o x x.c" does the job.
The most important thing is compatibility with humans. The main reason because I tend to use a pretty standard setup is because I know that these tools are the standard, I need to ssh into a server to resolve some issue, well I have a system that I'm familiar with.
If for example on my system I have fancy tools then I would stop using the classic UNIX tools, and then every time I would have to connect to another system (quite often) I would type wrong commands because they are different or totally not present or have to install the new fancy tools on the new system (and it's not always the case, for example if we talk about embedded devices with 8Mb of flash everything more than busybox is not possible).
To me the GNU tools have the same problem, they got you used to non POSIX standard things (like putting flags after positional arguments) that hits you when you have to do some work on systems without them. And yes, there still exist system without the GNU tools, MacOS for example, or embedded devices, some old UNIX server that is still running, etc.
Last thing, if we need to look at the future... better change everything. Let's be honest, if I could chose to switch entirely on a new shell and tools, and magically have it installed on every system that I use, it would probably be PowerShell. For the simple reason that we are no longer in the '80 where everything was a stream of text, and a shell that can manipulate complex objects is fantastic.
Absolutely. We had decades to work with a fairly stable set of tools and they are not going anywhere. Whoever needs them, they are there and likely will be for several more decades.
I am gradually migrating all my everyday DevOps workflows (I am a senior programmer and being good with tooling is practically mandatory for my work) to various Rust tools: ls => exa, find => fd, grep => rg, and more and more are joining each month. I am very happy about it! They are usually faster and work more predictably and (AFAICT) have no hidden surprises depending on the OS you use them in.
> we are no longer in the '80 where everything was a stream of text, and a shell that can manipulate complex objects is fantastic.
Absolutely (again). We need a modern cross-platform shell that does that. There are few interesting projects out there but so far it seems that the community is unwilling to adopt them.
I am personally also guilty of this: I like zsh fine even though the scripting language is veeeeery far from what I'd want to use in a shell.
Not sure how that particular innovation will explode but IMO something has to give at one point. Pretending everything is a text and wasting trillions of CPU hours in constantly parsing stuff is just irresponsible in so many ways (ecological included).
I suspect that the endless parsing may be even economical compared to the complex dance needed e.g. to call a Python method. The former is at least cache-friendly and can more easily be pipelined.
Though it's ok to be slow for a tool which sees mostly interactive use, and the scripting glue use. Flexibility and ergonomics trump the considerations of computational efficiency. So I expect that the next shell will burn more cycles in a typical script. But it will tax the human less with the need to inventively apply grep, cut, head, tail, etc, with their options and escaping / quoting rules.
Sure, but it was a different time. People there were down for anything that actually worked and improved the situation. Like the first assembler written in machine code, then the first C compiler being written in assembly, etc. People needed to bootstrap the stack somehow.
Nowadays we have dozens, maybe even thousands of potential entry points, yet we stubbornly hold on to the same inefficient old stuff. How many of us do REALLY need a complete POSIX compliance for our everyday work? Yeah, the DevOps team might need that. But do most devs need that? Hell no. So why isn't everyone trying stuff like `nu` or `oli` shell etc.? They are actually a pleasure to work with. Answer: network effects, of course. Is that the final verdict? "Whatever worked in the 70s shall be used forever, with all of its imperfections, no matter how unproductive it makes the dev teams".
Is that the best the humanity can do? I think not... yet we are scarcely moving forward.
> I suspect that the endless parsing may be even economical compared to the complex dance needed e.g. to call a Python method. The former is at least cache-friendly and can more easily be pipelined.
50/50. You do have a point but a lot of modern tools written in Rust demonstrate how awfully inefficient some of these old tools are. `fd` in particular is times faster than `find`. And those aren't even CPU-bound operations; just parallel I/O.
Another example closer to yours might be that I knew people who replaced Python with OCaml and are extremely happy with their choice. Both languages have no (big) pretense that they can do parallel work very well so nothing much is lost by migrating away from Python [for various quick scripting needs]. OCaml however is strongly typed and MUCH MORE TERSE than Python, plus it compiles lightning-fast and runs faster than Golang (but a bit slower than Rust, although not by much).
> Though it's ok to be slow for a tool which sees mostly interactive use, and the scripting glue use.
Maybe I am an idealist but I say -- why not both be fast and interactive? `fzf` is a good demonstration that both concepts can coexist.
> Flexibility and ergonomics trump the considerations of computational efficiency.
Agreed! The way I see it nowadays though, is that many tools are BOTH computationally inefficient AND not-ergonomic.
But there also reverse examples like GIT: very computationally efficient but it's still a huge hole of WTFs for most devs (me included).
Many would argue that the modern Rust tools are computationally less efficient because they spawn at least N threads (N == CPU threads) but the productivity gain earned from that more aggressive use of machine resources is IMO worth it (another close example is the edit -> save -> recompile -> test -> edit... development cycle; the faster that is, the bigger the chances that the dev will follow their train of thought until the end and will get the job done quicker).
---
So TL;DR:
- We can do better
- We already are doing better but the new tools remain niche
- Old network effects are too strong and we must shake them off somehow
- We are holding on to old paradigmae for reasons that scarcely have anything to do with the programming job itself
Stream of bytes is the only sane thing to do. Nothing keeps you from not having a flag in your cli programs to choose the input/output format. In fact many programs already has this and json seems pretty popular.
Having standard serialization is just gonna be boilerplate and unneccessary for many programs. Having user choosing how to interpret the input/output is the best way.
And for a lot of programs having to keep generating and parsing strings is just a bunch of unnecessary boilerplate.
> Stream of bytes is the only sane thing to do. Nothing keeps you from not having a flag in your cli programs to choose the input/output format. In fact many programs already has this and json seems pretty popular.
Having an API like this is a great idea... as a thin wrapper on top of a tool with a standardized serde protocol or binary file format. The different API's can be exposed as separate tools or as separate parts of the API of a single tool from a user's POV.
Furthermore: JSON is not just a stream of bytes nor text, neither is CSV, nor any other text format. Handling text as properly typed objects makes a lot of sense.
I have little experience with shells that can work with objects like Powershell. But I'e seen screenshots of new objects-based shell developed in Rust some time ago and it was early days still so I didn't actually even try it out, but looked downright great compared to the current stream of text-based shells of today in terms of ergonomics and capabilities.
I think the situation with GNU coreutils on Solaris/BSDs is a bit different: a lot of the time the BSD/Solaris tools on one hand and the GNU tools on the other were incompatible in some ways. Some flags would be different (GNU ls vs. FreeBSD ls for instance have pretty significant flag differences) and then you have things like Make who have pretty profound syntax differences between flavours.
As a result if you needed to write portable scripts you either had to go through a painful amount of testing and special casing for the various dialects, or you'd just mandate that people installed the GNU version of the tools on the system and use that. That's why you can pretty much assume that any BSD system in the wild these days has at the very least gmake and bash installed, just because it's used by third party packages.
So IMO people used GNU on non-GNU systems mainly for portability or because they came from GNU-world and they were used to them.
I know that there's a common idea that GNU tools were more fully featured or ergonomic than BSD equivalents, but I'm not entirely sold on that. I think it's mostly because people these days learn the GNU tools first, and they're likely to notice when a BSD equivalent doesn't support a particular feature while they're unlikely to notice the opposite because they simply won't use BSD-isms.
For instance for a long time I was frustrated with GNU tar because it wouldn't automatically figure out when a compression algorithm was used (bz, gz or xz in general) and automatically Do The Right Thing. FreeBSD tar however did it just fine.
Similarly I can get FreeBSD ls to do pretty much everything GNU ls can do, but you have to use different flags sometimes. If you don't take the time to learn the BSDisms you'll just think that the program is more limited or doesn't support all the features of the GNU version.
An other example out of the top of my head is "fstat" which I find massively more usable than lsof, which is a lot more clunky with worse defaults IMO. It's mainly that lsof mixes disk and network resources by default, while fstat is dedicated to local resources and you use other programs like netstat to monitor network resources. Since I rarely want to monitor both at the same time when I'm debugging something, I find that it makes more sense to split those apart.
I completely agree with you overall point btw, I just don't think that's in an either/or type of situation. Rewriting the existing standard utils in Rust could provide some benefits, but it's also great to have utils like ripgrep who break compatibility for better ergonomy.
That was the original feature that made me want coreutils on other systems.
But it's the behavior on GNU systems, and it's even the behavior of most applications using getopt_long on other non-GNU, non-Linux systems (because getopt_long permuted by default, and getopt_long is a de facto standard now). So it should be supported.
I'm talking about interactive command-line usage, for which the ability to put an argument at the end provides more user-friendliness.
I won't deny the convenience. (Technically jumping backwards across arguments in the line editor is trivial, but I admit I keep forgetting the command sequence.) But from a software programming standpoint, the benefit isn't worth the cost, IMO.
And there are more costs than meet the eye. Have you ever tried to implement argument permutation? You can throw together a compliant getopt or getopt_long in surprisingly few lines of code.[1] Toss in argument permutation and the complexity explodes, both in SLoC and asymptoptic runtime cost (though you can trade the latter for the former to some extent).
[1] Example: https://github.com/wahern/lunix/blob/master/src/unix-getopt....
I completely agree; most commands should behave the same on the command-line and in scripts, because many scripts will start out of command-line experimentation. That's one of the good and bad things about shell scripting.
> Have you ever tried to implement argument permutation? You can throw together a compliant getopt or getopt_long in surprisingly few lines of code. Toss in argument permutation and the complexity explodes, both in SLoC and asymptoptic runtime cost (though you can trade the latter for the former to some extent).
"surprisingly few lines of code" doesn't seem like a critical property for a library that needs implementing once and can then be reused many times. "No more complexity than necessary to implement the required features" seems like a more useful property.
I've used many command-line processors in various languages, all of which have supported passing flags after arguments. There are many libraries available for this. I don't think anyone should reimplement command-line processing in the course of building a normal command-line tool.
I personally don't think permutation (in the style of getopt and getopt_long, at least in their default mode) is the right approach. Don't rearrange the command line to look like all the arguments come first. Just parse the command line and process everything wherever it is. You can either parse it into a separate structure, or make two passes over the arguments; neither one is going to add substantial cost to a command-line tool.
So, this is only painful for someone who needs to reimplement a fully compatible implementation of getopt or getopt_long. And there are enough of those out there that it should be possible to reuse one of the existing ones rather than writing a new one.
Not so sure about that. ls has been changing its output format based on whether it is being used interactively or not for as long as I can remember at least. Both GNU and BSD versions.
So for instance if you run a bare `ls` from a script that outputs straight into the terminal you'll get the multi-column "human readable" output. Conversely if you type `ls | cat` in the shell you'll get the single column output.
It can definitely be surprising if you don't know about it but technically it behaves the same in scripts and interactive environments.
"for commands to not alter their behavior based on whether they're attached to a terminal or not"
ls is a clear counter-example of that.
I think the behavior is a good thing. I'm pushing back against this notion of what is "best practice" or not. It's more nuanced than "doesn't change its output format."
Things many programs do if attached to a TTY: add color, add progress bars and similar uses of erase-and-redisplay, add/modify whitespace characters for readability, refuse to print raw binary, etc.
Things some programs do, which can be problematic: prompt interactively when they're otherwise non-interactive.
Things no program does or should do: change command-line processing, semantic behavior, or similar.
Arguably ripgrep breaks this rule. :-) Compare `echo foo | rg foo` and `rg foo`. The former will search stdin. The latter will search the current working directory.
In any case, I bring this up, because I've heard from folks that ripgrep changing its output format is "bad practice" and that it should "follow standard Unix conventions and not change the output format." And that's when I bring up `ls`. :-)
No. Gmake, maybe. But not bash.
Being licensed under the GPL is an essential part of the GNU project and their philosophy. Of course everyone is free to do as they please. However, IMHO, if one appreciates the GNU project and the ideals it stands for, then, maybe, it could be preferable not to rewrite parts of it under weaker licences that go directly against their mission!
// a few global constants as used in the GNU implementationAlternately, to solve the problem yourself and mention it to someone and have them say "that's how GNU split does it, so why not?"
https://en.wikipedia.org/wiki/Chinese_wall#Reverse_engineeri...
Nitpick: they don't need to, they can just attribute the Rust developers if needed. They can also just take MIT code and publish it under GPL (reverse is not true).
http://undeadly.org/cgi?action=article&sid=20070913014315
I'm not sure that is even legal.
LLVM did it. Switching from "University of Illinois/NCSA" to "APL 2.0".
But not without some big discussions.
A similar thing happened to OpenOffice / LibreOffice: [0]
> OpenOffice uses the Apache License, whereas LibreOffice uses a dual LGPLv3/Mozilla Public license.
> For some legal reasons, then, anything OpenOffice does can be incorporated into LibreOffice, the terms of the license permit that. But if LibreOffice adds something, take font embedding, for example, OpenOffice can’t legally incorporate that code.
[0] https://hackaday.com/2020/11/02/openoffice-or-libreoffice-a-...
Having "won the desktop from Windows" and "finished a next-gen OS (Hurd)" was also part of that plan, but it didn't really pan out, did it?
I'm glad for MIT software since I can use it in more places I can use GPL (which companies often wont touch, in it's later license version).
In what context are you using Rust Coreutils? Windows? MacOS? BSD?
They allow them. I had cases where clients forbade some GPL code (banks, etc) - worse with GPL3.
It's the tragedy of the Commons.
ApacheV2, MPL, MIT, ISC and BSD licenses are gaining in popularity.
No, they'll just find a MIT/BSD or probably proprietary alternative.
That's also why some GPL software is dual licensed. GPL for the masses, and a proprietary license allowing you to do whatever without needing to follow the GPL if you can afford it.
I thought you meant they would just use copyrighted code that wasn't under a GPL (which would be just as illegal and probably more dangerous in terms of enforcement).
Corporations choosing to not use GPL software mostly don't know enough about it, to know it's ok. So they ban it outright.
No, they might not. Just like Microsoft, they might demand that you don’t release your plugins. Addidionaly, GNU software gives you the alternative of releasing your plugins. But you don’t have to, and they can’t make you do it.
arch - GNU
chgrp - Descended from chgrp introduced in Version 6 UNIX (1975)
comm - Descended from comm in Version 2 UNIX (1972)
dd - Descended from dd introduced in Version 5 UNIX (1974)
du - Descended from du introduced in Version 1 UNIX (1971)
factor - Descended from sort in Version 4 UNIX (1973)
head - Descended from head included in System V (1985)
join - Descended from join introduced in Version 7 UNIX (1979)
ls - Spiritually linked with LISTF from CTSS (1963)
mktemp - GNU
I have access to an incredibly high quality easy to use router firmware in the form of OpenWrt by virtue of GPLd coreutils. I wish FOSS developers weren't so afraid of it these days
+1. It's very sad to see user freedom being thrown away as a goal and replaced by software that becomes unpaid labor for FAANGs. Especially since the SaaS takeover.
So it is worth having some deeper thoughts about what the implications are for different licenses of coreutils. When does using coreutils create a derivative work that requires GPL?
"However, we must recognize that this strategy did not succeed for Ogg Vorbis. Even after changing the copyright license to permit easy inclusion of that library code in proprietary applications, proprietary developers generally did not include it. The sacrifice made in the choice of license ultimately won us little."[0]
[0] https://www.gnu.org/licenses/license-recommendations.html
The people in the past had all the time in the world to tinker and invent. Maybe I am mistaken though, past is usually looked through rose-tinted glasses right?
But the fact remains: nowadays answering the above questions is beyond my pay grade: in fact it's beyond anyone's pay grade. Services like GitHub are deemed a commodity and questioning that status quo is a career danger.
I really do wish we start over on most of the items you enumerated. But I am not paid to do it. In fact I am paid to quickly select tools and never invent any -- except when they solve a pressing business need and are specific enough for the organization; in that case it's not only okay but a requirement.
Beyond anything else however, we practically have no choice. If I don't host a new company project on GitHub I'll eventually be fired and replaced with somebody who will.
We here on HN have both the time and ability to set up things like CLI Git, and Matrix. But for a new language, forcing people onto esoteric (& superior) platforms makes them less likely to use them.
It would be nice if Matrix and self-hosted Git were the default, but when acquiring users/programmers is your goal, Rust doesn't have that luxury.
Rust uses Github, but could easily switch to a self-hosted platform if Microsoft became opposed to Rust's goals; (and yes, Microsoft is a Rust sponsor, but not an essential one). Cargo has support for alternate registries built in.
Some of community is on Reddit and Discord, but most technical discussion takes place on Discourse and Zulip. The official "user questions" forum is a Discourse instance. Most subcommunities forming around Rust projects use Zulip instances.
The Rust community uses proprietary services when convenient, but it's hardly dependent on them.
Disagree, just look at bors and the use of Azure CI. It would be a huge PIA to switch.
> hardly dependent on them
I wouldn't say hardly, I think it's more like kinda. GitHub's network effect is pretty strong. Compare the number of contributors to golang for example which is hosted on Google code.
I use it heavily and think it's extremely underappreciated, so instead of reinventing it, I would like to build on it. But - trying to extend the old C codebase is daunting. I'd even be happy with a reduced featureset that avoids the exotic stuff that is either outdated or no longer useful. The core syntax of a Makefile is just so close to perfect.
(I wrote about some of this in the remake repo: https://github.com/rocky/remake/issues/114 )
> The core syntax of a Makefile is just so close to perfect.
I'd argue what is close to perfect is much of the underlying model/semantics, and what is terrible is very much the syntax. I've long wanted to make something similar to Make but simply with better syntax...
target: prereq | ooprereq
recipe
In my eyes, the most important thing when building something that is complex is the dependency graph and it makes sense to make the syntax for defining the graph as simple as possible. I think the make syntax just nails it and most of the other approaches I have seen so far add complexity without any benefit. In fact, most of the complexity they introduce seems to stem from confusion on the side of the developer being unable to simplify what they're trying to express.At the level of variable handling and so forth, make is slightly annoying but manageable.
Anything beyond that - yeah, I'm with you, a lot of that is terrible.
Gnu Make is almost pathetic.
1 module Main(main) where
2 import Development.Shake
3 import System.FilePath
4
5 main :: IO ()
6 main = shake shakeOptions $ do
7 want ["foo.o"]
8
9 "∗.o" %> \out → do
10 let src = out -<.> "c"
11 need [src]
12 cmd "gcc -c" src "-o" out
Maybe there is a need for super sophisticated graph handling, but for most use cases the way more complicated syntax of Shake is not a worthy tradeof.
> Large build systems written using Shake tend to be significantly simpler, while also running faster. If your project can use a canned build system (e.g. Visual Studio, cabal) do that; if your project is very simple use a Makefile; otherwise use Shake.
For what it's worth, if I remember right, Shake has some support for interpreting Makefiles, too.
> [...] the way more complicated syntax of Shake [...]
For context, Shake uses Haskell syntax, because your 'Shakefile' is just a normal Haskell program that happens to use Shake as a library and then compiles to a bespoke build system.
Also:
> The original motivation behind the creation of Shake was to allow rules to discover additional dependencies after running previous rules, allowing the build system to generate files and then examine them to determine their dependencies – something that cannot be expressed directly in most build systems. However, now Shake is a suitable build tool even if you do not require that feature.
I had trouble with the one class in uni that used Haskell and I've been working in a Python shop for some years now, so I kept expecting to encounter some impenetrable section that would make my eyes glaze over and my hand close the tab. I was probably closest around page 9!
But the writing was excellent, and clear, and satisfying, and little concepts I was so close grokking kept catching my eye and pulling me back in until I understood them, then their neighbors, then the section, then I was done.
Guess I've picked up some things since college. Wish I could go back and take that class again. I think learning to use Rust's Option type in anger on a personal project helped me understand monads more than anything in that class.
Also, I'm happy to see the method described by the paper does seem to have become the official GHC build system [0].
0. https://gitlab.haskell.org/ghc/ghc/-/wikis/building/hadrian
Other than that, I think Ninja is perfect.
Small enough to even checkin into git. Your compiler, cmake and all the other tools you use is going to be a much bigger problem.
While ninja is great for many uses i would not recommend to use it for hand written rules. In fact any simple dependency-resolver like make or ninja will be lacking a lot of language context about includes and other transitive dependencies so you always want some higher level abstraction closer to the language as your porcelain.
Anyway, not having any other automagical behavior and hidden rules and the fact that it can do incremental and correct job re-runs even when the rules change is exactly why I like to use ninja.
Here's one such use case, maybe not the cleanest:
https://megous.com/git/p-boot/tree/configure.php
I especially love it in projects involving many different compilers/architectures/sdks at once (like when doing low level embedded programming), where things like meson or autotools or arcane Makefile hacks become harder to stomach.
As opposed to say cmake, bazel or other higher level abstractions where you just ask it to take all c-files in this directory and solve the rest.
Yes, ninja basically handles it for you, compared to what you have to go through when using Makefiles, to have autogenerated dependencies.
Yes, there are workarounds for some situations, but none of them work in all cases and no-one ever applies them consistently.
Spaces in filenames is a perfectly reasonable thing to want to do. They are supported on all three major platforms, but make just throws up it's hands and doesn't care. Yes, it is workable around, but it would be nice if it wasn't.
It's an annoyance with make overall, but if we're talking about the syntax specifically then yes I would call that a "major insurmountable problem".
It caused so many issues. But mostly with developer tooling, like Make, that assumes these sorts of things. Most programs worked just fine.
Well there's your problem right there. That kind of shit needs to stop, and making companies that perpetrate it less able to develop software (and thus less profitable) is one of the only things people who don't directly interact with such a company can do to discourage it. (Not that that's in any way a good strategy, but to the tiny extent it has any bearing on makefile syntax in the first place, it's a argument against supporting spaces in filenames.)
It's a build-system in Rust with some really cool features (and an even simpler syntax).
If you haven't, I'd suggest having a look at pmake (sometimes packaged as 'bmake' since it is the base 'make' implementation on BSDs) - the core 'make' portions are still the same (like gmake extends core 'make' as well), but the more script-like 'dynamic' parts are much nicer than gnumake in my opinion. It also supports the notion of 'makefile libraries' which are directories containing make snippets which can be #include's into other client projects.
freebsd examples:
manual: https://www.freebsd.org/cgi/man.cgi?query=make&apropos=0&sek...
makefile library used by the system to build itself (good examples): https://cgit.freebsd.org/src/tree/share/mk
It's make-like but supposedly improves on some of the arcane aspects of make (I can't judge how successful it is at that as I've never used make in anger).
edit: Yeah, I can't see how I can make it work for a file-based approach.
So while it's replaced one use of Make for me, I can't rightly call it a Make replacement.
if it was opt-in to him monitoring i might be more open. maybe i should fork and change the filename to keep this from happening
That makes make a great entry point for any project. Does anyone have a general suggestion about how to work around this? The goal being, sync a project and not need to install anything to get going. One thing I do sometimes is to use make as the entry point, and it has an init target to install the necessary tools to get going (depends on who I know to be the target audience).
E.g. in one project it will install rbenv, bundle, npm, yarn etc. Call Rake, npx, that odd docker wrapper, etc. Or deploy with git in one project and through capistrano or ansible in another.
As a dev, in daily mode, all you need is 'make lint && make test && make deploy'. All the underlying tools and their intricacies are only needed when it fails or when you want to change stuff.
edit: Yeah, doesn't seem like it.
The one thing that I'm currently struggling to find information on is dynamic dependency handling (dyndep in ninja terms).
Is that something that build2 covers as well? Any resource you could point me to?
Would it be appropriate to open a github issue for discussing this further? I would like to share some example for how my current setup is working and having the github syntax available would be helpful.
https://github.com/google/kati/blob/master/INTERNALS.md
This takes your existing Makefile(s) and produces a ninja build file. Not sure if Android still uses it or it is all soong now.
That's about the last thing I would call it.
GNU Make is hurt by its minimalism. Doing anything interesting in pure GNU Make is a herculean effort. Even something simple like recursing into directories in order to build up a list of source files is extremely hard. Most people don't even try, they would rather keep their source code tree flat than deal with GNU Make. I wanted to support any directory structure so I implemented a simple version of find as a pure GNU Make function!
At the same time, a reduced feature set would be nice. GNU Make ships with a ton of old rules enabled by default. The no-builtin-rules and no-builtin-variables options let you disable this stuff. Makes it a lot easier to understand the print-data-base output.
> The core syntax of a Makefile is just so close to perfect.
It's very simple syntax but it has its pain points. Significant spaces effectively rules out spaces in file names. It also makes it much harder to format custom functions.
Speaking of functions, why can't we call user-defined functions directly? We're forced to use the call function with the custom function's name as parameter. Things quickly get out of hand when you're building new functions on top of existing ones. I actually looked up the GNU Make source code, I remember comments and a discussion about this... It was possible to do it but they didn't want to because then they'd have to think about users when introducing new built-in functions. Oh well...
What's wrong with shelling out to find?
Okay, I just wanted to see if I could do it in pure GNU Make.
true := T
not = $(if $(1),,$(true))
directory? = $(if $(1),$(wildcard $(addsuffix /.,$(1))))
file? = $(and $(wildcard $(1)),$(call not,$(call directory?,$(1))))
glob = $(sort $(wildcard $(or $(1),*)))
glob.directory = $(call glob,$(addsuffix /$(or $(2),*),$(or $(1),.)))
recurse = $(foreach x,$(3),$(if $(call $(2),$(x)),$(x),$(x) $(call recurse,$(1),$(2),$(call $(1),$(x)))))
file_system.traverse = $(call recurse,glob.directory,file?,$(or $(1),.))
find = $(strip $(foreach entry,$(call file_system.traverse,$(1)),$(if $(call $(or $(2),true),$(entry)),$(entry))))
sources := $(call find,src,file?)
Yes.Then I discovered GNU Make supports C extensions. It will even automatically build them due to the way the include keyword works. It might actually be easier to just make a plugin with all the functionality I want...
I would love to see this completed to the point of passing the GNU make testsuite. Having make as a modular library would be wildly useful.
https://metacpan.org/release/CWEST/ppt-0.14
Back then it was a simple way to get some things running under Windows, and I guess it would have also been called "memory safe" albeit not in a way that rust is!
So while most people involved seem to supportive, including Torvalds and GKH, it will take time for any non-driver code to be written in Rust.
[1] - https://lore.kernel.org/lkml/CAK8P3a2VW8T+yYUG1pn1yR-5eU4jJX...
[2] - https://github.com/fishinabarrel/linux-kernel-module-rust/is...
Of course it means that they aren't going to rewrite the core kernel logic in Rust tomorrow, but they wouldn't do it anyway even if GCC supported Rust.
There's already a Rust OS[1] built on top of a brand new Rust kernel, but even if it gets all the traction possible, it won't have proper hardware support before several decades (see the current state of Linux, which is better than ever, but still far from perfect).
What is missing is someone with big pockets to pay very competent engineers to work on such kernels full time and setup a decent test lab with different hardware. Hobbyist might take a long time to finish the job.
Maybe they could take the opportunity to add that. Otherwise I end up needing to
sed -i 's/sed -i/perl -pi -e/g' *There is https://github.com/chmln/sd written in Rust, but it's far from a sed replacement – it's reducing it to search and replace for fixed strings, as far as I can tell.
https://www.oilshell.org/release/latest/doc/eggex.html
It really annoys me to have to remember multiple syntaxes for regexes in shell scripts!
Also, even if you added Perl syntax to sed today, it wouldn't be available everywhere. Oil isn't available everywhere either but there are some advantages to having the syntax in the shell rather than in every tool.
(It's also easy to translate eggex to the BRE syntax or Perl syntax, but nobody has done so yet)
FWIW I think this deserves a submission of its own, so here it is:
Edit since the parent was in fact speaking about the end-user, which I misunderstood: I don't see the problem either. The manufacturer has no obligation to prevent the end user from updating his car's software. There is no locks that prevents the car owner to just disable his airbag[2], or remove the safety belt. It's illegal to do so in most countries, and if the user do do and injure himself or somebody else because of that modification, they are on their own. I don't think it should be any different for software actually.
[1]: https://news.ycombinator.com/item?id=26397176 [2] Edit: in fact, this is a bad example, because you need to be able to disable the airbag to put an infant car seat next to the driver.
The chain of software delivery often looks like this:
Small subcontractor delivers parts of system→ big company provides ready to use solution → hardware vendor uses the solution and gets their devices certified → end user uses the final product
In this case, the hardware vendor is mostly interested in having their devices work as intended. Everyone up the delivery chain has to meet their requirements in some way to basically get paid. That's not a position where it's easy to make demands regarding certifications, since the hardware vendor may just go to someone else.
drm.
ps: airbags not working is less of a software problem than airbags misfiring.
GPLv3 says that manufacturers have to release all the information needed to run modified software on the device, it doesn't mean that there is one (and one only) certified version that can legally run on the device for safety reasons.
GPLv3 in this case would force manufacturers to release the information so that the owner of the car could run modified software, but legally if you do it, you, the user, not the manufacturer, are violating the law.
It's the same thing that happens with electronic blueprints, you can modify the HW, it will void the warranty if you do it.
--------------------------------------------
Protecting Your Right to TinkerTivoization is a dangerous attempt to curtail users' freedom: the right to modify your software will become meaningless if none of your computers let you do it. GPLv3 stops tivoization by requiring the distributor to provide you with whatever information or data is necessary to install modified software on the device. This may be as simple as a set of instructions, or it may include special data such as cryptographic keys or information about how to bypass an integrity check in the hardware. It will depend on how the hardware was designed—but no matter what information you need, you must be able to get it.
This requirement is limited in scope. Distributors are still allowed to use cryptographic keys for any purpose, and they'll only be required to disclose a key if you need it to modify GPLed software on the device they gave you. The GNU Project itself uses GnuPG to prove the integrity of all the software on its FTP site, and measures like that are beneficial to users. GPLv3 does not stop people from using cryptography; we wouldn't want it to. It only stops people from taking away the rights that the license provides you—whether through patent law, technology, or any other means.
But as I've said I'm no law expert and I wouldn't put my hand on fire about it.
In many projects there is the requirement of a fixed release for 3rd party dependencies, versions for which all tests have been checked to pass (this is what is done in NodeJS with packages.json). There is even a requirement of reproducible build sometimes (like with the ongoing project to reach full reproducibility in Debian builds).
Wouldn't these fit the same thinking pattern as the requirements of certification of software for the industry?
I'd love to hear RMS on this subject, maybe he would, too, say that the solution exists inside of GPL3 rather than outside of it.
Sure. Starting price for a car would be several years salary but at least you could abide by GPLv3 licensing should any part manufacturer choose to use it.
then you do not need to make it modifiable under GPLv3. It explicitly states that.
anyways, it's generally not, expect any Linux system in a car to run from flash memory.
I deal with manufacturers and OEMs putting software in cars for a living. I am not speculating, I am just reporting the reasons they tell me they will not accept any GPLv3 software for anything that gets loaded into their target devices.
1. They are lying, and spreading FUD about regulations and licenses as an excuse to hide the real reason why they don't want to let user install their software. (Eg, because options costs extra and if it was free software one could install it for free)
2. Or, they are mistaken and don't understand the license. Maybe the cost to use alternative is less than the cost of figuring out.
3. They are right and that is a sad reality that the government give more power to companies than power to end users for things they owe.
Since you mentioned something about the ROM which was clearly false, that could very well be option 1 or 2.
I mean, why would the gouvernement want to restrict users to update the GPS software or the media player?
Some governments require that automotive manufacturers implement these kinds of certificates to prohibit the installation of custom software in automotive applications for safety reasons, as they do not wish that users could install their own, potentially buggy software, at the potential cost of human lives.
Older coreutils were licensed under GPLv2, which has no such restriction.
If there was a will to do something about it, GPL3 wouldn't be in the way.
Unfortunately, challenging the status quo is difficult when your customers are not the final users, but other companies: if you don't deliver them a free software firmware without GPLv3 components, someone else will, or someone else will deliver something that's completely proprietary.
good.
It means that GPLv3 works as intended.
EDIT: as intended by the software authors, that chose freely GPLv3 as license, they were not forced to.
I'm quite sure they knew what they were doing.
If car manufacturers want to use GPLv3 software, they simply need to respect the license the author released their software under or rewrite the software.
That's false.
It simply prevents the Tivoization.
GPLv3 was created exactly with the purpose of preventing free software from becoming a commodity.
There's a cost involved when you use free software:
- the software must stay free
- if you include software licensed under a FOSS license, you have to adhere to the license terms
simple as that.
If the authors of software X or Y chose the GPLv3 as license I imagine they were completely aware and agreed to the terms of the license they used, including the limitations it enforces.
That is certainly not universally true. It is common, but not universally true.
And we should assume that the authors were NOT aware of real ecosystem implications of licenses - because nobody understands these in detail. I've been trying to understand them since the mid 1990s when I started contributing in a BSD environment, and I won't say that I properly understand them. I understand parts of them, but I don't understand all of them.
As developer of free software, I don't care about those, who don't care about me.
and always is. and people like you are part of the problem, since you want to enable them in their ways.
Why? why cannot we just enable companies to behave mor ethically?
Please read the rest of the sentence, it answers the question.
> people like you are part of the problem
Thank you for turning the licensing question into ad hominem.
That isn't allowed -- you fail verification if you give users a documented way to disable required safety features.
That is quite interesting, but since one can modify one's car to begin with to make it unsafe, it is also rather futile.
Rather, a sensible system would be that after such modifications, a car would have to pass inspection again to be deemed road-worthy. — one may change the software, but one must pay to have it certified again ere it be allowed on public roads, and if that not be an option, one can always simply drive it on private property only.
In France, if you want to modify your car (in a significant way, not specified by the car manufacturer) you need to send your car to a specific administration were engineer will inspect your vehicle before you can get it a plate number. This is called [Passage aux Mines](«https://fr.wikipedia.org/wiki/Passage_aux_Mines»). You typically have to do it when you decide to repurpose a cargo van as a camping van.
That would shock me if you had to do the same after updating the embedded software of your car.
You can disable airbags, ABS and every other electronic circuit in the car and it's all explained in the car manual.
There might be a good reason to do it (for example the airbag is malfunctioning)
It has nothing to do with the terms of GPLv3
Then the car/vehicle is no longer road-worthy/road-safe and has to be repaired before going on public roads again. If the airbag isn't e.g. discovered in a repair shop, during an inspection or any other place where the car can be fixed without rejoining the public roads it has to be safely towed/transported to a place where it CAN be fixed.
And incidentally it's not a problem of "muh car == muh freedom".
If you want to drive a vehicle not safe by the standards everyone has to adhere to you're free to do that on private property.
If you want to drive a vehicle on public roads where probably nobody knows about anything stupid you've done to the vehicle, i.e. a 1-2+ tonne lump of metal, glass and plastic probably with some sharp points/edges moving at a significant speed quite close to other people without any protection whatsoever. It's just horrifying accidents waiting to happen.
Let's take the example of someone fucking around with the airbags in the car. If they're not known good, they might as well go off at any moment during the drive, possibly incapacitating the driver while the car is moving at normal road speeds, making the car veer around wildly and generally being an extreme hazard. There is a reason cars have to certified and adhere to standards of road-safety/road-worthiness.
P.S. When a topic like this comes up I always have to think back to my father when e.g. a commercial/TV-programm about super-/hyper-cars came on. He almost always said "Dafür bräuchte man eigentlich einen Waffenschein." ("One SHOULD need a weapons license for that thing.") in the sense that in our country (Germany) prospective gun owners need a license which IIRC requires amongst other things a psychological examination/certificate to ensure that no irresponsible, no mentally-ill, no mentally-challenged etc. people get the license to own guns. MEANING driving around on public roads with something that's essentially a road-going mix of a missile and a door wedge should probably something like an idiot-test. I mean it's already standard practice to deny RENTALS of cars above a certain power threshold to people under something like 23 or 25.
"You can" doesn't mean it's legal.
But it's documented because you might need to, for example carrying babies on the passengers' seat or if the airbag is malfunctioning and you need to go to the repairing shop it might makes sense to disable it.
> for taking away the operating license of that particular car.
why not the death penalty then? :)
The airbag arguably only saves the driver's life, disabling it has the same effect of smoking cigarettes, except cigarettes are vastly more dangerous.
They don't take away your license for smocking (I guess)
Users will be forced to use outdated, buggy software, or vendors will write minimal closed implementations covering only what they need, without the benefit of the accumulated good work in foss (bugs, lack of features, etc)
EDIT: I'm actually outraged by this. I want to love the rust rewrite of coreutils, but I don't understand why they had to taint this beautiful work with stupid license politics.
How do you know if the platform tools are faster\slower more featureful or less if the first thing you did is replace them?
GNU coreutils won on portability of knowledge, no need to learn how the platform's tools work when I can just roll with GNU. Which is great when you are dealing with multiple different varieties of Unix that are all slightly different. In general the base os tools where like all software that competes in the same space, it did somethings better and somethings worse.
From where I'm standing it looks like you are the one introducing the stupid license politics. The Rust Coreutils developers have chosen a license that they like, and you are the one making arguments about how terrible this is for the world and how they are tainting the GNU Coreutils.
If clang wast named "Rust GNU Compiler Collection rewrite" I would certainly be!
Yes.
Stallman himself was largely responsible for GCC not exposing intermediate representation for analysis.
https://www.reddit.com/r/programming/comments/2rtumb/current...
There is a world in which rustc might have been built on gcc instead - a world where GCC was interested in being a platfom rather than a Maginot line trying to prevent any possible theoretical use by proprietary software even if it means preventing free software from doing the same.
This is wildly offtopic for this thread, but it is not clear that Stallman's stance on this issue was an error. The wide existence of LLVM-based compilers for proprietary architectures is, for many of us, a tragedy that Stallman correctly anticipated. His insightful and courageous steering of the GCC development was a crucial step in avoiding this problem for GCC.
Whether that's still true in terms of the trade-offs being worth it to maintain the policy in the current day, and if it's no longer true what value of N accurately describes when that changed, is something that I think people can reasonably disagree about.
(my extremely boring take being "there's so many counterfactuals here I'm really not sure")
See the email thread I linked for several such examples.
https://lists.gnu.org/archive/html/emacs-devel/2015-01/msg00...
https://lists.gnu.org/archive/html/emacs-devel/2015-01/msg00...
https://lists.gnu.org/archive/html/emacs-devel/2015-01/msg00...
https://lists.gnu.org/archive/html/emacs-devel/2015-02/msg00...
https://lists.gnu.org/archive/html/emacs-devel/2015-02/msg00...
Most compiler research and tooling development now happens in the LLVM ecosystem, precisely because they dragged their feet for so long.
Interestingly enough, if Stallman didn't miss the email[1], LLVM was intended by Latner to be given to the FSF[2]. So yes RMS did commit a management mistake, but it's not the one you think.
> The patch I'm working on is GPL licensed and copyright will be assigned to the FSF under the standard Apple copyright assignment. Initially, I intend to link the LLVM libraries in from the existing LLVM distribution, mainly to simplify my work. This code is licensed under a BSD-like license [8], and LLVM itself will not initially be assigned to the FSF. If people are seriously in favor of LLVM being a long-term part of GCC, I personally believe that the LLVM community would agree to assign the copyright of LLVM itself to the FSF and we can work through these details.
[1] https://lists.gnu.org/archive/html/emacs-devel/2015-02/msg00... [2] https://gcc.gnu.org/legacy-ml/gcc/2005-11/msg00888.html
https://lists.gnu.org/archive/html/emacs-devel/2015-02/msg00...
What's worthwhile is to share and talk about why the license selection may result in proprietary and bloated applications built upon the stack, and the resulting inefficiencies and security issues that could occur years in the future as a result. Perhaps it'd even be possible to speculate about what incentives create and support the conditions for those.
If enough people agree and understand the situation, perhaps we'll see change. If they don't, we'll continue on for a number of years and maybe those hypothetical problems will occur, or maybe they won't and the concerns will have been unjustified.
It's pretty simple reasoning: Rust evangelists want to maximize Rust usage. Period. Replacing GPL-licensed tools with alternatives written in Rust is an easy way to bootstrap this. Compare with: paperclip maximizer (https://www.lesswrong.com/tag/paperclip-maximizer)
I know that the law requires some radio equipment to make it hard for users to modify so that they can't mess with some parts of the spectrum.
But for cars, the responsibility is typically of the user, that is, if you modify your car in a way that is not what the manufacturer certified, you can but you have to do by the rules or face the consequences, the manufacturer is not responsible.
But by locking down software hard, it is a way for manufacturers to make sure that they can't be held responsible and the side effect of locking down features is a nice bonus.
Sure, but they'll also have an easier time of ensuring they're "pretty close" to fully reproducing the original.
At least, in comparison to having to reverse something black box. :)
Rust coreutils replacing GNU coreutils in these programs by implementing the same interface is the same issue.
I’m not sure if it’s possible to spell it out clearer than this.
I guess you're working on credit card terminals?
Anyway: why not have a secure hardware element / TPM that has a GPIO with a pull-up resistor that can be queried by the device? Then, have the bootloader check as part of the boot if the TPM attests with a digital signature that the GPIO is still high, and the application also regularly checking the TPM? Or an e-fuse similar to Samsung's Knox Guard?
That way, a user can remove the pullup resistor (and thus, as it's a hardware modification, has to break the seal of the device) to "unlock" custom firmware loading to fulfill the GPL requirement, but at the same time your application can be reasonably certain at run-time the device hasn't been tampered with.
EDIT: So, I don't think this is actually a technical question. It's a legal question that boils down to the fact that a premise and its inverse can not both be true at the same time.
how do you ensure beyond doubt that the vehicle was not "tampered with" by an "unauthorized unqualified third party"?
would I want to own a car where the oem can turn around burden of proof and can easily claim I have modified the vehicle software, and thus broke vehicle behaviour?
oem: you've patched coreutils, that's what killed your wife, your fault!
me: no. that's technically impossible. the drm disables third party modifications.
now if you provide means to do it anyhow, you'd need to make forensics crystal clear.
so you as a customer want the physical modification to be dead obvious even on a burnt vehicle. especially on a burnt vehicle. so you as a customer can show that you did in fact not unlock software modification.
and state regulations for that reason require drm from car makers, to make it impossible to evade responsibility with flakey claims of third party modifications.
Looks like that device might not be covered by the tiviozation close since it might not be an user product as defined in the GPL
> A “User Product” is either (1) a “consumer product”, which means any tangible personal property which is normally used for personal, family, or household purposes, or (2) anything designed or sold for incorporation into a dwelling.
If you sell to business, that should be alright.
Also, it is fine to make it impossible to modify, if you also don't have this possibility:
> this requirement does not apply if neither you nor any third party retains the ability to install modified object code on the User Product
To start, both of them were write with embedded applications in mind, busybox is feature complete since distros like Alpine and OpenWRT use it (not sure about Toybox, but I know Android ships it).
BTW, both of them cover way more tooling than coreutils too. Busybox for example have a primitive init system and a vi/vim clone. You can also of course disable them during build time.
More seriously. I agree that you can't call rust-coreutils mature yet. But over time maintenance on it will probably be considerably easier and if it ever gets popular enough to be thoroughly battle-tested, it almost certainly will be more reliable and secure. In addition, odds are that even if still immature, it already has fewer serious bugs. Especially memory safety, but also in things like string-handling since C++ -- and especially C -- string handling is notoriously difficult to get right. Then again, maybe not, since coreutils and presumably also rust-coreutils take the stupid approach that strings are just streams of bytes and there's relatively little the Rust std and compiler can do to fix that mess. Oh, wait! Actually it can help a bit by allowing you to be "flexible" in your interface but also parse and validate at the interface, then use better abstractions internally.
Case in point: OpenSSH vs rustls: OpenSSH is a lot older than rustls, has been extensively battle-tested by a very expert community, has been audited, fuzzed, etc to a significant degree, and we're still now starting to see security experts advising using rustls in stead of openssh, or projects deciding to switch to rustls because rustls is likely (and according to a growing body of real-world experience and evidence) less buggy and more secure than openssh. (And as a bonus it's also faster.) Not all of openssh's issues are due to the language, and neither is rustls's choice of language the only reason for its reliability. But in both cases the implementation language does play a crucial role in their security and reliability.
Most different implementations of the coreutils have different goals in mind. GNU's implementation, as is common of GNU is to treat memory for performance, and busybox's attempts to achieve the converse. — so what is the purpose of this one? simply being written in Rust?
How does it compare to the established ones?
https://lists.freebsd.org/pipermail/freebsd-current/2010-Aug...
This is an interesting article that outlines why GNU grep outperforms almost any other grep quite significantly, at the cost of more memory.
In your second link ripgrep is even marketed as having "the raw performance of GNU grep" (though it managed to exceed it).
Also, to be clear, in a strict apples to apples comparison, GNU grep and ripgrep will tend to have comparable performance. ripgrep does edge it out a bit in many cases, but it isn't always earth shattering. However, if you don't limit yourself to apples-to-apples comparisons and instead look at the user experience, then yes, ripgrep will likely be "a lot" faster. Primarily because of automatic parallelism and its "smart" filtering.
> Why?
> Many GNU, Linux and other utilities are useful, and obviously some effort has been spent in the past to port them to Windows. However, those projects are either old and abandoned, are hosted on CVS (which makes it more difficult for new contributors to contribute to them), are written in platform-specific C, or suffer from other issues.
> Rust provides a good, platform-agnostic way of writing systems utilities that are easy to compile anywhere, and this is as good a way as any to try and learn it.
For some value of "anywhere".
Rust: 73M source: https://packages.debian.org/experimental/rust-coreutils
C: 17M https://packages.debian.org/unstable/coreutils
Note that no size optimization that been done and there is a lot ways to do it in Rust https://github.com/johnthagen/min-sized-rust
In term of performance, well, it depends on the binary. For example for cp, most of the time is spending copying the files. For the factor command, some work happened to make it faster https://github.com/uutils/coreutils/commits/master/src/uu/fa... (getting closer).
Using Rust also opens some great capabilities like parallelism. But this makes sense only for complex commands. For example, doing a parallel "df" isn't super interesting (I tried it was too expensive just to start threads to do it).
Anyway, performances should be a focus but only after correctness is implemented.
The Debian version ships with separate binaries though (109 of them), and is much larger. There is probably a lot of duplication as many tools use the same Rust libraries for flag parsing and whatnot which are compiled in the binaries. This is why those 73M compresses down to just 8M.
Also, that way every command executes with the same privileges because it's the same binary. And at least in Linux capabilities are bounded to a binary. So for utility that necessitate of special capabilities (e.g. ping, passwd, sudo, etc) this approach will not work.
Yes, it works on busybox since everything is owned by root on the typical system so nobody cares about privileges.
Finally, you need to load in memory every time you run a single binary all the thing needed by all the other binaries, even for the `true` binary. To me is stupid.
The solution exists, and exists since ever, and it's encapsulating common functionalities in shared libraries, and have very small binaries linked to only the libraries they use. Unfortunately Rust choose not to support them for stupid reasons like problems that were resolved decades ago.
Rust has a stable a.b.i., simply not it's own stable a.b.i. as distinct from C's.
`repr(C)` is what gives a stable a.b.i. in Rust.
Actually I would enjoy that. Sometimes slow filesystems (e.g. stuck nfs) prevent the whole list from being displayed, or displayed only up to some point.
If ctrl-c let me break out the command yet list all the collected data (preserving non-parallel order), it would be nice at times.
Bonus: list skipped mountpoints to stderr..
As it turns out, cp can be made significantly faster by using (very) modern kernel APIs: https://wheybags.com/blog/wcp.html
20 512MiB files
wcp 3.97s, 2579.59 MiB/s
cp 8.44s, 1213.38 MiB/s
rsync 17.26s, 593.33 MiB/sThere's a lot of ways to "copy a file" though. Finding the right buffer sizes, making concurrent read/write ops for different backend devices, and many other things can influence this. A good cp is not trivial.
[0]: https://github.com/uutils/coreutils/blob/master/src/uu/yes/s...
[1]: https://github.com/coreutils/coreutils/blob/master/src/yes.c
[2]: http://github.com//coreutils/coreutils/commit/35217221c211f3...
[3]: https://github.com/coreutils/coreutils/blob/ccbd1d7dc5189f46...
[4]: https://github.com/openbsd/src/blob/master/usr.bin/yes/yes.c
It's pretty naive - a simple linewise read_until loop, a conditional to avoid word splitting and such if it's not needed, and for some reason it collects results into an array and prints when it's done rather than printing as it goes.
It doesn't support --files0-from like GNU wc, so isn't a drop-in replacement from that perspective. It also has the sadly common Rust trope of only supporting filenames that are valid UTF-8.
It doesn't seem overly slow considering its simplicity - usually trading blows with GNU and BSD wc. Perhaps the most glaring omission is the lack of a fast path for -c, which should reduce to a stat() call. Also unfortunate not to use the excellent bytecount crate to provide a very fast -l/m path.
The read_until loop also makes its memory use unpredictable compared with other wc's. If you run it on /dev/zero it will try to eat your computer.
$ wc -c /proc/self/cmdline
25 /proc/self/cmdline
$ stat -c '%s' /proc/self/cmdline
0FreeBSD just uses fstat: https://github.com/freebsd/freebsd-src/blob/e4b8deb222278b2a...
[0] https://github.com/edolstra/nixpkgs/blob/master/pkgs/stdenv/...
Do we already have enough quality Rust workforce to guarantee that?
I wonder if this can be made simpler in order to spur adoption.
There's a less ugly solution - Debian/Ubuntu have the battle-hardened "alternatives" system: https://wiki.debian.org/DebianAlternatives
In this case, diversions would be more suitable, but in this case probably not necessary, since at the moment it’s only for testing.
Is it just my version of Firefox, or does anybody else get a line wrapping justification algo that ends the first line with "Debia" and begins the next line with "n"?
It appears hyperlinks are split opportunistically to match the line justification.
Two things:
* I have never seen a single web page do this
* Something smells soooo right that the only page I've ever seen with this weirdo readability regression involves the word "Debian." It's as if when I read the word "Debian" in a blog/article, I get this sixth sense shiver that something soon will be broken with a default setting because "nobody ever said you couldn't do it the other way."
Edit: it's as if this choice were made specifically to keep me from clicking the FF Reader button. Because now I'm scanning over the entire document counting the number of words in hyperlinks that are broken across lines-- I see 8, which includes another instance of "Debian"-- this time it's broken into "D" and "ebian."
Edit2: OMG the text doesn't wrap to remain in the viewport when I zoom in. I successfully see "Debia" at the end of the first line every time, and if I zoom in far enough I can enter horizontal scrollbar readability hell.
a {
word-break: break-all;
}Otherwise that would be pretty exciting.
The right thing to do is to install your own version of tools you need. The OS vendor tools are only for configuring the OS and downloading your tools.
A project existed[1], but was abandoned since then, so I don't think it will happen anytime soon.