Cross-platform Rust Rewrite of the GNU Coreutils
github.com
github.com
One of the hardest parts of coreutils is keeping it working everywhere, including handling that buggy version of function X in libc Y on distro Z. That's handled for GNU coreutils by gnulib, which currently has nearly 10K files, and so is a significant project in itself.
Some stats:
coreutils files, lines, commits: 1072, 239474, 27924
gnulib files, lines, commits: 9274, 302513, 17476
I believe the relevant part of the GPLv3, regarding "whole package" is:
> A compilation of a covered work with other separate and independent works, which are not by their nature extensions of the covered work, and which are not combined with it such as to form a larger program, in or on a volume of a storage or distribution medium, is called an “aggregate” if the compilation and its resulting copyright are not used to limit the access or legal rights of the compilation's users beyond what the individual works permit. Inclusion of a covered work in an aggregate does not cause this License to apply to the other parts of the aggregate.
Absolutely not. MIT is a FSF-approved GPL-compatible license and you're free to use it with whatever GPL-licensed packages you wish[0].
However, if you package and distribute the Rust coreutils with the GPL test suite, then the users of the package are obliged to either use GPL, MIT, or some other FSF-approved OS license.
But it's simple enough to package the Rust coreutils without the test suite, which puts the Rust coreutils users under no GPL obligations.
GPL2 is all about the distributing of software, not how you use it.
I do agree that any use of a GPL package will make some folks using this package in commercial software nervous, given the very few (none?) actual court cases that have decided these issues.
I don't think this is right. The programs in question are the test suite, and there's can certainly distribute complete programs licensed under the GPL alongside with programs under even non-free licenses (otherwise Mac OS X couldn't ship).
Since the test suite doesn't link against the individual utilities, it won't affect them, license-wise.
It is well understood that the GPL doesn't cross executable boundaries (again, otherwise Mac OS X is in trouble).
I'm free to write a proprietary coreutils test suite and charge $100 for it, so long as it doesn't hook directly into the source of the tools. If it's only interacting with the product, so be it.
Also, comments like this make it seem this is not "fully baked" yet and still needs some dev time:
> fn parse_date(str: &str) -> u64 {
> // This isn't actually compatible with GNU touch, but there doesn't seem to
> // be any simple specification for what format this parameter allows and I'm
> // not about to implement GNU parse_datetime.
> // http://git.savannah.gnu.org/gitweb/?p=gnulib.git;a=blob_plai...
That said, I am FAR more concerned about shipping a solid 1.0 than on a maximum performance one. It'll be a while, but we'll get there.
Still, there seems to be something similar since VS2005. »__restrict is similar to restrict from the C99 spec, but __restrict can be used in C++ or C programs.«: http://msdn.microsoft.com/library/5ft82fed.aspx
> __restrict is similar to restrict from the C99 spec, but __restrict can be used in C++ or C programs.
Rust definitely needs more dev time, but if coreutils already has such an excellent test suite, this sounds like a great way to test Rust in action.
But I was expecting it would take a while before people had the ambition to start doing these things, even when it comes to the smaller ambition of rewriting the GNU coreutils. Rust is a great language already, but it's not a stable language yet.
I wonder if my imagination will become reality eventually. There's really good buzz around Rust now, and it's not even 'production ready'.
And I happen to like having a preprocessor.
Any adoption of Rust in areas where people otherwise use C/C++ is an improvement I think. It doesn't have to be total displacement, and that will never happen.
You just need a successful OS vendor pushing the said language as only means to target their OS.
So yeah, I disagree that any systems programming language can displace C. Other system languages simply don't come with the same amount of improvements that Rust offers.
So, given my premise, how do you target super-cool-everyone-wants-it-os that only offers new-cool-language as their system programming language in the OS SDK?
Are you going to write a C compiler in new-cool-language to be able to keep on using C?
In case case the C runtime has to be implemented on top of the new-cool-language ABI, as it is the only way the vendor supports creating software for their system.
What I am describing is not that different from thirty years ago when C was a UNIX only language, and other OS had other languages as their systems programming language.
I'd suggest that these are fundamentally incompatible statements. "Only offers new-cool-language" is another way of saying "it's a pain in the ass to port to - you have to rewrite your whole app!" if you're coming from anywhere that's ever touched not-quite-as-cool-language.
D is not a viable threat to C/C++ (this is what I said!), as noone is really thinking that D will actively displace the use of C/C++ in many projects in the future. The same goes for all the other languages I mention. In my perception there's a buzz around Rust that indicates it might be viewed differently.
It's not that different to using C++ objects in the stack.
If by no GC you mean to use heap allocated objects with manually managed destructors, then I think the answer is no.
It's true that D's stdlib still uses GC, but he D team has been working on eliminating that.
Ultimately is a security vs. freedom tradeoff, and I think we've sacrificed too much of the latter already.
Freedom should be achieved in different ways. Please don't make freedom and security contradictions.
> Freedom should be achieved in different ways. Please don't make freedom and security contradictions.
In an ideal world, they wouldn't be. Unfortunately that's not how it is in reality.
This guy wants to intentionally leave software exploitable, ready to be unknowingly hacked by criminals or governments, and call it freedom. And he gets upvoted for it!
Also, it seems that "criminals" are the new "terrorists".
Available information indicates that NSA used various coercion tactics to ensure cooperation from various companies and that it has the methods to ensure hardware modifications in its own interests.
So the damage to NSA from improving security too much is less than damage to my ability to use the computing device I have paid for as an actual computing device.
Of course we should also consider smaller threats.
But the better the platform is locked down, the less limitations there ae for advertiser-friendly platform design, which makes people expect a lot of permissions requested by applications, so phishing and trojans become easier.
The problem here is about the effects of scale: it is hard to produce just a few thousands phones with good specs.
unsafe { … }
Someone should invent such a thing.Edit: …yeah, I was being tongue-in-cheek. This is exactly what Rust provides.
http://doc.rust-lang.org/rust.html#unsafety
"...When a programmer has sufficient conviction that a sequence of potentially unsafe operations is actually safe, they can encapsulate that sequence (taken as a whole) within an unsafe block. The compiler will consider uses of such code safe, in the surrounding context…"
> I want my function to be marked as unsafe as well (because it is)
This misunderstands the nature of `unsafe`. `unsafe` means "I promise that this is actually safe, even thought it can't be inferred." Take, for example, a type which has mutable, shared state, but enforces all access through a mutex[1]. This type presents a safe interface, but the compiler can't know that it's safe, so you have to use `unsafe` internally.
1: http://static.rust-lang.org/doc/master/std/cell/struct.RefCe...
Pretty much every function is going to be transitively running unsafe code, since the core libraries use it to implement the base primitives.
unsafe fn foo() { ... }
Any function marked as such is then allowed to call other unsafe functions: unsafe fn bar() { foo(); }
But there is a way to break the chain, which is to use an `unsafe` block without marking your function as unsafe: fn qux() {
unsafe {
bar();
}
}
That said, it's incorrect to think that Rust is any more unsafe than any other language because of this; most languages simply defer this behavior to their FFI. By pulling it into the language itself, Rust is actually safer than e.g. calling C from Python, because Rust can do the low-level fiddling while still retaining at least some of the safety checks of normal Rust code. Even unsafe Rust is safer than C.This is an important point. `unsafe` blocks only let you do a few extra operations[1], not anything you want. A lot of safety checks still happen inside of unsafe blocks.
1: http://static.rust-lang.org/doc/master/rust.html#behavior-co...
[1]: http://doc.rust-lang.org/master/rust.html#behavior-considere...
Consider, in a "type safe" law system, intent doesn't matter. Only what you did. Now, look at the laws that run on that principle.
Now, I personally think this is probably orthogonal to type safe programs. But I'm not completely convinced, yet. Type safe programming seems hamstrung by the fact that a type system only really protects that which is specified in the type system. Which seems to mean you can't easily have a variety of ways to do things, as ultimately the type system has to grow to encompass all of the system.
Now, I grant I'm probably just soured by some bad systems in the past where a change to one part required a rework of the entire system.
Also, IME, type safe languages don't hobble me as a programmer. However, the systems and things I work on (and am trying to get work on) have generally well considered specs associated with them. Be it radio protocols, file formats, or whatever. A type safe language (haskell, sml, ocaml, ada, rust (it seems, not explored it much yet), etc.) can really help with these projects. Instead of needing a dictionary translating what magic # 134 means if it's used in this context (god, these old fortran programs break me), we can actually specify things in code in sane, clear ways that map from spec/design to code cleanly. OO, in theory, can help us with this, but it also adds a lot of overhead to create a ton of classes when really we want integers that the compiler knows some are meant to be temperature and some are meant to be pressure. If you start adding temperature and pressure, it flags it. If there's some formula where that actually makes sense, you deliberately, intentionally, explicitly handle the conversion. This is a good thing. Slightly more verbose code, but such a boon for mainenance and v&v work.
So basically it's a tradeoff. If people working on the Linux kernel were very careful and good hackers, writing in C would be much less of a risk than letting fresh-out-of-college programmers tinker with memory on the low-level. But even very careful and good hackers can screw up sometimes—see what happened with OpenSSL.
I don't see a reason not to use Rust except for performance, or when the unsafe features of C are required to get the job done. I can't say a lot here because I've only seen Rust's syntax and feature list and never programmed in it. (On the other side, I've written a fair amount of (functioning) code in C, and I feel that debugging race conditions coupled with memory management bugs has shaved a few years off of my expected lifespan already.)
Idiomatic Rust will drop a little bit, to the level of idiomatic C++, but there's nothing stopping you from optimizing that to C levels in exactly the same way when necessary.
On the other hand, a lot of bounds checking is probably going to happen where you'd manually enter it into C/C++ code anyways (if you wanted to avoid certain errors).
When we build the next generation of free operating system, I want it to be as secure as humanly possible. Your device isn't free if the NSA or some criminal has owned it.
> Your device isn't free if the NSA or some criminal has owned it.
It's not free if you can't own it either - and that's where the problem lies: how do we stop the secure systems we create from being used and secured against us.
> That's the wrong place to attack the problem.
Then what do you think is the - currently practical - way to attack the problem?
I think there is a bit of misunderstanding I created in my original post; I'm not completely against security and hate buggy software as much as anyone, but only pointing out that we rely on insecurity for many good things and freedom we have today, and thus we should give more consideration to the implications of using more secure languages before they get to the point of becoming so popular that there is no turning back. They're just such a good fit to be used in trusted computing/DRM systems, and that's what scares me the most.
Ultimately, the only way to correct this is the active encouragement of DRM-free content. The active development of FOSS software. The active development of open-specced hardware. Let's take these tools that let us develop things with fewer errors (or where errors are made blindingly obvious) and make this free world we want.
Also, memory safety is hard to enforce, I reckon, while controlling the memory hardware, which is what an operating system like Linux* has to do. Rust can be used, but still, there are many unsafe things that OS has to do, like moving and assigning pages, managing which programs use which pages, swapping, etc.; which, I reckon, requires knowledge and usage of physical memory addresses and access to individual bytes.
Rust may be an appropriate and advantageous language to use for the non-core parts of a microkernel, though.
* I acknowledge that Linux is not an operating system per se.
C has no memory safety by default. So if you want to audit your C program for memory safety problems, those problems can be anywhere. Even if that C program/component doesn't actually need such low level access in the first place!
[1]: http://static.rust-lang.org/doc/master/rust.html#unsafety
[2]: Here's a list of (hobby) kernels written in Rust: https://github.com/mozilla/rust/wiki/Operating-system-develo...
With your comments I have to concur. I have also read a comment on name-mangling somewhere on this thread, probably one that you've written. If C can be replaced with Rust for most-if-not-all applications in future, it is a step forwards IMO.
It is possible to expose Rust data and functions to Python, Ruby, etc, since Rust is ABI-compatible with C. Unfortunately, you lose most of the Rust standard library since the Rust runtime is not freestanding. And I'm not sure how -- or if -- Rust will fix this issue. It's certainly desired, but it will be a lot of work.
The only bright side to all of this is that Rust is such a modern and pleasant language to use that scripting language bindings are almost not needed. Almost. But it's still a enough of a low-level language that bindings are highly desirable for quick prototypes, exploratory programming, and other scripting tasks.
> I'm not sure how -- or if -- Rust will fix this issue.
We're addressing this as we speak! We've split out the stdlib into many constituent libraries that depend on each other, and then just made the stdlib a convenient facade for these libraries. After opting out of the stdlib, you can then opt into any of the "core" libraries as you like, a la carte. Many of these smaller libs don't rely on the runtime (they don't depend on librustrt) nor do they rely on allocation at all (liballoc), and can thus be used in a freestanding environment. See the other comment in this thread for more details:Not really, at least not in the sense that you mean here: The generation of machine code in shared objects. There is a distinct lack of definition between C syntax and the resulting machine code.
> The best advantage of writing libraries in C is that it allows the library to be usable via FFIs in many other languages.
You can create C shared objects from Rust (or from a ton of other languages which generate native code). No big deal. It can do what you want.
And more importantly: there is a distinct lack of similarity in behavior in this respect between vendors, versions, architectures, language features, ambient temperatures, and times of day. Change any of those things and you can get different output from your C compiler.
Rust has allowed this for over a year at this point. For example, see the (old) blog post "Embedding Rust in Ruby": http://brson.github.io/2013/03/10/embedding-rust-in-ruby/
EDIT: Thanks for the informative responses! This sort of effort might ought to receive the occasional audit, which probably could be done in semi-automated fashion.
[1] in a copyright infringing way, as defined by the courts in your preferred jurisdiction
I did not mean to hurt any feeling. I was just surprised to see the license of a software where "GNU" is in the title. I naturally assumed that it will be in GNU license. So it was a surprise to discover MIT.
Now I understand. I should have written "surprising" instead of "funny". Sometimes writing about a subject gives you the answer.
Sorry for those that I hurt.
GPL has neither of those goals and actively works against the latter.
For your definition of 'anyone.' BSD software often does not let the end user use it as they see fit, as they no longer have access to the source code.
Someone else may add some modifications to my code and I may not be able to see those modifications, true. But nothing's been removed.
If I'm going to release code under a license, I need to understand every line in that license. GPLv2 is pretty understandable, GPLv3 is ponderous in places and ambiguous in others (needs lawyers and lawsuits to understand), and AGPL is simply around the bend. Loopholes you could drive a truck through.
I don't want to use GPLv3 or AGPL until their ambiguous bits have been clarified, but their adoption has been pretty low so that may never happen. Whether that's a good thing or a bad thing depends on your point of view.
Also, the GPLv2 vs GPLv3 fragmentation has turned off a lot of people.
I want to spend my time coding, not worrying about license issues and lawsuits.
However, you can easily define a function with a C ABI (and a non-mangled symbol name):
#[no_mangle]
pub extern "C" fn foo() { ... }
I'm not sure if you regard this as better than C++'s situation.Considering how quickly C++ compilers are iterating, it might even be available in GCC/Clang before Rust 1.0.
cinfo() { xdg-open "http://www.gnu.org/software/coreutils/manual/html_node/$1-invocation.html#$1-invocation"; }
Which you could use like: cinfo ddThey are rendered in less, which supports vi-like searching.
Info, on the other hand, is one of those odd pre-www hypertext systems. Advanced for its time, it is now clumsy and nonstandard.
Its not that I want to discourage anybody from using their preferred documentation resource. If a developer would rather spend their time writing info pages, that's okay. But at least dump it into a man page. Even an oversized man paged like the one for mplayer is preferable to a pithy page telling me to use info. If I want something with hyperlinks I'll check the web.
I don't get how on the earth can using CVS for version control be an appropriate reason to consider a software project bad. Yes, CVS is old and centralised, but, is it that big of a deal that it's usage by a project per se projects the project old and inactive?
However, I don't think you should ever avoid writing something you'd enjoy for the sole reason that there are prior attempts (successful or unsuccessful). If you aren't motivated by money, you won't lose anything if your project doesn't get a single user, but you will always gain enlightenment and great experience.
My point is that the argument for CVS is probably just a forced, faulty rationalization by the author so he/she can justify writing a new implementation. But I'm repeating again that there is nothing wrong in reimplementing something just for the sake of it—you have nothing to lose and a lot to gain :)
CVS still works just as well as it ever did, but it is super crusty at this point. If an open source project hasn't transitioned off of it, that's a sign to me that the maintainer either doesn't know any modern DVCSes, or doesn't care about the project enough to transition to them.
It's not a guarantee that the project is dead or outdated, but it's a smell.
Of course, I consider still being hosted on Sourceforge to also be a smell.
Past re-implementations have focused either on the learning experience, code size (for embedding or Unix "purity" reasons) or the license (i.e. not GPL).
There is something to be said for the nice clean implementations of some of these tools. There is a little to be said for the language itself.
Really, something bigger and better is needed. A new HTTP server, a new sendmail or something, a new DNS server. I'd love to see Go or Rust take on Bind and produce a safe, secure, high performance implementation. The Ada guys never stepped up and produced anything interesting to show their tools' superiority, there is a lot more interest and community in Rust and Go. What's the leakiest, buggiest part of the equation right now? Go make a better one.
[1] http://onlinelibrary.wiley.com/doi/10.1002/spe.4380111102/ab...
Why?
Many GNU, Linux and other utils are pretty awesome, and obviously some effort has been spent in the past to port them to Windows. However, those projects are either old, abandoned, hosted on CVS, written in platform-specific C, etc.
Rust provides a good, platform-agnostic way of writing systems utils that are easy to compile anywhere, and this is as good a way as any to try and learn it.
I've nothing against rewriting coreutils in Rust as part of a project to produce a "Rust operating system" in the same way that UNIX is the "C operating system". I do find "project rejected due to not on github" to be silly and a slightly unpleasant trend.
Maybe cygwin is OK if you are fine with the costs and want to maximize the amount of Unix-like functionality (and you don't hit the bugs - I have seen some very nasty ones in the cygwin dll - deadlocks, random crashes). But something like MinGW/msys does a much better job of compiling for Windows.
"Cygwin does a very good job of trashing your HD on windows." It cost me 90 euros for two hard drives.
*by replace I mean occupy some of their current marketshare.
(I realize your post may have been tongue in cheek, but it can be hard to tell on the interwebs.)
If you watch this excellent panel discussion from this year, you can hear Rob Pike express how they later regretted how their poor choice of terminology created such confusion:
http://channel9.msdn.com/Events/Lang-NEXT/Lang-NEXT-2014/Pan...
(approx 6 minutes 45 seconds into the talk)
On the other hand, it's also in version 0.10, keeps changing all the time, and the stdlib doesn't contain nearly everything that Go's does, by design.
The "hello world" from the tutorial compiles to a 1MB executable (dynamically linked with the C library but statically linked with the rust library). The equivalent C program creates a 6.5KB binary with gcc.
High-level description the runtimes: http://doc.rust-lang.org/master/guide-runtime.html
Avoiding the standard library/runtime: http://doc.rust-lang.org/master/guide-unsafe.html#avoiding-t... (don't miss the "Using libcore" subsection)
If we're going to ditch C for something proper, lets do it properly.
Go just seems like another half-assed minimum effort Google project.
I gave you an upvote, because I gave you benefit of a doubt and assumed that was an honest question.
Rust is sort of like a cross between C++ and Haskell. It has better facilities for abstraction and enforced safety than Go, but at the cost of more cognitive overhead and being harder to learn. People who plan to use it generally use C++ now.
These languages are actually going after different use cases, and for my robot there are some parts of the codebase I'd write in Go if we were re-writing, and some in Rust.
For example, I really dug the decision to remove GC from the Rust language and move it to the standard library.
It does mean that our GC will never be particularly optimal. And that's fine, because if you really need shared ownership you should be using our really great refcounted pointers instead. :)
Now that the refcounting is also moved to a library, I am not sure.
Or are those types known to the compiler and the respective optimizations applied?
fn just_a_ref(x: &T) { ... }
fn rc_by_val(x: Rc<T>) { ... }
fn rc_by_ref(x: &Rc<T>) { ... }
let some_rc_pointer: Rc<T> = ...;
just_a_ref(&*some_rc_pointer); // no ref counting
rc_by_ref(&some_rc_pointer); // no ref counting
rc_by_val(some_rc_pointer.clone()); // ref count incremented
// last use of a value (statically guaranteed that
// some_rc_pointer is never used again):
rc_by_val(some_rc_pointer); // no ref countingRust is far more extensible than Go - for example all of the concurrency primitives are built as libraries, and users can create their own depending on their unique use cases. Go's on the other hand are built in, and are difficult to extend.
Rust aims for generic, zero cost abstractions and control over allocation without compromising on safety, but this means the type system is more complex than Go's. Go's static type system is quite simple, and easy to learn, but is hard to write generic abstractions over unless you want to resort to using `interface {}`. Go also uses a mandatory GC which makes low level programming and interfacing with C difficult, but abstracts away from up-front memory management.
Rust's compiler is slow, but generates very fast code because it performs a great deal of static optimizations via LLVM. Go on the other hand builds blazingly fast but performs little to no optimization at compile time, nor compensates for that with a JIT.
If you want raw performance, control over allocation, and an expressive type system, choose Rust. If you want simplicity without having to think about allocation, choose Go.
I wrote one of the first utilities for this when it was first opened up for collaboration, so I hope it succeeds :) I need to go back and write tests for my util.
Any chance of Rust taking over embedded systems programming? That's still mostly done in C and quickly devolves into horror.
I don't know about 'taking over' but there are some people who are using Rust for this use case.
Freescale's Freedom boards are so cheap, so capable, and have such miserable tooling (mbed, DIY usb, Processor Expert, CodeWarrior, ugh)... I wonder if Rust could make them attractive.