Cross-platform Rust rewrite of the GNU coreutils
github.com
github.com
> Rewriting SQLite in Rust, or some other trendy “safe” language, would not help. In fact it might hurt.
(see link for expansion on that matter, which is a question of tooling and testing)
Yes, all programming languages allow the programmer to write bugs. But languages very much vary in how many, and what kinds of bugs programmers write in practice. Saying "well, Rust doesn't eliminate all bugs" is attacking a straw man. If you want to argue that Rust isn't worth it, you need to convince me that C plus gcov results in fewer bugs in the important areas in practice than Rust (plus kcov [1] if you like) does. I think that's going to be pretty hard. (Especially if memory safety issues are the most important bug class you're concerned about: I think it's completely impossible for any C-based solution to compete with Rust here, regardless of how much tooling you add.)
Drawing an equivalence between undefined behavior and compiler bugs also doesn't make sense. Compiler bugs are way way less commonly encountered than undefined behavior in C. Also, they're qualitatively different: compiler bugs get fixed in new compiler versions, while UB is by design and doesn't get fixed.
[1]: https://users.rust-lang.org/t/tutorial-how-to-collect-test-c...
I don't have a dog in this fight, but I don't see how the burden of proof is on Hipp rather than the folks proposing the change. In other words, shouldn't the "rewrite it in Rust" folks have to prove that the cost of their proposed rewrite will be justified?
That a programmer who has produced such high-quality and rigorously tested software as sqlite should be portrayed as either cavalier or naive about software quality is something I find profoundly mis-guided.
Yes, it's very much akin to saying "well, a seat belt won't eliminate all deaths". I would understand an argument of "I'm not prepared to rewrite the software at this time" or "I'm not familiar enough with replacement X to assess whether it's a good choice", but at some point more people will have to acknowledge C's failings with more than lip-service. That doesn't mean Rust has to be the solution/replacement, but something does.
Is it theoretically possible to rewrite Sqlite in Rust and still maintain compatibility with programs written in C?
The promise of rust, which may or may not be realized is pushing some very common problems down to the compiler. All code has bugs, so the compiler probably does things wrong in some cases. As the tools mature, these will get scrubbed out, just like every other project.
That said, array bounds checking is responsible for so many problems, it seems worth it to raise the minimum for a language. Not every developer is elite. In fact, they're pretty rare. C requires you to be really smart all the time, or at least be aware of when you're not smart enough to get a chunk of code right. Looping over some bytes from a file shouldn't be that risky. rust lets me, a less than elite developer, save my few moments of brilliance for the hard part of a program, rather than having to worry about the evaluation order of foo(i++,i++);
Maybe, it'll turn out the only way to make good software in the future is to find the best 100 developers in the world, and get them to make stuff. I doubt it though, being able to leverage the other million of us to make stuff, and have some (a lot?) of confidence it'll be free from the most common C errors is valuable.
Nobody should be forced to use tools they don't like. rust is trendy, but it has some very good ideas. Trendyness isn't reason alone to dismiss its approach. blah blah rust cheerleading blah blah.
(The one exception may be, like, DJB. But djbdns/qmail are very unusual C programs in many ways.)
It's not so much that it requires smarts. It's that is requires you to be ever-vigilant and to never make any mistakes. (That's why UB and bounds overflows are so devastating to software security. Almost any slip-up by the developers can be exploited.)
Incidentally, the ever-vigilant bit is also why we really want compilers to be doing the bounds-checking (or proving that it isn't necessary). Compilers are really good at being ever-vigilant. Humans (no matter how smart)... not so much.
Amen to that! If I could just add one feature to C - at least as an option - it would be array bounds checks.
I have no trouble with manual memory management - a garbage collector is nice to have, but I have not had many problems with memory leaks or dangling pointers. And the ones I had were relatively easy to locate and fix.
But array bounds violations are so easy to commit and so nasty to track down... When I still wrote C code for a living, I would have gladly sacrificed quite a bit of performance to get bounds checks on all array accesses, at least for testing and debugging...
At some point some popular compiler is going to make a subtle but important change to some undefined behavior that's not going to be immediately obvious as to it's repercussions, and the fallout will be massive. It boggles my mind the mental contortions people will go though to justify what is essentially an argument of "it hasn't caused a problem yet" while ignoring that it's caused many problems already, just not that they've noticed or that have affected them.
Wait until you realize that the length of a byte in C is not clearly defined. Someday a processor will come along where a byte is 6 bits, and the fallout will be massive. (Really, it's happened before).
No, actually that processor would not become popular because on one would use it.
The reality is unless you are using a formally-defined language like ML, you are relying on undefined behavior in your language.
I don't think the argument is that these cases should be ignored -- they are now corrected, after all. It is that most of these cases should be treated as low priority compared to issues that are creating observable problems.
This discussion seems to miss a couple things. SQLite is embedded quite often, by programs written in C; would embedding a rust library and possibly runtime fix things? The parent program would still potentially have defects and those defects could impact SQLite.
SQLite is also rather mature, I do t see how you can compare a rewrite to nature code. Take OpenSSL as an example, they are t tossing it, they are fixing it, it's a much more shallow lift to fix it.
I'm all for some big rust programs to prove its case though, a mailer, a dns daemon, some sort of database. Something useful cut from whole cloth, and ideally something we have historically not done well. I don't think a rewrite is it though.
Testing "every single instruction" does not guarantee that all C-level UB has been eliminated, because some bugs can be input-dependent. For example, even if your coverage tells you that this function has been tested, it could still trigger undefined behavior for other inputs that trigger overflow:
int f(int x) { return x * 2; }If the inputs here are coming from something external to the program, such that the compiler can't know the value of x then the machine code should be faithful to the language statements, it's a bug if overflow occurs which may have other runtime implications, but the compiler isn't going to remove some chunk of code, etc. because of it.
From their testing page SQLite say they also test boundary conditions (I don't know to what granularity that is though) which which may catch something like this, which is agreeing that instruction level coverage isn't enough. They also run their tests with all the various sanitisers enabled.
Isn't this the same case in Rust, the non-release builds would need to see a suitable test case for the overflow check to cause a panic.
I guess I'm asking is this example relevant to concerns about undefined behaviour and optimising compilers vs. Rust? Since the function is input value dependent in practice isn't this more like implementation defined behaviour at runtime, in that a platform will alway provide a consistent behaviour, e.g. overflow, trap, saturate, etc.?
I've followed regehr's blog for a few years, I've read a lot about Rust, but I'm mostly working in higher level dynamic languages and don't have a lot of hands on experience with C or Rust and just wondering if I'm missing some subtlety here.
For all the complexity of the SQL Language, and efficiency constraints SQLite needs to have, plus all the algorithms it has to implement, it still got a pretty well-defined task. It is not easy to transfer that experience to other systems. Say, your company's backend API. Completely different constraints.
It could still be the case that, had Rust existed when SQLite was being created, that it would have taken much less engineering effort. Which is the metric that matters, as given infinite manpower, you can write anything in any language.
I do agree that rewriting SQLite now, which is a very battle-tested piece of software, would probably do more harm than good. I bet people will still try, for fun if nothing else.
Realistically, you rewrite the components piece by piece in the new language. You make sure each compiles right with testing and review. You report any problems to the compiler team, who fixes them. Eventually, you're whole app is in the safer language without a lot of work. You might even swap them where safer one becomes reference code with other one there for any platforms not supported yet or too buggy. A diversity benefit as Hipp mentions.
With all of this, the program becomes immune to most memory & concurrency issues while being easier to maintain. Its undefined behavior will probably be a fraction of C's in number and severity. That's a net win.
Note: I'd have told him Ada/SPARK instead of Rust given it's been stomping C in embedded safety for a long time w/ lots of tooling for verification activities already there. Counters his compiler maturity argument, too.
That is not at all what Hipp's comment says, and the original author's response to it gives it a much more charitable reading than you seem to be.
I agree in principal, however the timing is wrong.
Rust is a very new language (less than 6 years old), and is undeniably totally unproven. It is the latest "buzz" language, and may not be around long term... nobody knows.
What-more, a full-fledged re-write of CoreUtils in Rust is unlikely to be used by any production environment, because this new CoreUtils will also be undeniably totally unproven. This is the same growing pains LibreSSL have been experiencing, lots of gung-ho fans, but very few actual users (outside of OpenBSD/FreeBSD) - and their undertaking is arguably a lot easier, since they're just cleaning a codebase, not starting from scratch.
Average users who just consume distro's are not going to switch to a new unproven CoreUtils (even if they knew how), and distro maintainers are not going to switch until it's proven either. It will take a huge company with a huge install-base switching and testing it in production for many years before others start to feel comfortable... however this is also an enormous burden on said mega-corporation, for little-to-zero perceived benefits.
Yes, in principal, it's "safer" code, but to a mega-corp with thousands of installs, the risk is too great. New bugs, language pitfalls, behavior potentially changing, etc. Perhaps Rust dies, perhaps it's replaced with an even better alternative. It will take a LOT of time to work all this out.
Flatly, re-writing a several decade's old, matured codebase in today's flavor-of-the-week language is not a good idea. It's a waste of time and effort.
Let the languages mature more, do more systems work that doesn't involve replacing the foundation we all stand on... and maybe, in 5-10 additional years, we'll see where Rust goes.
It's odd that so many in the computing industry are unwilling to move on from a language from 1978. We think of ourselves as one of the most fast-moving industries, but we have this odd reverence for early C and Unix that makes us stubborn and resistant to change. The fact is: we didn't know how to do some things properly in 1978. We know more about programming language design now. (Even Rob Pike would presumably agree—that's why he created Go!)
To be sure, we shouldn't just rewrite things in new languages for no reason. I think many segments of our industry are too fad driven. But to me the right thing is simple: Let's evaluate new technologies on their merits. Rust may well be worse than C! But if it is, let's figure out why, and say so explicitly.
Actually we did know how to make it properly, as Extended Algol in Burroughs B5000 in 1961 was being used, just to cite one example from many others that were ignored by the UNIX authors, because they didn't want to spend too much effort designing a proper compiler.
We're fast-moving because our foundation is solid and not changing (ie. CoreUtils and gang). It's an assumption that these things "just work" with zero fuss and weirdness between systems.
We build on-top of these systems, so changing them out from underneath us all is a dramatic shift.
Perhaps Rust is the key to making these things better. I never claimed Rust is bad. I've only claimed that Rust may or may not be the right choice here, and since it's so young and unproven, we should wait before trying to re-write "all the things" in Rust. Today, Rust is a pet language... tomorrow, maybe not.
Remember, it took C many years to "catch on", and even longer to become the de facto standard for systems work. We can't rush this sort of thing... especially given the sheer magnitude of things depending on this code.
We should also be careful who actually does the re-write when the time comes. New CS grads who cannot understand the old C code and therefore feel [insert-new-hip-language-here] is better... are not the best ones to tackle this sort of endeavor. This requires deep, deep understanding of the entire package, how all the components interact, legacy behavior and the reasons behind design decisions, etc...
Seems like this is a good opportunity for Mozilla to demonstrate the resilience of Rust and its ecosystem by adopting Rust as the language of choice, where possible, and deploying in-house a custom Linux system in which Rust-written components are plugged-in, as and when they are written and ready, with perhaps a full-fledged switch to Redox sometime in the future. If Rust and software written in it are pushed to their limits within Mozilla, then its not an unreasonable recommendation for it to be deployed on an even wider scale.
I don't see Mozilla doing this, really. There's no direct benefit, and there's a nebulous future benefit for Rust.
Mozilla is using Rust components in Firefox though (as well as Servo being mostly Rust). "Adopting Rust as the language of choice" seems to be happening already -- I've heard a lot of folks enthusiastic about (re)writing in Rust. This stuff takes time, though.
That's the sort of thinking we'd benefit from having less. Legacy is a terrible burden.
But again, performance is complicated. People have put tens of thousands of hours into c/c++ optimization. rust is young, so not so much time there. On the upside, rust has room to grow.
https://benchmarksgame.alioth.debian.org/u64q/which-programs...
Please don't generalize that into language X is beating language Y.
Honestly using higher-level languages is the same mentality as taking a pill to magically lose weight. It's quick, but detrimental (to programmers ability) in long term.
Programming becomes easy but in the long term most people forget how algorithms and data structures work, in addition to cache mechanisms and other optimizations.
Until hardware changes drastically there's no sense rewriting everything (unless of course it's just for fun). It's better use of time to study Math and lower-level concepts instead.
That being said, most businesses will take the quick pill instead.
Do you know how the lifetime/borrow check system works?
I've heard plenty of criticisms of it, but I've never heard "the lifetime system doesn't make you understand manual memory management".
That's not a valid argument for why better tooling can't help alleviate some of the difficulty.
> Honestly using higher-level languages is the same mentality as taking a pill to magically lose weight. It's quick, but detrimental (to programmers ability) in long term.
If that were true, the most effective programmers would only code on assembly.
> Programming becomes easy but in the long term most people forget how algorithms and data structures work, in addition to cache mechanisms and other optimizations.
Why would higher level languages obviate knowledge of any of these things? If anything, I think it would help by reducing "noise" from incidental complexity; e.g. ownership in C vs Rust.
No, because C is as fast as hand-coded assembly in most of the cases. The same can't be said of any high level language in comparison to C (save for C++ and Fortran).
The mental capacity saved can be invested in higher level design issues that gone get you a lot more in the long run.
However, lower-level understanding is paramount. For instance, and most-importantly today, taking advantage of multi-threading requires understanding of cache-coherence, memory-alignments, et cetera.
Therefore, while a proper serial algorithm today may be correct its scalability is going to be limited without lower-level understanding. Although, I'll admit that can be built into the higher-level languages (like concurrency in Clojure for instance), but I personally prefer to understand what is going on rather than blissful ignorance :)
The evidence is pretty clear that programming in C does not confer enough skill to prevent disastrous mistakes despite its near-hardware level of abstraction, so I'd really like to know what you're saying here.
Although I cannot refute there are additional memory-related issues to be aware of in C, it is precisely such awareness that make someone a better programmer because its how the underlying hardware works. For instance, parallelizing algorithms must take into account cache-alignment boundaries and so on.
Abstractions are good. However, too much abstraction is bad. Just as too much of anything is a bad thing. When people rely solely on such abstractions, it is in fact ruinous to ability.
That being said, higher-level languages certainly have their place. But I firmly believe that C is a high-enough abstraction for systems programming and with proper idioms and testing strategies, it can be just as safe as the plethora of garbage-collected languages out there in the wild.
I cannot speak to SQlite in this regard.
EDIT: Thanks to replies for clarification that it's just one allocator, aborts, and others are available. Still feel weird about it but that's better.
Clarifying note: this is configurable.
None the less, this is what I typically end up doing when writing Rust code as much of th Rust standard library is simply not ready for "serious" usage.
In my experience, the Rust standard library is extremely convenient and it handles corner cases very well. There are definitely still holes, but that's what "cargo add $CRATE_NAME" is for.
My biggest annoyance with Rust (and it's not a huge one) is that if I'm doing something off the beaten path, I'm probably going to need to wrap a C library or two that nobody has wrapped yet. There's a lot of great stuff on crates.io, but it's only a minuscule fraction of the total C ecosystem.
Can you elaborate on this? I've been getting serious usage out of the Rust standard library for years...
Really, this is true of any programming language and any involved enough program. It's just that with C and security-critical programs, there's some unfortunate concordance between the errors you want to avoid and the errors that are harder to avoid.
I have seen such a paragraph in another projects README, IIRC a Go rewrite of standard utilities. I do not understand why a project would be obsolete because it's on CVS or is old. CVS is simpler than Git, albeit less capable. I, for one prefer it over Git for this reason, and others may do so too. Why would the end user care?
And why would we care about the age of a programme if it works?
Now, that said, the authors need not justify anything, they are free to do whatever they want, and I guess it's fun to code this stuff. I tried this just to play with Golang when it was 1.1.
And lastly, that Makefile is really a bunch of shell scripts, and some common environment variables. There is no real dependency tracking in it, and it is easier to maintain a bunch of shell scripts than a seriously ugly and complex Makefile like this. N.b. that when I say dependency tracking I mean dependencies among input and output files of processing commands, not tasks. I guess Cargo would know how to do that, and how to not build if the build artefact is already there. It's a useless use of make. And that makefile is very GNU-specific, not complying with a project whose purpose is to be cross-platform. Also, I guess, tho I'm really unfamiliar with Rust and Cargo, if that makefile was removed, maybe on Windows they'd be able to drop development dependencies on Cygwin or Msys.
If a person can't build the thing, they effectively have no power to contribute to changing it. If they have no power to contribute to changing it, it's... spiritually, missing the most essential parts of open source.
This is the fundamental issue driving comments like the quote above.
You can tell me that not having a dependency management system, not having a sane version control system, using "some" C compiler (sans a matrix of tests specifying a range of expected good compilers), arbitrarily fine-grained platform specificness meaning an average OSS contributor can never reasonably test their changes against all targets..... all of these things can be "worked around". But at some point, the litany of issues -- some of which take a new contributor dozens of hours to work around -- becomes a simply overwhelming barrier to contribution.
It's time to admit that a foundation of workarounds in FOSS development processes at the very core of our systems is a problem.
Re-writing it in language $x may not be the solution, but it's certainly understandable that there's a widespread desire for simply getting better toolchains underneath our most basic essential systems.
I've said nothing about dependency management as in fetching code that the project depends on. I'm talking about compilation dependencies, i.e. file a.o depends on a.c, a.h, b.c and b.h. Make is for this:
a.o: a.c a.h b.c b.h
cc -o ${.TARGET} ${.ALLSRC}
But nowhere in the projects Makefile the rules are in this fashion. They are like shell aliases. What I wanted to say is that the Makefile could be replaced with a bunch of shell scripts that would be easier to use and maintain.For the rest, it seems that we mostly agree, tho I do not think that CVS is not sane, it's perfectly usable.
https://github.com/linuxfoundation/cii-best-practices-badge#...
a) SVN seems daunting and complex, tho I didn't ever dive into it. CVS is so simple and easy, a half-arsed programmer like me can actually understand it. Things like git and mercurial are way more complex.
b) RCS is real handy for single files, e.g. a free-standing text file or shell script. But when the thing grows up, it is very easy to integrate the fileset into a CVS repo preserving it's history: move the ,v files to $CVSROOT/$MODULE/.
c) The repository model of CVS is as transparent as it gets.
d) The keywords like $Id$ are really useful.
e.g. I keep my system configuration in "~/Checkouts/system-config", and I have a script that cp's the files to appropriate locations using a map file. When I'm not sure if the active config is not up to date, I can verify very easily. And I can be sure that dirty files won't be active as long as I don't expressly copy them. I know that SVN has this too, but I find CVS easier to use in general.
I guess for fast paced, very active development, yes CVS is sub-par, but for personal stuff, or for something that is patched say at most two-three times a month, it's O.K. It boils down to personal preference.
There were plenty of things I didn't like about SVN (separate folders per branch? yuk), having used CVS, and didn't find it 'trivial' to pick up.
1, atomic commits. I edit ten files, that's one checkin, rather than the per file checkins of cvs. On a low volume project, not a big advantage. if you've ever conflicted on a bigger project with cvs, it can be kind of a pain to resolve. seeing the whole commit of the other guy is helpful. If you don't run into this more than, say, monthly, it's not worth it.
2. offline diffs. svn has a whole copy of the repo, so you can compare history even if the central repo is down, or you're working from the beach. This one is pretty nice regardless.
svn is a pretty nice upgrade, if you're working with a distributed team.
I believe that one should use the best tool for the case, not the overall best tool in every case.
That said, I'll give a look at SVN. I can consider switch when I have the time if it is easy to import from RCS, because I do use it a lot here and there, mostly for plain text documents. I do not like maintaining unrelated things in a single repository.
Heck, the increased difficulty in creating feature bloat could even be considered a feature in and of itself, too!
(Incidentally, while I think that part of the reason why Git is this way is due to design differences from CVS, this isn't essential to it being true. Git is also better than Monotone or Fossil or Bazaar, despite being much closer in design, because it has network effects that those other systems don't.)
It is very easy to contribute if you know git. But you can mess it up if you don't know. I made a two-line bugfix patch to flycheck, and heck, I was nearly pasting the patch into a comment in issues because I didn't know anything about how to make a pull request on github and how to commit so that the puller would be happy. I didn't want to do sth. embarassing and spent two hours reading and reading how to submit my patch properly. And the whole fix was made and tested in about two minutes. If only I was able to submit a dumb patch, I wouldn't have to care if they used git, cvs, or tarballs and quilt.
That said, GitHub does let you use a web-based editor to create a git branch behind the scenes and open a pull request, with no VCS client required at all. So that's a point in GitHub's favor. (Again, it's not inherent to git, and a hypothetical CVSHub could do that, but GitHub exists.)
https://github.com/redox-os/coreutils
I am excited to see this as I was working on a similar project last year (rewriting the BSD userland in Rust), but it's from pre-1.0 Rust so not really as idiomatic as what is coming out of the Redox project.
If all of these utilities are running in userspace, per Redox's microkernel architecture, then what is the advantage of intentionally not making them feature-rich?
POSIX defines echo as only taking string parameters and no options, but notes that behaviour facing `-n` is implementation-defined.
BSD and GNU echo implement `echo -n` as not printing a trailing newline, but `echo` commonly calls to a shell builtin which may or may not follow that behaviour (and may switch behaviour depending on whether the shell is in "sh mode" or not), so `echo -n` could print `-n<newline>` or nothing whatsoever (empty string and suppressed newline) depending on the utils set, the shell, and the shell's runmode.
GNU echo also supports -e, -E, --version and --help options, and much like -n shells and other utils set may or may not support these.
For instance on my machine (OSX 10.11)
* zsh (builtin) interprets -e, -E and -n as options (but not version or help)
* bash (builtin) also does, except when invoked as sh in which case it does not and all parameters are literal (this may also apply to zsh)
* dash (builtin) interprets -n, but none of the others. bash note may also apply to it.
* BSD echo interprets -n but will print -e, -E, version and help literally
* GNU echo interprets all of the above
So if you use echo with any non-literal parameter, or with one of the parameters listed above, in a script you distribute to un-controlled third-parties as an sh script (rather than e.g. a bash or zsh script specifically) you will suffer from portability issues.
And that's just for measly trivial echo (and incidentally why you should always use printf rather than echo in scripts you try to make portable).
The point of coreutils is to have utilities that make up the ability to write scripts for and interact with your operating system, right? Well, what operating system?? A POSIX-compliant one? Or just a mostly-POSIX-compliant one? Or one with POSIX extensions? How would your utilities know the difference? How would the OS know how to deal with these utilities? Would the user know the difference?
Ultimately, each platform has quirks, and it is up to the developer to port and test their script or application to a platform and make any necessary changes. This extends to far more than just POSIX compliance.
> Ultimately, each platform has quirks, and it is up to the developer to port and test their script or application to a platform and make any necessary changes.
That is not humanly feasible and that's why specifications exists. You can't "port and test" your script to a platform which doesn't even exist yet, but if you follow the specification and the platform implements it (assuming it does so correctly) your scripts will run.
ECHO(1) FreeBSD General Commands Manual ECHO(1)
NAME
echo — write arguments to the standard output
For this definition, -n should mean merely a sequence of two bytes to be written to stdout. Why not just use printf instead? It is way more flexible and powerfull, and "printf x" always prints "{'x', 0}";.This is also why the old Windows POSIX subsystem was so useless. It implemented only the minimal amount needed to check off the box on a feature list, and none of the stuff you need to actually make a system usable.
That said, POSIX has a fair bit of braindamage baked in and people can be excused for ignoring the worst parts and instead doing the right thing.
Note that the -n option as well as the effect of
`\c' are implementation-defined in IEEE Std
1003.1-2001 (``POSIX.1'') as amended by Cor. 1-2002.
-- http://www.freebsd.org/cgi/man.cgi?echo
So there's probably some system (or shell) out there where -n doesn't work?https://www.gnu.org/software/coreutils/manual/html_node/Stan... https://www.gnu.org/software/coreutils/faq/coreutils-faq.htm...
I would hope your cross-shell scripts are also conforming to a certain shell script language, and are also resetting all environment factors which change the function of various commands. Not that endianness is ever a worry with a shell script ............
(Also note that POSIX supports printf, which you can use to insert any character string you like, basically)
How is that the most obvious response to parent's observation about echo?
I mean we see tons of comments with "small amount of information"/cryptic references that might similarly puzzle people in HN, but not equally many "I'll bite".
Why is echo -n risky?
Also Rust doesn't have exceptions so you have to wrap almost any function call with let/match/Ok/Err. Ugly.
I looked at one random file which turned out to be a `du' command implementation: https://github.com/uutils/coreutils/blob/master/src/du/du.rs...
Are they really starting a new OS level thread for every directory found? Looks like an easy way to exhaust system resources to me. Also I don't see the code that would collect error information if if the thread panics.
respectfully, This is pure nonsense.
First of all rust does not have runtime and AFAIK for providing exception you should have runtime to manage stack.
Second not every language should be like high-level languages, it is not the rule to be like C#,Java,Python,etc. I use a lot of them for my work when I need simple thing to do, but rust designed to do low-level stuff, and I cannot understand how having not having exception makes a language ugly (specially when you code in lowlevel).
Not quite.
1. you're supposed to handle errors around function calls which can fail, which is a strict subset of "every function call"
2. rust has a number of higher-order constructs to facilitate that handling[0][1][2], not just raw `match` statements or expressions.
That aside, for rust explicit error handling is considered a feature both at the language level (allows for less runtime requirements and much stronger guarantees — check out exception-safe C++ for what happens when low-level meets exceptions) and at the user level (by forcing a conscious and explicit decision, whether it's crashing the system, handling the error or passing the ball upwards)
> the code quickly gets bloated, doesn't it?
Does C code quickly get bloated? Because you're also supposed to check for error codes after each function call which can fail, and C doesn't provide much abstractive power to mitigate that.
[0] http://doc.rust-lang.org/std/result/enum.Result.html
You can use convenience functions: expression.expect("panic with this message if expression evaluates to an Err")
You can also use macros: try!(expression) makes it so if expression evaluates to an error it returns the error immediately, otherwise it does nothing.
This approach is still more verbose than using exceptions but on the flip side dealing with errors up front can make it easier to write reliable code.
FWIW I also share their opinion that rust is unapproachable.
Do you have a specific symbol you would like to change in Rust, and what would you like to change it to?
The only example I've seen (in a child comment to yours) is effectively a complaint that Rust has lifetimes and Go doesn't, which is effectively saying "you should have a garbage collector like Go does", which is an argument against a fundamental design decision of Rust. If you want to argue that you should always use a garbage collector, argue that directly instead of making vague negative comparisons between Rust's and Go's syntax.
It's a minor quibble to be sure, but it's the only language symbol that bothers me when writing Rust. Not sure what I'd suggest replacing it with...backtick, maybe? Pipe? @? ~?
There aren't many other special characters on a QWERTY board that aren't already used in Rust. Which I think gets at one of the stumbling blocks that I see in the various Rust syntax bikesheds among those who haven't worked in the language. It's just alien until you've used it a bit, especially if you're writing a lot in pseudocode-y dynamic languages.
That's an argument against introducing new notation for anything. That can't be right.
> static NAME: &'static str = "du";
Also, idiomatic code doesn't match on errors; it uses try!. And nightly now supports a ? syntax that's even shorter.
This way lets you achieve something similar to checked-exceptions, and in addition to gain composability of Err (and the like), without adding another language construct (which a language without exceptions would have to do). Got a great deal of experience in other languages that does this, and it's worked out pretty good so far for at least for me once you get the hang of it (i.e. functional coding).
Where I am sitting, that would be a wise choice for Rust, at least if I am understanding it correctly, as it tries to be a safer alternative to other system languages. However, arguably (:), it is OK in dynamic languages or similar that aims trades in correctness for conciseness and, arguable (again :), speed of development to simply have unchecked-exceptions.
Also, it is possible I misunderstood your comment and/or how that code was ugly, in which case I hope the downvotes/replies won't be too harsh :)
EDIT: typo/clarity
What you end up with are much more flexible "exceptions" that don't need a lot of extra compiler support.
With all these node.js / Go / Rust CoreUtils implementations I'm still hoping for one of them to actually match the efficiency of the original implementation.
edit: confirmed. http://stackoverflow.com/questions/930044/how-could-the-unix...
Shouldn't the derivative work still be covered by the GPL?
The project isn't terribly far along. I wonder if just starting a GPLed fork and building on that instead wouldn't be a better idea.
My general perspective on code I write that isn't for work - it has to be GPL. I refuse to have my code be yoinked by random corporations for their profit without having the code shared downstream.
Just my two cents.
If anything would be a 'tumor' by your reasoning around ideas, it would be proprietary software, which is what GPL prevents.
Also, do you think that Linux, for example, is draining the world's resources? If so, how?
Something like: Recipient of this software can do what ever they want so long they agree to a contract that bounds them to never do enforcement, in legal and technical form, for copyright and patents.
A complete ban on lawsuits for copyright infringement and patents, including the use of DRM technology that manage and enforces copyright and patents. It is also not limited to just the work that I distribute, but as a contract would cover everything the recipient create or has created, be indefinite, and with harsh fines if broken. I do not think a single person or company who refuse to use GPL would instead accept that deal. Will you be the first person to accept such deal and forever stop creating tumors on the worlds resources by putting software under proprietary licenses?
arguments over derivatives, clean room implementations...
something yanking out copyright notices to replace with "this code is now GPL"...
In that case, they should have no problem sharing it back under a copyleft license either.
One of the authors of the Python requests HTTP library has called out Uber for using Python, and almost certainly using requests, and not paying any money
Copyright protects the code itself from being copied, but the ideas, abstractions, overall design, and even individual APIs are not eligible for copyright protection
https://en.m.wikipedia.org/wiki/Oracle_America,_Inc._v._Goog....
Back in the BIOS clone days, they had someone read the IBM BIOS code, and write a detailed specification from it. They had someone else never look at the IBM BIOS, read the specification, and write code to implement it. (This was the "clean room" approach - the IBM BIOS was never in the room of the implementers.)
But that's still "derived" in the sense that it implements the same functionality. But it's perfectly legal. So "derived" doesn't mean "implements the exact same functionality as the other, and we examined it in detail to make sure".
Why did they do the clean room approach? So that IBM could never claim that they had copied the IBM BIOS, even by re-typing rather than electronically copying.
Well, if you're re-implementing it in Rust instead of in C, you're not copying it, either. You're making a completely new implementation. (Rust doesn't take C code as valid syntax, so far as I know, so typing in the same code from memory wouldn't get you anywhere.)
It's important to remember that copyright law protects works from being copied, not from being read. Clean room is a legal tactic used against an aggressive adversary, it's not something that's at all necessary or appropriate in the general case.
For GNU Octave, we stress very strongly that anyone who has read Matlab's source code is ineligible to contribute to Octave. This is because, should it ever come down to it, we want to be able to ascertain that our implementation is completely original, because nobody has read Matlab's source code. In a similar vein, I'm still waiting[1] for someone to implement the medcouple for Python's statsmodels, because I cannot do it myself.
--
[1] https://groups.google.com/forum/#!topic/pystatsmodels/LpsmIJ...
While you're free to invent any contribution rules you like for Octave, there's really no need for such drastic measures. It might give you a nice piece of mind, but it's not legally necessary - it's perfectly possible for someone who has read the Matlab source code to contribute to your project without copying anything. You'd just rather not have to think about it, which a is pragmatic, but heavy-handed restriction.
It's not surprising that your lawyer implied otherwise, as its "best practice" to guard against every feasible risk, no matter how unlikely. Understand that your lawyer is protecting you against a hyper-zealous misinterpretation of the law, not the actual law.
I appreciate that this is how it works today, but isn't that a completely outrageous idea? A well read, well travelled person will have seen countless things that will influence their future behaviours. It is not uncommon to completely forget a particular source of inspiration (sometimes we falsely attribute to someone else, and even other times we attribute it to ourselves!)
Lawyers tend to fuel the FUD around this with "best practice" concepts such as "clean room", which is a drastic overreaction. It's akin to saying "if you want to be an author you should never read any books, in case you accidentally copy one of them". Sigh.
I have sent in a few PRs to this project, and I have never done more than maybe glance at the source of coreutils, and it was for unrelated reasons. can't speak to the regular contributors, though.
But if having glanced at GPL source code prevents implementing similar functionalities in an entirely different language, that's a pretty darn strong argument for me to never look at GPL code again.
Not according to the lawyers who have advised us GNU Octave developers to never read Matlab source code.
This is only an issue if your derivative work is under a different license.
https://github.com/uutils/coreutils/blob/master/src/whoami/p...
looks suspically simmilar to:
http://code.metager.de/source/xref/gnu/coreutils/src/whoami....
The GNU coreutils are themselves rewrites of earlier BSD tools. The Rust version AND the GNU whoami.c both look similar to the BSD original https://github.com/weiss/original-bsd/blob/master/old/whoami.... Hell, the rust version is clearly a lot closer to the BSD original than to the GNU clone.
Honestly if you're going to have an opinion on this issue, you should try to educate yourself on all of the Unix that predates GNU.
I would really like to know if licensing agreements (and copyrights) still apply to ported code (assuming there's no patents on the algorithm).
It would if they did indeed read GPL'ed code to implement this. The SFLC has advised us to make sure we do not read Matlab code when implementing Octave code. The only thing that is known to legally work is clean-room reverse engineering. It's ok to read independently-written specifications of how the software works and reimplement that. Reading the software itself and reimplementing it constitutes a strong case for derivative work.
Improving the general security and reliability of all systems seems like it might potentially be a more valuable goal.
This is a bit of a bugbear. The most common way to use the coreutils is through standard Unix pipes, which does not create a derivative work. I don't know of anyone who has found the copyleft of the coreutils prevents them from doing anything they would like to do. The situation with busybox and Linux is different, as the coupling there was much tighter, and without it we would not have OpenWrt.
Also a little out of topic; Can more light be shed by fellow HNers on the debugging tools for rust, the debugging experiences from the security research point of view?
Rust works with GDB, so you end up debugging like anything else. IDE integration is being actively worked on, and sorta-kinda works in my understanding.
On Solaris, as one example, there is no stable syscall layer -- libc is your only interface to the system. There's good reason for that too; on Solaris, libc is updated with security and performance improvements on a regular basis for each new hardware generation and other improvements in the operating system. It automatically accounts for things that might cause pipeline stalls (think memcpy, etc.) in newer processor generations, and so on.
Now with that said, what you could advocate is that libc, etc. be rewritten in RUST, but exposed via the "C" ABI convention that RUST provides. It wouldn't be as great as native RUST, but it would still be an improvement.
There are no compelling arguments for getting rid of libc that I've heard yet for platforms where libc is well-maintained. On Solaris, there are strict compatibility guarantees for every interface, so applications can assume changes in libc will never break them as long as the interface they're using lists an appropriate stability level in the manual page.
A world where every language implements its own version of libc seems like it would increase the number of defects, not decrease them. We could of course argue that those implementations might be better than libc, but I think whatever possible benefit there might be there is lost in the likelihood that each language will have its own flaws in its implementation.
I believe the best view of an operating system is as a fully-integrated set of components. Everything from the kernel up to the system libraries should look and act consistently, and that's impractical to do if every component is viewed as interchangeable. Integration from top to bottom in an OS stack can produce incredibly great results and provide unparalleled reliability, availability, and performance.
My guess is that the Rust team consider that relying on a well tested C codebase is safer than writing tons of unsafe platform-specific code, to replace the functionality offered by libc.
Anyway, I think the Redo project do have a Rust standard library that calls directly the OS and doesn’t need a C library, since they are writing a 100% Rust based stack.
Someone did start a C stdlib implementation in Rust, but it appears to be inactive: https://github.com/mahkoh/rlibc
Anyone have any pointers on where I can find a good IDE for Rust? Should I just start praying to lord-JetBrains for something that works?
[Good lord, a downvote on this? I'm being completely sincere.]
1. You learn the command line interface to your compiler, which is invaluable and something you'll have to learn at some point, even if you start off on an IDE.
2. Similarly you learn the command line interface to the compiler's support tools like make, linker, debugger, profiler, source code revision control, grep, strings and so on, and more importantly how the whole process of 'write-compile-execute-debug' cycle is done.
That being said, nowadays I use an IDE because it is extremely helpful to have autocomplete (which is okay with Racer+Vim, but kinda hacky) as well as the hover for type information (name of types, type signature, etc.). Without the hover information, I usually end up doing ridiculous things like writing bogus explicit types to see what type an undocumented function from a library will return after compiling (Is it Option? or Result?). That is really annoying.
I didn't downvote you BTW.
An IDE lets me focus on the problem, not the language.
What exactly is there to learn? That you have to spend more valuable time on memorizing unimportant things?
There are now some pretty decent autocomplete plugins for emacs, so it is possibly an option if you already an emacs user. But emacs is emacs.
- Autocompletion
- Mass rename
- Source formatting
- Integrated debugger interface with breakpoint insertion and overlying of state on source
- Integrated VCS control
- Automated deploy
- Error display
- Automatic importing of modules
- Source cleanup (Automatic loop transformation)
- Automatically building my project
That's just a few things that I like an IDE to have. Some provide even more features that I like.I understand that there are other tools that do the tasks better, maybe even faster, but that's not what I want. I want to be able to learn one thing, and learn how to use all of it's features well.
Source formatting: rust.vim + rustfmt
Automatic building: https://github.com/passcod/cargo-watch
Error display: tmux, terminator or iTerm2 split planes
Rust has pretty good tool support in vim and Atom, for example (can't speak for the rest). We don't have stellar IDE support yet, but it's on the to-do list: https://www.rust-lang.org/ides.html
Personally, I don't miss IDEs. I like the approach of small, single-purpose, composable command-line utilities more.
You probably noticed how useless racer is on the command line: type code in your text editor and then when you want a completion, you go to your CLI and type "racer myfile.rs row,column" and there are your completions! :) Jokes aside: it's clearly superior (and intended) as an integrated tool than as a command line tool.
I'd argue that a debugger can offer the same kind of power being an integrated debugger as the (admittedly silly) racer example.
Split panes are a poor substitute for actual error display integration; highlighting the errors in the source-code viewer, combined with a cross-linked list of errors is far more useful.
Not only do you have a package manager, but a build system.
People love C# for that very reason, once they get used to Visual Studio, convincing them to switch to something with very limited IDE support is a tough argument. Most people have better things to do with their life. Rust will get there though, just give them time.
About the project, not being GPL is serious bummer for me. I don't want to GNU project isolate more than this and anyone other who has simple understanding of politics behind the scene will understand GNU as foundation have done a lot of good things for programmer/developer/hackers community.By far more than any other player in this area.
And since I have worked on Glibc for a while and I wanted to use rust just for experiment,I was seriously considering to work on this project if it were GPLv3.
vscode and textadept are good
In comparison, Eclipse starts in under a minute and uses about half that on my system.
Atom when I tried it was much to large for my system. I might look at it again.
Though if you want an IDE a plugin for IntelliJ exists too.
Now obviously any Windows port would have to call Win32 APIs instead of POSIX APIs (unless you're writing for a Cygwin-like layer).
e.g. type c:\home\foo.txt -> cats it out type c:/home/foo.txt -> 'The syntax of this command is incorrect.'
whereas cd will except either forward or backslash
cd c:\home/foo cd c:/home/foo
work as expected.
Secondly, Cmd is deprecated. Powershell supports both slashes and "type c:/home/foo.txt" works perfectly.
(on Windows use MinGW/MSYS or Cygwin make and make sure you have rustc in PATH)
This made me chuckle - if you have MSYS installed, which already comes with a windows port of coreutils, why would you want to use it to build a windows port of coreutils?Cygwin has a complete, up-to-date port of the Coreutils for Windows.
The POSIX + C language provides a a good, platform-agnostic way of writing systems utils that are easy to compile anywhere.
Can you show me a Rust program that calls CoCreateInstanceEx at all? Also, what is that for? The MSDN page isn't particularly helpful.
Doesn't this imply there needs to be the Rust compiler / Cargo (if it needs that) installed on the system? And what versions - stable / unstable ?
For someone starting at zero, would you recommend C rather than Rust or the other way around?
https://github.com/uutils/coreutils/blob/master/LICENSE
GNU coreutils is GPL, this is MIT.
I hope it's not Rust version 2 MB, C version 6 KB.
It gives a good idea if the replacements can actually be comparable replacements.
The second part is then to compare the number of features implemented and the behavior.
EDIT: Yea, this was meant as a reference to the coreutils-in-nodejs post on the frontpage.
It's akin to saying "ARMv7 assembly beats using Java for Android programming".
I think neither of these beats actual coreutils and their support base (still under impression of Termux capabilities...), at least not yet.
But like I said, both implementations suck compared to actual coreutils.
Not to be mean, just a genuine question (and yes I read the Why section of the readme).
Your second question presumes that just because something is old means it is the best. Traditional things can indeed work out well, but challenging tradition is how progress is made. And you can't know if the coreutils can be improved if you don't try, right?
I guess it doesn't intend to be a 1:1 rewrite of coreutils then ? The public available function to users will be rewrote.
Why non-GNU? To completely and thoroughly dispel the notion of "GNU/Linux".
I didn't count, but that may have taken less than 30 seconds.
> The above copyright notice and this permission notice shall be
> included in all copies or substantial portions of the Software.
As those have now been stripped away, it seems like a license violation.I'd like to think that in the future when "the open-source bubble" bursts and companies fall back to 70s-era mentality, developers would rediscover why free software is important and a new movement would would be born.
Sadly, I think hardware be permanently locked down by then. Perhaps by FCC-style law.
https://opensource.org/licenses/MIT
https://www.gnu.org/licenses/license-list.html#X11License
Maybe "copyleft" or "reciprocal" licensing?