Rust once, run everywhere
blog.rust-lang.org
blog.rust-lang.org
https://github.com/crabtw/rust-bindgen
(Caveat: I think it would need to be fixed up to avoid unstable APIs in order to run on stable releases of Rust, as opposed to nightlies.)
> I think it would need to be fixed up to avoid unstable
> APIs in order to run on stable releases of Rust, as
> opposed to nightlies.
The code generated by rust-bindgen won't be using any unstable features, so even if rust-bindgen itself happens to require a nightly build of Rust (which would be less than ideal), it can still be used to generate interfaces that can be used by stable Rust.Examples of manual wrappers on top of bindgen - there are many:
http://erickt.github.io/blog/2014/12/13/rust-and-mdbm/
https://github.com/kenz-gelsoft/wxRust/blob/rust-servo/src/b... (I don't like the design of this one, actually)
https://github.com/rust-gnome/gtk/blob/master/src/widgets/ac... (better)
For perf, 30% of the CPU is burned in malloc/free, due to them copying strings around. The system has an arena allocator built in, and many of these strings might be able to go in there. Except no one is sure exactly how long the lifetime is on these things. So everyone copies everything just to be sure. Rust would force addressing this kinda thing up front. (And this project coulda added an refcounted structure or something, but it's a few hundred kloc of C and C++ so...)
Essentially ropes are character strings represented as a tree of concatenation nodes optimized for immuatability. A rope may contain shared subtrees so is really a directed acyclic graph where the out-edges of each vertex are ordered. Kind of search trees that are indexed by position.
[0]: https://www.sgi.com/tech/stl/Rope.html [1]: [pdf] http://www.cs.rit.edu/usr/local/pub/jeh/courses/QUARTERS/FP/...
Also, chrome being a web browser means a huge part of the program is strings.
I was just talking with someone on Twitter about this today, actually. I agree with you 100%, and want to fill that kind of stuff out. You gotta pick priorities, and since we'll have more new people than experienced people at first, I've been focusing on the basics. Now that the language isn't changing all the time, shortly after 1.0, I'll be able to start writing stuff like this, that covers topics in more depth.
I also think that libraries will pop up to help with this specific kind of issue, too. We'll see.
That said, I do personally think that Rust will one day be an excellent language to take the place of C in implementing fundamental cryptographic libraries. But first we need domain experts to hammer on the language and provide feedback.
The results probably won't be well-optimized, though; C isn't actually very portable at this level, there's too much implementation-defined behavior that needs to be taken into account. You'll probably spend as much effort trying to robustly handle the quirks and deficiencies of some obscure platform's compiler than doing an LLVM backend for the processor...
I wonder if anyone has updated the Cell SPU backend that was eventually removed. That is one of my favorite architectures. =)
[0]: https://github.com/draperlaboratory/llvm-cbe
[1]: http://thread.gmane.org/gmane.comp.compilers.llvm.devel/7594...
While Ted's post has been criticized by a few Rust programmers (see: http://tonyarcieri.com/would-rust-have-prevented-heartbleed-... ), it still brings up the argument that one's language choice isn't a magic bullet. There's an adage (I have no idea who originated it) that's along the lines of "every 'idiot-proof' system underestimates an idiot's ability to break things", and that's worth considering here. Rust prevents a lot of bugs, but it's not a guarantee of security. One needs to understand the actual problem that needs solved, and this is where the OpenBSD devs are focusing with LibreSSL (and their various other projects); their security track record stands as a good testament of that notion of understanding trumping language choice.
This isn't to say that I disagree with you. Rust is readily poised to replace C in a lot of contexts (IIRC, one of the obstacles right now is the use of a different allocator, so trying to use a Rust library from C results in some performance overhead), and it would be nice to see it used in a lot more projects. Just bear in mind that moving to Rust isn't the only step that would be required to fix these sorts of security bugs.
It will though stop ordinary random code having remote code execution or spurious data leak bugs.
That is really significant. It needs cheering and encouraging now, and soon it needs forcing.
Memory safe languages aren't bug free in terms of security bugs, as there are also logical errors, as shown by many Java exploits.
However, removing the typical C memory corruption errors out of the picture is already quite an improvement.
Rust and other system programming languages have more probability to succeed outside classic UNIX systems though, as UNIX and C go hand-in-hand and I doubt that will ever change.
The only non-classic UNIX is Mac OS X and it already has its own "Rust".
Zero overhead vs. the case where you'd also have function call in C.
But in some cases, especially with numerics, if you have a C library that is designed to be inlined, you get slowdown even calling C from C, if you prevent inlining.
A random number generator, SIMD-oriented Fast Mersenne Twister [1], was such a case when I tried. Just calling with `_attribute__ ((noinline))` makes it 2x slower, even when calling from C.
Could the Rust compiler make use of design by contract style annotations in the C header? Maybe it could reason that an unsafe block is in some ways safe.
There is a specification language for C called ACSL (which was designed for and is mainly used by Frama-C). Annotations look like:
/*@
ensures \result >= 0;
assigns \nothing;
*/
int foo(int x)
{
if (x < 0)
return 0;
return x;
}
I noticed there are annotations for Rust with the LibHoare library (https://github.com/nrc/libhoare). The example above might look like: #[postcond="result >= 0"]
fn foo(x: int) -> int {
if x < 0 {
0
} else {
x
}
}
It would be really cool if ACSL could be translated to LibHoare.We still use glibc, yes. Being able to use musl instead is something that the community and team have been experimenting with lately.
> can one build the compiler on Solaris/illumos, NetBSD and OpenBSD?
The last two have community members that keep the build green. I _think_ illumos works, but I'm not sure.
If this is true, it's because support for others hasn't rolled out yet, not because it has some inherent dependency on that version of libc.
Im still new to rust, so forgive me if this is a stupid question.
Now there are some languages, namely Go, that skip libc and just implement directly against the syscall interface. Go has the advantage of being able to draw from Google's vast experience interacting deep within the system, so it was comparatively cheap for them to do this.
For rust, it never really felt like it was worth the effort for the benefit we'd get out of it. It was more important to get the language done.
There's already been some tinkering on various operating system kernels written in Rust (even if they don't do much besides printing "Hello, world!" quite yet), so I figure a pure-Rust standard library is the next step in that.
MinGW links the system CRT by default, which is one reason that it can be fussy to compile Python extensions with it.
(I'm more curious about what a survey would reveal than I am trying to argue with you about what you said)
I was a developer on the Delphi compiler for some years. Statically linked, a minimal Delphi executable had no dependencies other than the *32.dll libraries. And there is no system functionality that requires use of some libc, or significant duplication of effort. The Win32 API doesn't even use the C calling convention.
Many C APIs take a callback function pointer, which in rust has the type `extern fn()`, and this is the primary use case for an `extern` function defined in Rust which doesn't have `#[no_mangle]` on it.
Yes, but it's a C++ API and includes a bunch of inline bits in the form of RAII classes and so forth.
Making those non-inline would be a noticeable performance hit.
Similarly, using the non-inline jsapi.h versions of the various inline stuff in jsfriendapi.h would be a noticeable performance hit: those were added for Gecko to use based on performance measurement and profiling.
The main place where this inlining is needed is in the DOM bindings, where pretty much anything you're doing is overhead on the path from JS to the actual implementation. This is especially noticeable when the actual implementation is fast (e.g. many getters in the DOM). In modern browsers the binding overhead is on the order of 2 dozen instructions or so. A non-inline function call takes... well, it depends on the ABI. On x86, with cdecl, you only have 3 caller-save registers but have to push all the args on the stack. On x86-64, with the AMD64 ABI (so everything except Windows), there's a ton of caller-save registers that might need to get pushed/popped around the call. Either way, the chance that you add noticeable overhead to an operation that's already <30 instructions is high. And that's if you only have to make one call. If you have to make _several_ such calls as part of the binding code, you're just screwed. And then you start adding APIs that compute and return all sorts of stuff in as single call (see the SpiderMonkey typed array APIs) and other such ugliness.
It is not so easy to interface with the other languages, you have to do it through the C ABI.
Doesn't that suggest a good opportunity for all these actively developed languages (Python, Ruby, Julia, Rust, Go, plus the output of GCC and LLVM) to get together and agree on some kind of type safe ABI for calling functions in a generic way? Even if it is just in the form of some metadata annotation of function signatures to go along with the C ABI?
1) Better binding to C++ API
2) Compiling via MSVC on Windows. Windows is a bit of a red headed step child for lots of projects and Rust won't have that luxury
#2 is something we care about, and want to have the option to do in the future.
It's especially important as Mozilla intends to start shipping minor pieces of Rust code in Firefox this year, but Rust-based components will be relegated to nightly Firefox until Windows support improves.
Here's a patch-in-progress for the first Rust code to be integrated into Firefox (Servo's URL parser): https://bugzilla.mozilla.org/show_bug.cgi?id=1151899
Gecko, on the other hand, will gain Rust in modules that can be changed without replacing the entire browser. I'm happy that one of the first will be the code that supports the URLUtils DOM API. When the C++ code was changed from being just "URL parsing" to supporting segment changes, it came to have way more than its fair share of memory safety bugs. It needs a rewrite and it might as well be in Rust.
You realize that if this is perfect, you've basically implemented C++. It's easy to implement trivial bindings, but much beyond that requires an immense, immense effort and development.
I know both are areas of experiments in Rust. I think the intent is to use llvm to link(?!) C++ code in a format Rust can use.
It's difficult to see this working except with trivial bits of code. For instance, how would you link to a type generated from a template with a string parameter? The most I can see is linking with an explicit subset of C++.
With C++ mangling (decorating) is specific per vendor, per compiler, and it might be even per compile options. Then you have different way of handling exceptions, new/delete, location of the this pointer, deep/shallow virtual tables, etc. etc. And for templates it's possibly very awkward to support them in any good way.
It could be that wrappers like SWIG might help there, but would require some amount of work, or rewrapping the C++ interface to be "C" like with "objects". Some API/library writers go to the extent where even if the library is written in C++, a "C" interface is provided, and a new "C++" interface is written on top of the "C" (e.g. it's not using the original C++ one). This might be due to problems with dynamic (RTTI) symbols and different linkage versions (since you've asked about MSVC - different runtime libraries loaded in your application - yes this is pretty normal for Windows).
[1]: http://blog.rust-lang.org/2015/04/24/Rust-Once-Run-Everywher...
https://github.com/draperlaboratory/llvm-cbe
http://lists.cs.uiuc.edu/pipermail/llvmdev/2014-August/07593...
However, don't expect it to generate code that's portable between architectures with different pointer widths.
That's a limitation I can live with. If my only concession is that I have to have two builds (32/64), and everything else is portable, I'll be a very happy camper.
I make no guarantees that it'll all work however; I don't test Rust->JS like I do Rust->PNaCl.
[1]: https://github.com/DiamondLovesYou/rust.git
EDIT: Grammar.