Translating Quake 3 into Rust
immunant.com
immunant.com
> The Rust compiler also flagged a similar but actually buggy example in G_TryPushingEntity where the conditional is >, not >=. The out of bounds pointer was then dereferenced after the conditional, which is an actual memory safety bug.
1. https://github.com/immunant/ioq3/tree/transpiled/quake3-rs
2. https://github.com/immunant/ioq3/tree/refactored/quake3-rs
> we'd love to hear what you want to see translated next.
qemu [1] is all in C (probably C89 or "gnu89") and would make an interesting project, IMO.
The fact that they found a memory safety bug is excellent. Maybe translating old C projects to Rust via this method can be used to find memory issues as well as to aid rewriting in Rust projects.
EDIT: You can try the transpilation yourself on https://c2rust.com/
Here's an example of the output:
#![allow(dead_code, mutable_transmutes, non_camel_case_types, non_snake_case,
non_upper_case_globals, unused_assignments, unused_mut)]
#![register_tool(c2rust)]
#![feature(register_tool)]
#[no_mangle]
pub unsafe extern "C" fn insertion_sort(n: libc::c_int, p: *mut libc::c_int) {
let mut i: libc::c_int = 1 as libc::c_int;
while i < n {
let tmp: libc::c_int = *p.offset(i as isize);
let mut j: libc::c_int = i;
while j > 0 as libc::c_int &&
*p.offset((j - 1 as libc::c_int) as isize) > tmp {
*p.offset(j as isize) =
*p.offset((j - 1 as libc::c_int) as isize);
j -= 1
}
*p.offset(j as isize) = tmp;
i += 1
};
}It’s not. It should eventually improve but currently the goal is to
> build a rock-solid translator that gets people up and running in Rust.
A pretty good go would be the ability to recognise wide-spread high-level C patterns and be able to convert them to idiomatic Rust pattern, but even that would be difficult, generally speaking idiomatic C and Rust would likely be structured quite differently, especially on the data(-structure) side.
It's a much harder problem for a tool to recognise the design patterns C pushes you towards and refactor that to something more native to Rust.
The real issue with doing that isn’t imo, going to Rust, it’s understanding what still needs to be addressable in C. For example, a self-contained application has no external ABI dependencies, and so could assume everything that references that slice must only need to be Rust. A library on the other-hand might still need to present a C ABI.
https://github.com/not-fl3/miniquad/blob/c2rust/sokol-app-sy...
The generated code looks fairly straightforward, however I wouldn't call it "idiomatic".
What's exciting though is that (at least) both Rust and Zig get this ability to translate C code into the "native" language (and sometimes even finding bugs in the process). This basically turns C into a "meta-language" to create "bedrock" cross-platform API modules for a variety of modern languages, essentially the next step from "C as a cross-platform assembler".
And to be honest, this sort of "boring" low-level system API glue code which mainly talks to native C system APIs anyway isn't a joy to write in any language, and also doesn't benefit much from more modern language features, can just as well do it in C once, and then "transpile everywhere".
Lets say I want to expose a Zig library to Rust or vice versa, now I must call the Zig compiler from Cargo, or Cargo from the Zig build process.
With a transpiled library such a project is "clean", it only needs the Rust or Zig toolchain. The transpiling process only needs to happen once, and this can happen automatically on a CI service when the base library is updated and then made available automatically through the language's package repository.
Rust probably won't get the ability to transpile to Zig, and Zig probably won't be able to transpile to Rust. Some sort of Lingua Franca makes sense in a multi-language world, and C (or even a subset of C) makes sense for that because it already is the Lingua France, and has a very small feature set to consider. I'm aware that this is a controversial opinion though ;)
If we were going to create a Lingua Franca for such a purpose from scratch, it would probably look very little like C. It would probably look like a much reduced subset of Some compiler IR, or Haskell. We would want something that is typesafe and lowers trivially to an AST. As the barest of table stakes, a language with any undefined or implementation defined behavior whatsoever is obviously unsuitable for this purpose. After all, we need a wide variety of compilers to be able to easily generate their own IR from it, and for them all to guarantee exactly the same behavior when they do. C seems close to the worst choice possible.
Edit: Haha, I realized I invented WASM again. This happens frequently...
This project kind of makes me want to try porting Quake 3 into Clojure, but I want to know the limitations in doing so.
The biggest thing in java gamedev is avoiding GC pauses.
Languages like Clojure stress more the GC.
Java history is full of OpenGL and DirectX bindings as well.
Just not all of them have survived the original authors leaving the project.
Before Unity was a thing, Java already had the Java Monkey Engine.
What I'd really like to see is benchmarks that compare the rust version to C.
In many instances the Rust type system should allow the compiler to emit faster code because it has more data to use for reasoning about valid and invalid states of the program. Depending on how good c2rust is the Rust version should be roughly as fast, refactored to idomatic Rust it could be meaningfully faster than C.
But that leaves out other interesting questions that inform this comparison. For example:
* What about the average Rust programmer vs the average C programmer? Maybe two experts can produce identical results, but does one language vs the other make it so that a non-expert tends to produce faster code?
* What about "no unsafe Rust other than what's in the implementation of the standard library"? Most Rust code doesn't use unsafe, so is it fair to let the Rust use it? How much?
* What about "no code other than the standard libraries." I may not be an expert, but Rust's package ecosystem makes it easier to use code that's developed by them. This is sort of a slightly different take on the two previous questions at the same time.
* Should there be a time limit on development? If so, should the C have the same requirements of demonstrating safety as the Rust, that is, sketching out an informal proof of correctness?
* What about some sort of "how maintainable is this" metric. Mozilla tried to implement parallel CSS in C++ twice, and couldn't do it, because the concurrency bugs were just too many. Their Rust attempt succeeded, thanks to the safety guidelines. Is this comparison valid? Or could some sort of person who was better at C++ have implemented it in theory with no bugs? (this is sort of question 1, but slightly different.)
Lots and lots of ways to do this.
https://stackoverflow.com/questions/57259126/why-does-the-ru...
Theoretically I would argue that Rust could get faster than C, because it can make stronger guarantees about the code that would allow further and much more complex optimizations of the code. You could argue that the set of possible outputs of a piece of C code is much wider than a piece of Rust code and you could compile the code down to some sort of state machine that equivalently represents the original code in a more compact form. The more well defined the outputs are, the more minimal the output representation could be.
Finding latent logic and memory errors, gaining greater confidence, security and velocity.
> Has anyone demonstrated that well-written Rust code should run faster than C
Should? No. Can? Yes. See Brian Cantrill’s http://dtrace.org/blogs/bmc/2018/09/28/the-relative-performa... or Cliff Biffle’s http://cliffle.com/p/dangerust/
But last I heard Rust wasn't taking advantage of this due to issues with LLVM. The same benefits should also be achievable in C with pragmas.
c2rust has to be careful not to "restrict" pointers which can in fact alias.
> But last I heard Rust wasn't taking advantage of this due to issues with LLVM.
That's correct: https://github.com/rust-lang/rust/issues/54878
> The same benefits should also be achievable in C with pragmas.
I don't know that any compiler provides a noalias-all pragma. I don't know that it would be very sane either since C very much defines pointers as aliasing by default. Plus you wouldn't have any way to revert this behaviour and mark pointers as aliasing.
This is actually a huge issue for the CRuby JIT which generates C from the VM bytecode. Alias analysis on the generated C code, which is full of indirection, is basically useless and severely restricts the possible optimizations.
So I believe that when you are that close in term of performances, you can assume that they all have the same performance, and that it should not be a decision driver whatsoever.
1. Rust’s aliasing semantics are more powerful than C’s unless you use the ‘restrict’ keyword everywhere, which most people don’t. It’s well recognized that FORTRAN continues to generally outperform C in many numeric applications because FORTRAN compilers are free to perform optimizations that C compilers cannot or don’t, due to the fact that multi-aliasing memory is undefined behavior in FORTRAN. The rust compiler currently doesn’t pass the relevant hints to the generated LLVM IR due to a long history of LLVM having bugs in ‘noalias’ handling, in large part due to the fact that those code paths are so rarely executed while compiling C code, itself due to the relatively low usage of the ‘restrict’ keyword.
2. Implicit arena allocations. The Rust compiler has access to information about possible aliasing and overlapping lifetimes that it can use to replace multiple allocations with similar lifetimes with offsets into a single allocation that is then all freed together. This is a complicated topic, but work is ongoing to make this a reality.
Ooh, interesting; I hadn't heard about that but it makes a lot of sense. Do you happen to have an RFC or issue link? My google-fu is failing me.
https://static.googleusercontent.com/media/research.google.c... and https://static.googleusercontent.com/media/research.google.c...
The backends are getting so complicated now that I'm not sure low level languages can make that much difference assuming they generate clean IR.
Then you also get into architecture specifics etc.
dodobirdlord described the two major advantages Rust has on the language level very well so I won't repeat them, but the gist I'm making is that Rust simply allows more optimizations purely on the language level. Specifics of backend implementations are of course another topic completely (and is one of the reasons why I'm excited for the possibility of a GNU Rustc)
In the past performance infelicities have been fixed, but there are some that remain.
When I was working on compilers I spent as much time making such a script in a web-browser with scrolling etc as I did studying the results.
Facebook for some time did this with their "HipHop for PHP" transpiler which translates PHP into C++ and then fed it to gcc. That worked, but due to all the type conversion and similar magic required on all calls this lead them to compilation times of 24 hours and interfaces which were hard to interact with for language bindings, thus they went back to write a scripting engine (HHVM) which now runs their forked PHP.
I think in relation to Ruby, Python and other languages like them Rust's best application is leveraging the support for C FFIs to replace hot paths. While Rust has a significant learning curve it's probably much simpler, for say a ruby programmer, to pick it up and write safe performant code than C.
At the moment it wraps the mruby VM, but I think Ryan's plan is to switch to the CRuby YARV VM and then replace it with a VM written in Rust.
I'm also working on a CRuby JIT compiler written in Rust using CraneLift (a new compiler written in Rust).
Perhaps the Rust team should make their compiler work on reasonable resources, like 32 bit processors, before marching off to try to convert the world.
Or perhaps we should have rust2c instead, so people can take all this Rust code that everyone is pushing and compile it on modest machines.
mrustc is "rust to C", though it's not quite ready for the general case yet. It's primarily for translating the compiler, to ease bootstrapping.
It's time to let 32-bit processors go and move on.
The complaint above is about that. Running compiled Rust code on a tiny embedded microcontroller is not a problem.
https://medium.com/coinmonks/coding-the-stm32-blue-pill-with...
Eh?
The fact is, that there are a great many bugs in the wild that never would have happened in idiomatic rust in the first place. Yes, there is additional overhead to compiling. By comparison, if you use static code analysis on your C/C++ code there is additional overhead, and even then a greater chance of security bugs slipping by compared to rust.
In the end, used server hardware is relatively cheap in most of the world, which does pretty well with this type of code. There are a lot of people using mid-upper range x86 hardware and even to cross-platform compile to lower end arm64 support. At some point, it's less worth maintaining legacy hardware either for pragmatic and/or practical reasons.
Rust was started well after x64 and other 64-bit hardware was in the wild, and 32-bit hardware was for the most part no longer being produced. Supporting such systems is work that practically speaking is less valuable to most people.
As someone that really likes the old DOS based BBS software and related tech, and even supports Telnet BBSes today, I don't always like seeing this left behind to some extent. Last year, a large portion of the community blew up when 32bit was going to be dropped from Ubuntu, mainly because of the need required by wine and the gaming community.
This isn't about supporting 10+ year old apps and programs in an emulated environment... this is about porting existing code in an automated enough way to improve stability and support-ability over time by leveraging a safer language.
Rust compiles for and runs on STM32, a 32-bit microcontroller:
https://docs.rust-embedded.org/discovery/
And here's an index of supported embedded hardware:
https://github.com/rust-embedded/awesome-embedded-rust#drive...
That being said you would need to implement the basics of OS routines, but it should not be too hard
>Rust compiles for and runs on STM32, a 32-bit microcontroller
Cool, but it can't run the compiler, right? I think that's what GP meant.
I know there have been some changes to shrink memory usage for a bunch of data types in Rust but it's not clear if that has changed this limitation.
I'm trying to learn Rust now, and one of my eventual goals is to look into heap analysis of the compiler, because I agree this is some bullshit.
I want to run stuff on ARM eventually and cross compilation and transfer don't sound like my idea of a good time.
In the past, people have reported these things, and we've tried to fix or help where possible; see the discussion on https://github.com/rust-lang/rust/issues/60294 for example.
Making rustc compilable in a 32bits host might be a good idea, but it is far from a priority. If that is a deal breaker, as it is for OpenBSD for example, Rust is certainly a language that you should look at in the short term.
Even low-end Android phones have multi-gigahertz, multi-core, multi-gigabyte memory, so actually, many people have that in their pockets
Unless you're talking about people holding on to hardware from 20+ years ago then yes - basically everyone with a computer, or a mobile phone has something that fits that classification.
People complain about Rust having long compile times, but I doubt it would be unusably bad even on something like a core 2 duo.
git clone https://github.com/hyperium/hyper
Full Build: time cargo build (1 min 42 seconds without download) (349K LOC total [0])
Recompile: time cargo build (2 seconds) (around 18k LOC)
I don't see the problem. The full build also includes all the 53 dependencies. So we are talking about a compilation speed of about 3.5k lines/s on a dual core 11" laptop from 2017 with turbo boost disabled.
I've personally worked on C++ projects where changing a single header can cause 5 minutes of compilation on an incremental build. Any slowdowns you have experienced may have been caused by a specific library or project.
[0] according to cloc ~/.cargo/registry/src/github.com-1ecc6299db9ec823/