C2rust: Transpile C to Rust
c2rust.com
c2rust.com
1. C tag name space
2. enum values are placed into the enclosing scope
3. structs defined within a struct definition don't go into that definition's scope
4. `if (a = b)` is not allowed in D
5. unprototyped C function declarations
6. C bit fields
7. C struct declarations without a definition
8. oddities of _Alignas and _Alignof behavior
9. _Generic
10. all the weird C compiler extensions
11. etc.
All these things wind up making translation about 95% feasible, and the rest is a mess.
FTA:
> We are developing several tools that help transform the initial Rust sources into idiomatic Rust.
It sounds like this translator they've built may be meant primarily as the first part of a pipeline of tools to help you eventually translate existing C code into idiomatic Rust, rather than just letting you compile C code using the Rust compiler.
The only use case I can imagine personally having for this tool _on its own_ would be in avoiding the hell of trying to target musl with a project that mixes Rust and C.
It sounds easy enough in theory, but in practice, it's a nightmare. For example, you have no clue if the enum values changed, or the api's acquired an extra argument. Being C, the linker won't give you a clue, either.
This problem simply vanishes when your XLang compiler can actually compile C files.
Examples: C and C++, Java and Kotlin. The latter advertises: "Write concise and expressive code while maintaining full compatibility and interoperability with Java" and has been very successful with that.
Rust-bindgen is. Which "just" tries to translate header files, since rustc can link to C object files just fine. And for what it's worth, I'm pretty sure most of your complaints do apply to that tool.
The consultancy company developing c2rust specifically helps clients translate their apps to Rust. IIUC these clients want to move from C to a memory and thread safe language without loosing performance.
c2rust is the first step in that process. It mechanically translates C into "C-looking unsafe Rust".
The engineers then go and start migrating from unsafe Rust to safe Rust incrementally.
This is a long process, c2rust speeds up a small fraction of it, but most of the engineers time is spent into translating unsafe Rust into safe Rust, and then refactoring safe Rust into idiomatic Rust.
Maybe it's relatively few, but they exist, and this is targeting serving their needs. There are other tools serving C/Rust interop use cases, and at a certain point the “this isn't the most common use case” dismissal is just tedious.
HN shouldn't be a place where things targeting real but narrow technical niches are dismissed because the niche is narrow or there is a much bigger superficially related one around.
(2019) C2Rust hasn't changed much in a few years.
The Rust you get out of C2Rust is a representation of the pointer semantics of C in unsafe Rust using a library of functions that emulate C operations. This is only useful if you desperately need to compile C with a Rust compiler.
Intelligent conversion of C to Rust is hard, and less useful as time goes on and more lower-level crates become available for Rust.
Was compiling into C a consideration? Seems like this side steps some of the compatibility issues (not to mention getting a lot of optimization "for free"), albeit while inheriting any compilation and ABI limitations.
Thanks! P.s. We met at GoingNative2013 - you were very kind with my many questions back then as well!
Glad I was able to help you at GN2013!
Basically all compilers should be multi-lingual. Especially, rustc given its abilities with the borrow checker. Ideally (within reason for maintenance reasons), even parsers as “plugins”. Still building the same HIR (or AST if applicable) within rustc, but with different parsers.
I’d love to see lot of experiments with Rust syntax in the vein of how CoffeeScript inspired the future of JavaScript.
But still getting all the benefits of incremental compilation, cargo, etc.
No. It means the custom C frontend inside the D compiler, with the semantic routines in the D compiler tweaked where necessary to support C semantics.
upd: Also, I'm not sure that implementing -funsigned-char in 2022 will be all that great for morale!
Adatran was terrible. Hard to impossible to edit if something needed fixing. Used none of the features that made Ada a decent language. The very limited code they did this to did work however. This was a huge project and included C and ada code.
What do we call C code transpiled to Rust. Crust?
Fun fact: in ancient Rust, what we now call extern "C" functions were called "crust" functions, pronounced like the word.
The goal of c2rust is not “regularly compile C to Rust and keep compiling that garbage”. The goal of c2rust is “migrate a C codebase to Rust”. That is the tagline of the Git repository, as well as the title of the RustConf 2018 presentation linked in TFA.
The idea is that you convert the entire thing to working (but unsafe and C-equivalent) Rust, then have your entire build pipeline and tooling ready to chip at it and perform the conversion to safe rust, as an alternative to performing the conversion piecemeal and having a C/Rust boundary which you have to keep moving about, and duplicated definitions for the interop.
Because the translation, while it works, is awful
Now, translating everything and then going function by function seems better
> We rely on Emscripten's Relooper algorithm to translate arbitrary C control flows.
Article on why Relooper isn't good enough and the superior Stackifier algorithm, which they probably should be using instead:
https://medium.com/leaningtech/solving-the-structured-contro...
Do you have an example where the read vanishes?
Notice that it's very easy for C programmers to write code they think is performing a volatile read, but isn't, whereas obviously the intrinsics reduce the scope for this error in the Rust (it is also, though that's not relevant here, easy to write C code that depends on imaginary semantics of volatile access and so the code doesn't actually always work, or it works but not for the reason the programmer expected)
void func(volatile unsigned *a) { *a; }Definitely something to be wary of. I assure you the c2rust source code does know about volatile access and Rust's intrinsics (I was looking at that code for other reasons), so if you work with this stuff I'd encourage talking to the people who wrote it to find out what the situation is.
This will be fantastic for helping encourage Rust on microcontrollers. The microcontroller world is very C heavy, so libraries for daughter board and other chips, and other example code, is often only published in C, despite the growing ecosystem of microcontroller chips and boards that have good support for Rust, you end up constantly pushed towards C due to the ecosystem basically only using C. Having a good tool to take that C code and give at least a mechanical, non-idiomatic Rust port is just fantastic and I'm looking forward to giving this a shot with a particular add on board SDK that I wanted to use on a Rust supported microcontroller.
The problem is that for complex enough projects, the architectural redesign is considerably more demanding than a "remaining 20%". I can imagine it also being quite irritanting, due to shortcuts one may take in C taking advantage of implicit application logic (independently of them being warranted or not), that can't be directly translated due to the Rust strictness.
If you're using unsafe functions to flag something other than Rust's safety considerations (e.g. Rust's core concept doesn't care that this flag bit disables the interrupt controller, and thus if you get it wrong now the product doesn't work, but you probably do so let's mark that "unsafe") the same likely applies for that too.
One of the things I like in Jon Gjengset's live coding Youtube videos is that he takes the time to write such comments, which means both the final code and the live session explain why he thinks this is safe, and once in a while there's a realisation while doing this - aha, this is the wrong way to do it, we need to change other things.
It is your job to then make it safe rust code :).
Edit: spelling
https://galois.com/blog/2018/08/c2rust/
https://github.com/immunant/c2rust/wiki/Known-Limitations-of...
There was a while there where they were trying to test this out on the cvs codebase, IIRC. It's a good candidate: upstream doesn't exactly move quickly, but is very much a real-world codebase, still in use.
It shows up in this document of contracts: https://www.esd.whs.mil/Portals/54/Documents/FOID/Reading%20...
* Contract Number: FA875015C0124
* Performer Name: GALOIS, INC.
* Agent Name: AFRL, Information Directorate MR
* Program Name: Cyber Fault-tolerant Attack Recovery (CFAR)
* Office: I2O
* Fiscal Year: 2015
* Obligated ($): 2,099,878c2rust also has a refactor command that helps with refactoring the generated Rust code.
Good test coverage is essential for this. Count how many bugs you've written when you were writing this code for the first time. Even if your bug-rate is 99% better during the rewrite, that may still be a significant number. Fine-grained tests aren't necessary, but end-to-end tests that touch every feature are crucial to catch regressions.
Once the rough conversion is done, it is necessary to refactor the code to take advantage of Rust's idioms to get safety benefits. Just 1:1 conversion is underwhelming, and it feels like replacing gcc with rustc. I did not realize just how recklessly pointer-heavy C tends to be until I saw it through Rusty lens.
The lodepng conversion was "meh". It's a good C code, but its structure was very different from what you'd do in Rust (e.g. Rust prefers generics over pointer casts, iterators over indexing or pointer arithmetic, has interfaces for steaming processing that C lacks). I don't know how far I can refactor the code to Rusty idioms and still call it lodepng :)
OTOH the pngquant codebase was mine, and I'm happy with the results. When converting I took advantage of Rust idioms, and the Rust version is nicer to maintain and even a bit faster.
[1]: https://lib.rs/crates/citrus [2]: https://pngquant.org/rust.html
Off course there's a different trade-off when bugs are involved: c2rust shouldn't add any bugs during the conversion, while citrus will.
I found refactoring the resulting Rust code somewhat error prone and didn't have great success with the automated tools. I'd recommend having a good test suite and suggest adjusting the C before the conversion to avoid using C features that don't translate well like the C preprocessor.
[1] https://delta.chat/en/2019-05-08-xyiv#the-coming-delta-chat-...
Discussed here: https://news.ycombinator.com/item?id=12056230
Basically F(C) = R such that eval_C(C) = eval_R(R) where = means something like "discernible effects".
Wouldn't an optimization pass be a transpiler then?
Haven't seen that terminology applied to individual passes.
All transpilers are compilers.
C2Rust is a transpiler and a compiler.
Not all compilers are transpilers.
Graal for example is a compiler but not a transpiler.
Just like all cheese burgers are burgers but not all burgers are cheese burgers. Nobody questions the term ‘cheese burger’ because we already have the word ‘burger’.
This term seems completely useless to me.
High to high translation. Limited lowering.
> Who decides on this layering of which languages are higher level or lower level?
You know it when you see it.
Who decides what makes a book a 'horror book'? There's no authority on that either. You know it when you see it.
> C2Rust doesn't call itself a transpiler, it calls itself a translator.
Yes a translator - a translating compiler - a transpiler.
> This term seems completely useless to me.
Not sure why the term seems to wind people up so much - it just add a little extra info.
It’s like saying cheese burger instead of burger - it’s a little more info. It’s not saying that it’s not a burger as well, or that the techniques involved aren’t broadly the same.
edit: I see it is part of a broader toolchain, which makes a bit more sense
Has anybody incorporated an AI into c2rust to do the heavy lifting and learning from the compiler errors to self fix the transpiled code?
(But seriously, I think it's fair to say "as of 2022, no, AI can't do that")