C2rust vs. Corrode
jamey.thesharps.us
jamey.thesharps.us
So here is me acknowledging Jamey's work: I personally did take inspiration from Corrode, and I was expecting to work on Corrode proper when I joined Galois. I've re-read the CFG module of Corrode several times (as well as the Mozilla paper, and some older literature).
All that said, I also want to point out that Corrode hasn't had any activity at all since last April - and that's not for want of PRs piling up. I'm not criticizing here, since I understand that managing an open source can be quite time-consuming and stressful, but I feel like this also does need to be mentioned. Also, c2rust can be freely forked. Once the DARPA funding runs out, it is my hope that the Rust community will become the maintainers.
Finally, regarding the many improvements that can be automated: that is next up on our plate!
My initial attempts used top of tree corrode, but I found that the relooper got in the way of writing idiomatic rust, so I restarted my translation efforts and actually pinned corrode to before the relooper and backported a lot of the improvements that happened since the relooping effort https://github.com/danielrh/corrode/tree/unprocessed_loops
On one hand I'm very glad corrode supported the relooper: it got me excited about the capabilities of corrode, but in the end I think any serious port of a project is going to need to think about idiomatic control flow and will require a minor rewrite of the C code. The improvements to the C brotli encoder happened here https://github.com/danielrh/brotli/commits/corrode to remove the gotos and some negative array access.
All in all, the process to get the brotli encoder into rust was very smooth. Much of it may be due to the google engineers being extremely careful about memory ownership and design.
Anyhow I'm excited about c2rust, but am nervous that corrode may be de-emphasized since I know it worked well for me.
If you want to know more about the experience of porting a large existing project to rust, I posted about the rust brotli decoder here https://blogs.dropbox.com/tech/2016/06/lossless-compression-... (subsection "The Port")
Offhand, here are some sizeable additions:
* Keep track of some extra information from the C source about what basic blocks were in loops or branch points, then use than information to try to extract back out similar looking Rust
* Support translating `switch` to `match`, complete with collapsing some patterns together
* Properly handle initialization
That said, c2rust can be invoked _without_ relooper enabled if you so wish. In that case, it will simply refuse to translate code with goto's.
I'm sympathetic to the terms of a non-fundamental research program funded by DARPA. Easy street with that agency is to get designated as doing fundamental research. If you aren't at a university, that is an uphill battle that most prime contractors won't be willing to deal with.
However, the pre-pub regime is not impossible to live under. You'd have to do something like: have a private repo/branch where you do the work that's funded by DARPA, rebase constantly against the public branch, and send once a month public release request with a diff to merge into the public one. Probably after the first few months, the PM will come up with a way to streamline the process.
You lose some energy to friction, but your other option is to just not take the money at all, so it's a question of do you spend 80% of your time working on this project and 20% of your time dealing with administrative bullshit, or 0% of your time working on this project.
It's a choice that's given to many.
That makes it incredibly difficult to collaborate with others.
My ideal funding model is to have someone slide a manilla envelope full of cash under the door every month, and never speak to me or compel me to do anything. There are not that many opportunities like that out there. If you can't find one of those, you're going to compromise.
As harpocrates mentioned, we've been taken by surprise at the interest in the c2rust project. We've updated the website to acknowledge the inspiration we drew from Jamey's Corrode translator. Our apologies for not including the acknowledgments from day one!
I want to take the chance to mention that c2rust goes beyond transpiling C to (unsafe) Rust. We anticipate that even in the limit, the automatically translated code will need human-in-the-loop refactoring to arrive at idiomatic Rust. Therefore, we provide compiler plugins that instrument the original C code, the new Rust code to check whether the two programs behave identically when fed the same inputs. The checking can either happen at runtime (using a multi-variant execution environment to run the processes side by side) or offline based on logging. We'd love to hear what everyone thinks of this aspect too.
At the same time, my respect for Galois and DARPA (and other contributors to C2Rust) mostly only went up from reading it. It's unfortunate that corrode wasn't acknowledged, but I feel fairly confident that was either for a real reason (like some government contract thing) or forgetfulness.
It seems they acted reasonably and well, and did quality work. Their offer to fund a few tech talks seemed generous (and mutually beneficial); I wouldn't be surprise if they'd have paid for even a less organized talk, out of gratitude/respect if nothing else.
In any case, I hope Jamey is able to feel better about this and find some great work soon; seems like a great person to work with all around.
>> certain “features” of C that Rust didn’t provide.
Which then links to a markdown document that lists, among other things:
>> bitfields
I've used bitfields to represent demuxers in memory constrained embedded programming and as far as I know, the devices still work today, so I am confused as to why the sarcastic quotes there.
Sounds more like he's talking about the "Likely Won't Ever Support"[0] section.
[0]: https://github.com/immunant/c2rust/blob/master/docs/known-li...
The key to this is figuring out the comparable representation for data. Mostly this is a problem with arrays, since C's array/pointer system lacks size info. All C arrays have a size; it's just that the language doesn't know it. The trick is to figure out how the program is representing the size info. Somewhere, there was probably a "malloc" which set the size, and you may have to track backwards to find it. Then you can replace the C array with a Rust array that carries size information, and maybe eliminate variables which carry now-redundant size info.
That would produce readable Rust. But it requires whole-program analysis. That's OK, that's what gigabytes of RAM are for.
There's more discussion in the replies to https://news.ycombinator.com/item?id=17382464 that you may've missed.
If that is the case, I would hope Galois would just given him some ex-gratia payment.
[0] https://github.com/immunant/c2rust/commit/e0d3adf656db000b1c...
Quite a few products on the market for this on top of free stuff like Qubes or Muen. One can also use multiple boxes with KVM switch if worried about breaks. Anyone with DARPA or Defense experience think this proposal has a good chance of working?
I don't think any one thing at either the DARPA or Galois level can change this policy. It's a question of whether or not you are doing fundamental research, i.e. are you a 6.1, 6.2 or 6.3 program. If you want to do fundamental research on an applied research contract, you're going to have a bad time.
> Anyone with DARPA or Defense experience think this proposal has a good chance of working?
Frankly, no. This is a technical solution to a problem that is much deeper.
How does having a compartmented mode workstation help with this problem? Let's say that there's the main fork A and the private fork B. Someone comes up with a sensitive change to B that is, I dunno, about 50 lines. Someone else working on both A and B sees this change to B and they think oh hey, I could do something similar in A, so they write, from scratch, a similar change and commit it to A. They do this without having any file cross any isolation boundary, they just type stuff with their fingers.
I remember seeing something similar but i cant find it. Somehow, it analyzes C code and produced C code with safe pointers, etc.
Can someone please help ? Thanks
(Not ASAN or binary-instrumentation projects)
I say "just", but actually it addresses what I believe to be (by far) the most difficult issue, which, as other commenters have mentioned, is determining in the general case whether a pointer is actually being used as an array iterator. And it also determines whether that array is a fixed-size array or a (potentially) resizable array.
For example, in this code snippet:
void foo1(int *p, int n) {
foo2(p, n);
}
does p point into an array, or just at an int? You have to deduce it from context. And there's no theoretical limit to the amount of deduction you might need to do.The tool doesn't automatically translate regular "non-array" pointers yet, but that can be a straightforward task as the SaferCPlusPlus library has a safe, general drop-in replacement for pointers. Translating to "performance optimal" safe pointers is another story. There you have challenges similar to translating to safe Rust I think.
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus-AutoTransla...
I don’t know how to parse this. Is the author talking about corrode 2.0 being stolen? Or that someone would pay him to write code that can’t be released or discussed in public?