C-rusted: The advantages of Rust, in C, without the disadvantages
arxiv.org
arxiv.org
So, in the examples, when something calls "read", what does the analyzer know about "read", and how much of its "char *" buffer it overwrites?
If you're going to do formal static analysis, the big wins come from automatically detecting that thing A is inconsistent with thing B way over on the other side of the project. That's the stuff humans miss. If you have to annotate everything to death, nobody will use your tool.
This is sort of an ad for the "ECLAR software verification platform", which is apparently some expensive tool. The site is all onboarding, no info, "ask for quote" B2B.
From the paper, you can't tell much about what they're doing. From the ECLAIR site, you can't tell much about what they're doing. So, no clue if this is a good idea or not.
Although honestly, the red flag for me was the mention of a "field-sensitive" analysis in C. My experience doing static analysis on C code is that it tends to be so reliant on type punning that field sensitivity ends up kind of not working, at least not without significant effort on detecting and handling type punning in some manner (to say nothing of all of the fun ways people figure out how to encipher pointers).
But if you start with modern and cleanly written C and avoid certain techniques (such as type punning) I think it is clear that this can work. Rust fundamentally does nothing else.
I think I see a flaw in this plan. Everyone is on the curve, but not everyone will be as far along the curve as you or I would like.
The advantage of Rust is that the compiler always enforced all the good practices unless you explicitly opt out, and frankly the error messages that come out of the Rust compiler are a lot more helpful than the ones from the C compiler.
It's hard enough getting folks to treat compiler warnings as errors. Expecting every professional (let alone hobbyist) developer to conform to best practices overall may be expecting too much, and when it comes to security and reliability, we really are all in this together and need to make the best of what we've got.
If you want your programmers to write safe code and are in a position to force this, it makes no difference whether you mandate the use of Rust or whether you mandate the use of a safe C subset enforced by a compiler flag.
Ignorance or forgetfulness leads to safety by default in Rust. A single failure of vigilance in C leads to the status quo of unsafety.
This is no small difference. If someone doesn't want all the guardrails of Rust, there are still options like Nim and Zig that provide better options than C for greenfield systems projects today.
Intra- is "the other one".
Also, isn't "inter-" vs "intra-" a Latin idiom? It's present in all Romance languages. English got it through French.
Luckily I think this is an area where tools like ChatGPT (and modern word processor tech) can help us all out.
Isn't this how rust works?
People seem okay with annotating everything.
It is also the reason why Microsoft has acknowledged their idea of lifetimes checker did not went as well as planned, and they had to rethink it,
https://devblogs.microsoft.com/cppblog/high-confidence-lifet...
In my (caveat-filled) analysis of verbosity, Rust is on the wordier side of languages (about as verbose as Java).
https://danuker.go.ro/programming-languages.html#non-math-ma...
Go is only a bit more verbose. Programs have a wide variability which is not graphed.
In addition to verbosity, C bothered me with mandatory types, undefined behavior and having to manage memory myself.
Rust is not much better. It still has types and memory management. A lot[1].
I don't want to deal with that. The computer should figure it out. Not having to specify all that also translates generally to shorter programs, and this is why I appreciate terseness.
if (!x) return;
you know that x is non-zero etc. Or after x = malloc()
you can know that x is - at this point - a pointer that has
exclusive ownership (if x is non-null).So I think it is a viable approach. I have a toy C front end that does something similar. While this project stalled due to lack of time, I think it is doable with a modest amount of annotations and not more pain than Rust.
Proof overhead for memory safety can frequently be on the order of 5-10 lines of annotation per line of code.
That said, it would still be very helpful to have true first-class support for the GhostCell pattern in Rust, so as to be able to prove safety for somewhat more complex designs (including safe doubly-linked lists) albeit contained within a single module.
if (!x)
return;
// statment here
do_something(x);
Which statements on the line saying "statement here" would invalidate the zero check for x?So much so that we have even had scalable context-sensitive interprocedural analysis for at least a decade longer than we have ever scalable flow-sensitive analysis in most contexts.
For example, pointer analysis, which matters here, is ni that class.
It is easy to find context sensitive flow-insensitive pointer analysis that scales very well, and has been easy for >1 decade.
flow-sensitive -- still much harder in practice. A few papers here and there, but ...
You can do halfway flow sensitive analysis with SSA and get it in a reasonable time bound but it's not the same for something like this.
This is more of a problem than something like the standard library handling, because the standard library is a library and can be pre-computed into summaries, and the relevant summaries change very infrequently
(IE to the original example, how much of the buffer read overwrites does not change every glibc release)
Theses sorts of annotation based schemes are also not new - we've also had summary based algorithms that can do this sort of thing for decades.
The algorithmic capabilities here have not really changed, and the improvement in computing power has not helped enough because lots of the sound algorithms are, worst case, cubic or exponential.
In the end just remember if it was cheap, easy, and usable, it wouldn't be a problem, because it would already be done.
Do you have any references on why that would be slow? That's of interest to me :)
Also, perhaps is there a distinction to be made between data flow and control flow sensitivity?
Seems like these kinds of analyses are increasingly added to languages for nullability tracking so it might not necessarily be that slow. Especially if the type system is of help so that the analysis can be incremental, per compilation unit.
> So, in the examples, when something calls "read", what does the analyzer know about "read", and how much of its "char " buffer it overwrites?
As the text says: “the annotations possibly provided in function declarations.”*
The PDF also says “the large part of the result is obtained without any annotation at all, thanks to the fact that the C Standard Library and the POSIX Library have been annotated once and for all (and the same can be done with any commonly used library).”
I.e.: You know the function body is, according to some definition, let's say aliasing, consistent. You annotate the function signature, thereby expressing formal contracts for that function any caller has to fulfill; given the contracts are fulfilled, the function call will then be consistent.
Meaning, you prove the function body is consistent, you annotate the function, and you then get the guarantee that the function is correct for any callsite that fulfills the contracts. If you want to prove this using a formal prover, you have to annotate the function, and annotate the callsite in such a way that the prover can prove that the contracts are kept, AT THE CALLSITE.
Or am I missing your point?
To say the C preprocessor complicates analysis is an understatement.
I've built a thing like this for a project that had zero tolerance for any of C's well deserved faults, and yet, had to be written in C. I borrowed heavily from various other solutions (notably: Erlang) to ensure that the project was as safe as I could possibly make it. But I was always acutely aware that it only took one team member to declare a void * to an unbounded array for things to break out of that carefully constructed illusion of a safe jail.
Most of the work was done using some pretty tricky (and nifty!) macros and I really liked the end result, but I would have very much wished for that end result to be based on something more solid than C as the foundation. If you go down this route in the present day, unless you absolutely have to use something else. You'll be far happier and you will not have to be eternally vigilant against someone taking a shortcut.
Even getting folks to reliably use a linter can be a chore. Getting all C coders current and future to use these new techniques seems wildly optimistic to me.
Things might be different if you could preserve useful invariants all the way down to binary code, of course. Then externally compiled code could also be meaningfully checked for safety.
The main issue is the culture, being more in line with ALGOL/PL/Ada/Wirth languages point of view of why security matters, even if there is a bit of unsafe core around.
This is why code reviews are a thing. If you have competent programmers, with a CR process as part of CI/CD pipeline, you really don't need language features to handhold programming.
Case in point, even SQLite still has CVEs relating to memory safety: https://www.sqlite.org/cves.html
> Grey-hat hackers are rewarded based on the number and severity of CVEs that they write. This results in a proliferation of CVEs that have minor impact, or no impact at all, but which make exaggerated impact claims.
Also how most CVEs require arbitrary SQL injection (and if you have this access, you're screwed already):
> Almost all CVEs written against SQLite require the ability to inject and run arbitrary SQL.
It also explains that:
> Few real-world applications meet either of these preconditions, and hence few real-world applications are vulnerable, even if they use older and unpatched versions of SQLite.
---
It also explains that many CVEs are denial-of-service due to division-by-zero, which in Rust can only be caught by using the opt-in `checked_div` so would be equally safe in C and Rust(?)
See for example - https://nvd.nist.gov/vuln/detail/CVE-2023-0687
The CVE is correct that the “gmon” component of the GNU C Library contains memory corruption bugs. In fact, I’ve found even more than the specific one that CVE is about.
But, the claim that this is a security vulnerability is a bit silly. These are profiler functions, which are usually only called in non-production builds with profiling enabled (gcc -pg), and even then only from CRT startup code. It is rather unlikely an attacker can exploit any of these bugs in them without already having the ability to run arbitrary code in the process, at which point they don’t need these functions, they’ve already got all these bugs have to give them.
All I'm saying is that the CVE process is a human one, with lots of nuance. You're pointing out one aspect of that, that some CVEs are very practical, and others are more theoretical. I was trying to chime in with a similar way in which this is expressed, in that an identical bug in the standard library of one language may be a CVE, while in the other, it may not be a CVE, even with the same problem in both of them. That's "not an issue" in the sense that yeah, of course, Rust cares about this a lot more, so it's a more serious problem in Rust, so this process is legitimate, but it is in the sense that depending on how you're trying to do comparisons, there's more complexity than simply tallying up CVE counts.
With that said, I wouldn’t bring up this distinction here. The problem here is identical but Rust bills itself as a language where it takes responsibility for correctness bugs while C++ is a language where correctness is something that the programmer is supposed to provide. That doesn’t mean the Rust CVE is any less valid or complicated than the other CVE, it’s just that the bug would get assigned to Chromium or IOAccelerator. So if you’re saying that you shouldn’t just look at the number of CVEs and claim that Rust is somehow less secure: yes, absolutely. But if you’re saying this because you want to point out that Rust CVEs are somehow lesser because they are self-imposed, then no, that’s not true.
Even more general than that, it’s not about Rust specifically, just that number of CVEs alone, with no other context, is a bad metric.
> That Rust CVEs are somehow lesser because they are self-imposed
I am not saying that, yes. That would be incorrect.
Additionally, the easier memory safety is, the more time those competent programmers can devote to work on new ideas instead of babysitting code.
Or C coding on Windows with SAL.
Still, I’d like to see how it works out in practice. But where’s the code? Judging by the website for ECLAIR, which is the basis of C-rusted according to the paper, it may not be open source:
https://www.bugseng.com/eclair
In that case, it may still be useful in the tradition-heavy world of safety-critical embedded applications (which the paper alludes to in mentioning MISRA and certified toolchains), where each codebase is a proprietary island. But it won’t be usable for general-purpose C programs or libraries, nor will it be possible for the community to evaluate on its merits.
That's a shame, as it seems like a pretty decent idea, even if I feel like the paper is written a bit like an advertisement (it feels a fair bit biased: it's hard to argue that you can't adopt Rust incrementally at all even if it doesn't allow for gradual adoption in the same manner as C-rusted or TypeScript. Also, it definitely doesn't demonstrate a similar competency towards being able to express things safely as Rust does.)
With composable I mean the ease in which you can include a external library without having to study it first. Whenever I had to use C this was always a bottleneck and you resort to copy paste hell, with slight changes to the original source.
I'm an advocate of Rust for most software, mind you. The safety you get from enforcing ownership at function and crate boundaries is allowing the package ecosystem to grow rapidly at a high quality. The safety win partially mitigates the safety loss you get from adopting hundreds of small transitive dependencies maintained by random people on the internet. But if you've spend tens of thousands of dollars having engineers test and validate some piece of safety critical code, you're going to need a really good business case for touching it at all.
That advantage may be temporary.
Having existing C codebases in production which are time-tested and stable and reliable is a huge advantage of having to rewrite them from scratch, don't you think?
Also, you can check out C code from 20 years ago and build it in your personal laptop with ease, and be able to do the exact same thing 20 years from now. Meanwhile, Rust still fails to put together a single specification, let alone a standard one, and relies on reference implementations that may or may not exist in a few years.
Language stability, tons of tooling, validation, decades of experience, plus the fact that you might have spent 500 million dollars to develop your existing codebase.
Not to mention Rust is still a moving target and it's not ratified for safety-critical uses.
- Cyclone: https://www.cs.cornell.edu/projects/cyclone/ (2002)
- SAL: https://learn.microsoft.com/en-us/cpp/code-quality/understan...
Not to mention languages that are generally safer such as Ada.
They also fail to present a code example that looks more ergonomic than Rust.
I agree entirely about the ergonomics, and I think the authors also oversell what “safe” means for C-rusted-checked code, versus what safe means in Rust. There are some nice ideas here (and I’d like to see the implementation!), but it’s a stretch to say that something as incremental as this can provide the same program-wide guarantees that safe languages can.
[1]: https://www.microsoft.com/en-us/research/project/checked-c/
[1]: https://ziglang.org/
Would unsafe Rust work for that, or is there a C footgun you need that that doesn't provide?
The biggest thing that is missing is a `->` operator (or the safe-rust solution of auto dereferencing with `.`, which is disabled for raw ptrs). Having to write (*foo).bar all over the place is not a good experience.
The second biggest thing that is missing is a footgun. Auto casting integer types all over the place. To different integer types. To offsets for pointers. Etc.
Disclaimer: It's been a few years since I really dug into the unsafe side of the language, maybe things have "improved"?
There are some proposals going around that would solve that problem.
Has there been a breaking change in the past few years?
At the same time, Rust has grown to a large and complex thing and deserves a large update to make it more dev friendly (what C++ did years ago). I expect at least one big update before it can settle.
No such "dead trees tome" exists. ISO will explicitly point out that what you're spending an eye-watering sum of money on is just a PDF.
Some of the National Bodies will tell you that you can buy this document on paper, but what they're actually doing is contracting out to a local Print On Demand service which either didn't tell them it has a document size limit or they didn't listen.
If you splash out anyway eventually what you'll get in the post, instead of your imaginary tome is a CD with the PDF on it.
> that everyone should certify their compilers against.
Try for a moment to imagine that all the GCC and Clang volunteers own legal copies of the actual ISO document. Now, stop laughing and read. Nobody cares about the actual document, in 1998 there really were academic libraries who bought C++ 98, and they might still have C++ 98, but that's a 25 year old document. Today those libraries would tell an interested user, "I'm pretty sure this is online, go away and search" and they'd be correct.
Because it's a PDF not a paper document, ISO C++ is a moving target, what's stable about that?
One of the few things about C++ that I unambiguously admire is std::format() introduced in C++ 20 -- but actually the most important feature of std::format() wasn't in the published C++ 20 ISO document, it's an erratum, pasted in much later because it's obviously needed.
ISO C++ is only a moving target if one ignores that there is an industry that certifies compilers for specific standard versions.
I imagine a few people at Red-Hat, or the dozens of compiler vendors that take advantage of upstream work in clang can afford a couple of copies of the standard.
After all, they have circa 40 years of experience doing so since ISO/ANSI C89 was released.
That's what is so funny. You can indeed get the rubber stamp, and it's worthless. Have you watched that scene in "The Big Short" where FrontPoint goes to the ratings agency and basically gets told the agency knows these products are bullshit, but they don't make money unless they say it's good so that's what they do ? Turns out a "ratings agency" wasn't legally promising anything at all, they existed solely because people felt like if someone else says it's good, even if they were paid to say so, well it must be good. Hilarious.
But that doesn't matter because nobody wants a C++ 20 compiler which does what the ISO Document published in 2020 says anyway, because that's bad you want a compiler which does what the current document now says about C++ 20. It's a moving target.
And I doubt very much that the "dozens of vendors" have even stayed awake long enough to read much of the readily available and up-to-date "draft" document, much less that they insisted their employer spend actual money on the useless and perpetually out of date official ISO document. It's not a great piece of technical writing.
A reply quite fitting of the whole RIIR folks.
Please try this:
1. Install rustup on a clean Ubuntu 22.04 x64
2. Open the rust embedded book, go to the getting started chapter
3. Try the first part
Early in the tutorial "cargo install cargo-geneate" will gail due to E0282 in git-glob...
https://learn.microsoft.com/en-us/cpp/code-quality/understan...
I have a question. Who asked for this? Like who exactly raised their hand and said yes, I want a systems programming language that won't let me compile syntactically and semantically correct programs because they don't follow an arbitrary ownership model?
No one, and yet we have C and people use it. No one "asks" for the disadvantages of a given language, they are just trade-offs in exchange for some advantages.
Which means that when you present me with a systems programming language that enforces an ownership concept that makes it much more difficult to write allocators or even doubly linked data structures, and then use that language's existence to justify adding the same restrictions into a version of C, "who asked for this?" might well be the least offensive response I can muster.
That's not true. Assembler and machine code predate C. And do the job much more effectively/efficiently.
For example Menuet OS an OS in pure assembly has like 2MiB and has stuff like GUI, graphics, etc. What Linux does in hundreds of MiB.
But they need to rewrite it for each hardware platform.
We use languages that impose other “arbitrary” constraints on us in order to save ourselves from our mistakes. For example, we avoid implicit conversions from strings to integers, or we use type systems that make sure that we don’t treat “dollar” type decimals as if they were “euro” types. The ownership model is a similar concept but to avoid memory bugs.
> We use languages that impose other “arbitrary” constraints on us
> The ownership model is a similar concept [to type checking]
Except we didn't. K&R rejected Pascal for their work on Unix precisely because it was too strongly typed and didn't provide sufficient escape hatches. Arrays were typed by length, there was no pointer arithmetic, there was little or no typecasting or coercion, no generic pointer types, no null pointers, etc.
All these restrictions that were in place to create a safer language made Pascal completely unsuited to writing allocators, buffered IO, and other OS internals. So instead, they ditched Pascal and created C and had Unix up in running in just a couple of years. Maybe there's a lesson there.
I have to admit, this question sounds very strange to me.
Corporations don't get work done. People do. Rust has been designed and implemented by programmers.[1] There's countless hours of talks and interviews and many articles and books written on Rust by programmers. They did all that work without being forced by anyone.
> All these restrictions that were in place to create a safer language made Pascal completely unsuited to writing allocators, buffered IO, and other OS internals. ... Maybe there's a lesson there.
Rust lets you do explicit type casting and write `unsafe` blocks in situations where you absolutely need to dereference raw pointers or do pointer arithmetic.[2] It lets you create null pointers or mark pointers as non-null. But not all code needs to do that. In Rust, you can isolate the code that does from the rest that doesn't.
Isn't that just some sort of static analysis?
It’s a great idea I hope it catches on.
Then after a bit of time one realize zeitgeist is a super powerful thing and accepts the idea that "having a great idea" isn't something that exceptional nor valuable. Implementing the idea is.