Building on Rock, Not Sand
robert.ocallahan.org
robert.ocallahan.org
Yet, everyday, I have to field questions about why our low-level, network-facing system code is not written in C... old prejudices die hard.
Really I think these days, in fact since the Morris Worm, people should be asking the opposite question: why is network-facing code written in C given that almost any error can lead to an exploitable security compromise?
I do not buy the performance argument at all any more. So few systems are really performance-constrained by network processing (usually it's the network itself or the disk), you have to write horizontally scalable code anyway if you want to make use of multiple cores, so almost all cases where it's a bottleneck are more cheaply dealt with by scaling rather than tinkering with the software.
There are people doing HFT in Java. The managed languages are quite capable of great performance with a little effort and imagination.
I'm asking because my own gut estimate of the cost of automatic memory management is significantly higher.
The cost of array bound checks is generally negligible with a modern compiler/CPU and the speed/safety tradeoff is not even worth haggling over for any code that's exposed to the internet.
In my experience of managed software 99.9% of performance problems are due to stupids and 0.1% are due to marshalling and GC costs. An average unmanaged developer seems to assume the opposite.
I mean, my experience is pretty much the same as yours, to the point where if someone blames the language for poor performance I'm immediately skeptical of the claim. But other cases do exist, is especially for problems that are CPU intensive.
It's a very legitimate question to ask how it was measured.
Obviously they're happy with their performance and we can't comment on that, but the other claims regarding speed vs. C were more general.
It's like saying that a reinforced door with a lock is more expensive and heavy than a regular cardboard-and-planks one, and is slower to open. It is! But is has other advantages, not attainable otherwise.
It's not correct to say that they use additional resources to provide memory safety or security.
Python for instance uses additional resources because it simply doesn't have a performant implementation and Java uses additional memory to gain execution speed.
Swift and Rust are memory safe, but don't need a GC, so those memory safety advantages are in fact attainable otherwise.
> I'm asking because my own gut estimate of the cost of automatic memory management is significantly higher.
Is reasonable and not dripping with bias? My gut estimate says it doesn't matter in many cases. So where does that leave us?
Based on submission history, I'd guess Haskell.
The typical factor mentioned in the Haskell world is 1/2 the performance of C.
Rust is super interesting to me because it prevents the most common vulnerabilities w/regards to memory safety that show up time and again, but it's accessible enough that it's got a big community and a ton of high quality libraries. I think that's a pretty awesomely potent combination.
On the other hand, it probably is less effort to rewrite PolarSSL in Rust than doing that proof.
It's probably even less effort to convert to SaferCPlusPlus[1] (essentially a memory-safe subset of C++). There's even an tool[2] (under construction, but functional) to do a lot of the conversion for you.
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus
[2] https://github.com/duneroadrunner/SaferCPlusPlus-AutoTransla...
For one thing, the "this" pointer problem is pretty much unsolvable in C++.
Well, you could just avoid/prohibit the explicit and implicit use of the "this" pointer. I.e. prohibit non-static member functions[1], right?
[1] https://github.com/duneroadrunner/SaferCPlusPlus#practical-l...
Why not just use a language designed for memory safety (which is most of them)? It's a whole lot easier, plus you get an ecosystem of actually safe code.
Well, passing a "safe this" pointer as the first parameter to a (static) member function isn't that hard to get used to, is it?
But I don't disagree with your gist. If memory-safety is your only concern and you're not trying to salvage an existing codebase, then Rust might be a better choice.
But if you're trying to add memory safety to, say, an existing C implementation of SSL, the SaferCPlusPlus route would probably be less effort. Even when non-static member functions are banned.
> (which is most of them)
Personally, I consider RAII (deterministic destructors) an essential feature for safe, efficient programming at scale. That leaves only two choices, C++ and Rust, right?
And if there's a memory-safe C++ option available, there are some arguments for choosing it over Rust.
That's different from directly writing C...
But I bet you'll end up writing tools to make creating a safe program of any non-trivial size feasible. A large class of these tools is usually called "programming languages".
Anyway, PolarSSL is definitely not transpiled. Its source language is C.
... and your point is?
EDIT: more recent bugs:
These bugs include double-free and use-after-free - which are not covered by the original analysis, right?
A lot of infrastructure software has not been built with security in mind, but the game has changed, and this old C software is getting eviscerated.
If one looks at the fixes for these buffer overflows in dnsmasq, no other conclusion can be drawn except that holes are plugged in a leaky sieve. This could work for software that's done, eventually. But if something's under maintenance, using plain C is security suicide.
Given that the number of C projects that use similar toolsets is... low (?), I'd guess that said comparison would favor rust, but I'm prepared to be surprised :-)
In the ACSL manual in Example 2.10, does the programmer have to add the asserts? What if they make a mistake in the annotation?
TrustInSoft could prove that part of code too. There's nothing special about it. It's just that they didn't bother.
Sure, you can get very close to 100% in practice, but even if you're fine with this level of guarantee, the amount of discipline and cognitive overhead required for doing that in C makes Rust's learning seem like a piece of cake.
Centralizing the memory management code to some well tested core---like the Rust compiler or some C lib---sounds to be the crux of wiring memory safe code, not the language.
The language has not been formally specified which is a concern, but this is an active area that is also being worked on in collaboration with universities. This should make it easier to write unsafe code in combination with theorem provers like Lean or Coq for even greater confidence, and ensure that any soundess holes in the type system are found.
My point is that it’s not possible to just isolate the hazardous chunks into a well-tested core; it needs to be pervasive in the language’s memory model and type system in a way that C and C++ just don’t have.
> Without these, you simply can’t have a fast,
> memory-safe language: either you surrender
> speed, by boxing everything, or you surrender safety.
Not only can you have speed and safety without Rust's rigid enforcement of exclusive ownership, you can do even better--even faster and with even stronger security guarantees:https://www.youtube.com/watch?v=zt0OQb1DBko http://www.ats-lang.org/
Empirically, programs written in memory-safe languages produce multiple orders of magnitude fewer memory-related bugs than C and C++ codebases do.
> Centralizing the memory management code to some well tested core---like the Rust compiler or some C lib
That isn't practical in C (or in C++). "Memory management" here means "pointers", and you can't program in C or C++ without pointers.
The ultimate solution to the problem is formal verification. Rust solves this for buffer overflows, but nothing else. It's a one trick pony. Someone posted this interesting nugget on HN the other day:
https://www.youtube.com/watch?v=zt0OQb1DBko
Aside from some of the awkward ergonomics, its mechanism for formal specification looks brilliant. And you get to keep C's bag-of-bytes object manipulation and pointer arithmetic (when you want it) without having to resort to unsafe{}.For projects where it's worth the effort to carefully declare precise typing semantics (because safety, performance, whatever), I want a wholistic solution. Otherwise my time and money is better spent throwing Javascript or Python at the problem, which solve the buffer overflow problem just as well.
Completely false. Rust's design prevents all memory safety problems (that's what "memory safe" means). Buffer overflows aren't even the most pernicious kinds of memory safety problems anymore. Use after free is worse, and Rust spends most of its complexity budget on preventing that.
But if you watch the video, the presenter makes a great point: Rust's borrow checker is an amazing piece of technology, but it's inaccessible to the programmer. It's an implementation detail used to provide proofs for a narrow constraint--memory safety. Imagine if Rust provided syntax and semantics which not only allowed you to effectively implement the borrow checker yourself using a more general declaration system, but implement any other kind of formal specification needed to prove the higher-level semantics of your code.
In other words, imagine if a more general formal specification system were as first-class as the ownership- and mutability-oriented syntax are in Rust; a language that unifies the annotation model of solutions like Ada SPARK and Frama-C, but which is properly integrated into the language.
And that's what ATS is exploring. As the presenter says, ATS might be ugly as a systems language (because misallocated complexity--some easy things are too complex, some complex things are too easy--increases cognitive load and reduces efficiency), but its mechanism for formal specification is brilliant. Improve ATS, or apply it's novel approach to a language designed as a daily driver, and you'd finally have a realistic answer to the plague of buggy infrastructure software.
I realize this sounds like I'm making perfect the enemy of the good here. But I stand by my point: major failures like Equifax rarely involve memory safety, per se. Arithmetic issues are far more common, and even those are on the long tail of a much larger issue; namely, an inability to [efficiently] provide verifiable specifications for higher-level semantics. We hyper focus on buffer overflows, arithmetic overflows, etc, because we understand them and we know (at least in principal) how to fix them. But those are psychological blinders that cause us to miscalculate relative risks. We tend to overestimate the cost of problems we can fix relative to the cost of problems we're unsure about how to fix.
And you are missing an important point here. Once you have a reasonably sound and rich type system, you can leverage the type system in library API design to eliminate classes of higher-level bugs. For example, Rust crypto libraries can leverage Rust's affine type system to ensure you don't use a nonce more than once. The Apache Struts vulnerability was about failing to distinguish trusted vs untrusted input; that distinction can be expressed and checked in type systems.
This kind of emperically-driven motivation is great! Would you mind pointing me to some sources? My google-fu didn't immediately yield anything helpful.
I also have to point out that these pesky C buffer overflows are trivial to avoid in C++: just use std::vector::at.
C/C++ is used to write safe code for medical and aerospace applications every day. The compiler for the languages like C, C++, Ada, Rust or whatever, is not enough.
You can get better static and dynamic code analysis and test coverage analysis tools for C/C++/Ada than you can for Rust.
"Safe" in the context of medical and aerospace means something very different, but is much closer to the meaning of "Secure" in this context. No compiler is ever going to prevent you writing insecure code - there can always be a logic problem, bad choice of crypto algorithm etc..
How comes we still catch lots of errors in reviews there? How comes that the best paying gigs for c/c++ coders are all code review? Best practices and an excellent toolchain don't help if they are not used. A compiler/language that enforces those is a giant leap forward.
> You can get better static and dynamic code analysis and test coverage analysis tools for C/C++/Ada than you can for Rust.
Of course, but comparing the toolchain of a relatively new language with those of languages into which - literally - billions of dollar were put does only make a temporary point. And with lessons learned from those billions incorporated into the design of the new language, closing the gap will be much, much less expensive and time consuming than the initial development for the languages you mentioned.
Not sure what you mean about code review. Security reviews? I guess that's because C and C++ are easy to misuse and most programmers, teams and companies aren't that good at writing correct or safe code.
But we already knew that and the solution is not as easy as switching to a different programming language.
Rust tends to push you away from using unsafe all the time. Unsafe is a pain to use, because you don't have all the nice pointer operators you do in C and C++, so programmers naturally default toward working in the safe language. Even if you use unsafe more than you should, Rust tends toward much safer code than C and C++ in the aggregate. (This has been observed empirically.)
> I guess that's because C and C++ are easy to misuse and most programmers, teams and companies aren't that good at writing correct or safe code.
If you replace "most" with "virtually every" (i.e. everyone who isn't writing avionics/defense/aerospace/etc. code), I agree.
> But we already knew that and the solution is not as easy as switching to a different programming language.
Programs written in C and C++ empirically have far more memory safety related problems than programs written in memory safe languages do.
This has not been my experience.
One of these things is not like the others, One of these things just doesn't belong, Can you tell which thing is not like the others...?
Crack is smoked by humans every day. This is an Argumentum ad Populum (https://web.cn.edu/kwheeler/fallacies_list.html).
And I dispute the statement anyway; what is "safe" code without a guarantee of memory safety?
Memory safety is the absolute minimum for safety critical code.
It seems to be surprise for most people that you can write memory safe code in C and check for that statically and that includes static stack and heap exhaustion checks.
Right now Rust is slower than C (though not always slower than C++), produces bigger binaries, and compiles slower. But more safety required a more expressive language, and a more expressive language contains more information for code generation and optimizer. It's certainly conceivable that in a few years the Rust compiler might generate faster code than C compilers.
Why do you think so?
Dnsmasq also includes a DHCP server and the ability to read a blacklist to act as an ad blocker. In contrast, the "trust-dns" project is more of a replacement for the "bind" program instead of "dnsmasq".
If your intention was to only show that "non trivial Rust code exists", that's fine. However, some others might get the wrong impression that it's a Rust version of dnsmasq.
Would it not be great to have a language as powerful as Rust but with the ease of Go?
Anyhow, must be like CAP. You can't have everything in a language.
The hard part of Rust is strongly tied to ownership and lifetimes. You can’t get rid of them and keep memory safety without introducing garbage collection on at least almost everything. And thus you’re roughly at Go.
Robert is now a Rust proponent because it works.
From today’s article, I think this is the money quote:
> My sincere hope is that people will at least stop choosing C for new projects. At this point, doing so is professional negligence.
1) They made incorrect claims about C++ in relation to Rust before: http://robert.ocallahan.org/2017/02/what-rust-can-do-that-ot... and http://robert.ocallahan.org/2017/04/rust-optimizations-that-...
2) The Mozilla C++ code base is old and very raw pointer/reference heavy and it doesn't seem to be written with safety in mind.
Their wikis also don't have any particularly good security guidelines. Maybe the good stuff is kept private, who knows.
Saying that someone was a distinguished engineer at Mozilla is not saying much about their abilities of writing modern or safe C++.
Yeah I made a mistake once. It happens.
If having a PhD in computer science (programming languages), being reasonably smart, and using the language for 20 years (up to and including most C++14 stuff) doesn't make you an expert in that language, then your language is far too difficult.
In fact, C++ is far too difficult and there are very few genuine experts in it. For example, who can explain why using push_back on a vector<map<T,unique_ptr>> is not conformant to the Standard, without looking it up? (I'll save you some time: https://bugs.chromium.org/p/chromium/issues/detail?id=683729...)
There's also a definitional bait-and-switch going on here. C++ proponents use "C++" to mean "the language that lots of projects have been using for 20 years and lots of programmers know" when espousing its popularity. But when necessary, the meaning changes to some "'modern', 'safe' subset of C++" ... that few programmers know well and few projects stick to rigorously. The exact definition of that subset changes depending on the situation, too.
C++ is difficult, and there are few experts. My thesis is that one doesn't need to be an expert to write safe C++ code, but they do need access to quality libraries focused on safety, and good practices focused on safety. Banning some unsafe C functions, saying "use smart pointers" or making a list of UB is useful, but not enough.
C++ can be written much more safely than it normally is, but it seems that's not happening. I'm not sure why, it could be that the performance loss of additional runtime verification is not acceptable, that the adequate learning resources are not available or that it's not an important topic for the C++ community.
P.S: I'll gladly have the kind of error you linked to. It's at compile time, I will try to figure it out and worst case rewrite my code. UB is the problem.
http://robert.ocallahan.org/2017/09/rr-50-released.html http://robert.ocallahan.org/2017/07/confession-of-cc-program...
The interesting question is not whether or not the author is a Rust proponent; it's whether or not Rust is an improvement on C. As neither a C nor a Rust programmer, it appears to me that the answer is unequivocally 'yes' — and this despite the fact that Rust is roughly as intelligible as Mandarin to me.
The bugs in dnsmasq were buffer overflows. In this case, the good practices would be always use std::array and std::vector and index with "at".
Tackling UAF is more complex. It involves using smart pointers exclusively with a runtime-check on dereferencing. This will result in some performance loss.
P.S: I'm not saying it trivial to secure C++ code, nor that every project out there is using these techniques. It's a worthwile task to try to make existing C++ projects as safe as possible, and that's an effort parallel to Rust which should have very beneficial results.
Ah, I see the Rust evangelism strike force is back. n-gate.com is right, as usual.