OpenSSL CVE-2016-0799: heap corruption via BIO_printf
guidovranken.wordpress.com
guidovranken.wordpress.com
We've had some cargo culted patches to call _mesa_error_no_memory() when malloc fails, but it was recently noted that this can't possibly work because it internally calls fprintf, and there's no guarantee that fprintf doesn't call malloc! In fact, glibc's implementation does.
Suggestions for best practices welcome.
Of course, it would be better to avoid getting into a situation where either a graceful return isn't possible, or poorly-tested recovery code has to be run to deal with the failure. I read somewhere that seL4's strategy is to divide the work in two phases. The first phase allocates all the resources it will need and checks all preconditions, but changes no state; the second phase has no allocations, does all the work, and will never fail. That way, any error recovery is a simple release of the resources it had already allocated.
You can also use that wrapping technique to test _mesa_error_no_memory(), by setting the state to "all mallocs fail" and then calling it. Hint: Don't call fprintf, don't use stdio, call write() directly.
Fault injection can't cover all of the combinations of all inputs and all places to fail, but it's a lot better than nothing.
One (naughty) option is to malloc a decent sized buffer at startup (50MB say), and then when malloc fails free it, and immediately display a warning box.
One most modern 64-bit systems, malloc will never fail anyway, you'll get killed due to memory overcommit.
An abort with a reason is a far better choice.
Library code should in general punt out-of-memory conditions back to the caller.
It is if you ever want to exercise the code paths. Overcommit means you never get an out of memory error. Your process just dies. Kind of like it would if you just aborted.
...but that doesn't even really matter, because in order to fully exercise those code paths you need precise control over exactly which call to malloc() returns NULL, which in practice means you need to interpose malloc() with a debugging framework to force a NULL return.
Except if a virtual memory resource limit is set ("help ulimit").
$ sh -c 'ulimit -S -v 1000 ; perl -e 1'
Segmentation fault
$ sh -c 'ulimit -S -v 2000 ; perl -e 1'
perl: error while loading shared libraries: libm.so.6: failed to map segment from shared object: Cannot allocate memory
$ sh -c 'ulimit -S -v 20000 ; perl -e "qw{a}x100000000"'
Out of memory!wouldn't a simple 'ulimit -m' set appropriate limits on memory for the process ? and then use that to simulate failures, and hopefully see what breaks, fix it, and rinse-lather-repeat ?
edit-1: this is ofcourse predicated on the fact that you are using a unix like system.
Yes, as Hoare stated in 1981, regarding Algol design:
"The first principle was security: The principle that every syntactically incorrect program should be rejected by the compiler and that every syntactically correct program should give a result or an error message that was predictable and comprehensible in terms of the source language program itself. Thus no core dumps should ever be necessary. It was logically impossible for any source language program to cause the computer to run wild, either at compile time or at run time. A consequence of this principle is that every occurrence of every subscript of every subscripted variable was on every occasion checked at run time against both the upper and the lower declared bounds of the array. Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to - they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
Note: Not to mention Unisys still makes tons of money off their ALGOL machines from Burroughs. IBM & Fujitsu still service PL/1. So, old ones too.
And at least in C++, thanks to the STL or similar library there is the option to always enable bounds checking.
So even if the market didn't got hold of Algol, except for C and languages that offer copy-paste compatibility with it, all other ones do offer sane defaults in terms of security.
I've always liked that part. Once they saw the problems, they'd put up with whatever sacrifices it took to avoid them.
In an ideal world, someone who did that should be taken out back and shot. In this real world, someone who did that should be immediately walked out the door.
It's astonishing to me just how lazy programmers can be.
To be clear, in systems designed with the "worse is better" approach, responsibility for correctness can be pushed up the stack, from e.g. kernel programmers to user-space programmers, if doing the Right Thing lower in the stack is hard. The example given in the paper is a syscall being interrupted -- in unix, it just fails with EINTR, and it's the application author's responsibility to retry. (I agree that malloc returning null is similar.)
But worse is better isn't saying that programs shouldn't be correct if it's complicated; it's just a question of where the responsibility for correctly handling API edge cases should lie.
Moreover, if one reads the "worse is better" paper carefully, you'll notice he doesn't say that "worse is better" actually is better, overall. The argument is only that it has better "survival characteristics". Meaning, as I read it, that they can gain and maintain momentum against competitors because they can move faster.
Which, tangentially, is all very interesting to me, because the three modern systems with the best survival characteristics are unix, C, and the web, the latter of which is definitely not worse is better. (However bad you may think the web is as an API, "worse is better" is a design philosophy about giving users simple, low-level APIs, whereas the web is all about high-level APIs where the browser implements lots of logic to prevent you from shooting yourself in the foot.) So maybe we need to reconsider the conclusions of the worse is better paper in light of the giant counterexample that is the web.
Modula-3 or maybe Eiffel would be closer to Right Thing examples in that category.
The right answer is, I think:
* Don't add checking code at individual malloc call-sites.
* If you have an allocation regime where you need to do something better than abort in response to failure, don't use malloc directly for those allocations.
* Run your program with malloc rigged to blow up if it fails.
With links to git commits:
CVE-2016-0705: https://git.openssl.org/?p=openssl.git;a=commit;h=ab4a81f69e...
CVE-2016-0798: https://git.openssl.org/?p=openssl.git;a=commit;h=59a908f1e8...
CVE-2016-0799: https://git.openssl.org/?p=openssl.git;a=commit;h=a801bf2638...
However, they're scheduled for release on Monday the 1st of March - so we might be seeing a few more bugs in the next couple of days.
EDIT: I've spent all morning teasing a friend whose birthday is today about "what if it were tomorrow" (I'm sure they get it every year from every single person they meet), only to then come here and forget all about it in my tally. Go figure. Yes, the 1st of March is Tuesday. Sorry!
Clarification from OpenSSL maintainer.
_dopr(&hugebufp, .... _dopr(&hugebufp, ...Outside of every single thread about C. Everyone at this point knows the shitty side of C.
IMHO dynamic allocation is something that is often overused, and the fact that it could fail leads to another error path that might be exploitable, like this one. I don't think it's necessary for a printf-like function to ever use dynamic allocation. (If there are cases where it's necessary, I would love to know more... and perhaps suggest how it could be avoided.)
I do wish them the best luck, though.
Also, I'm pretty sure that Oberon was quite a bit earlier than that.
I mean theoretically the issue still exists, because you can run allocation yourself and get back null. But without unsafe this shouldn't be possible.
Rust's standard library aborts on OOM.
- implicit type conversions errors
- memory allocations with incorrect size
- implicit decay of arrays into pointers
- lack of bounds checking for arrays
- null terminated strings with missing null character
- invalid pointers passed as argument for out parameters
- implicit conversion of integers into invalid enumeration values
Obviously not. We've been doing this for decades and are still making the same mistakes with costly effects. Humans are as much a part of technological systems as the language or platform; continuing to ignore that even educated, disciplined, skilled developers are prone to certain classes of errors -- errors that could be obviated by other components of the system -- is hubris. It is much more feasible to create tools that exclude certain classes of error than to create less fallible developers.
> (code size; execution speed; random latencies)
Rust is not GC'ed, has no inherent reason to be slower than C (as the toolchain matures), and I doubt is more verbose than C code that provides the same safety.
And nothing to say about the performance cost of these infallible tools? They are not free.
And now we have power saws that immediately stop when they encounter flesh. Obviously, disciplined users never had to worry about that system, otherwise we would have had a bunch of people missing fingers... like my wood shop teacher in high school. :/
> And nothing to say about the performance cost of these infallible tools? They are not free.
People have brought this up a few times, but I've never gotten an explanation that makes sense. My understanding is that Rust does most of its protections in the compile stage, so while not free, they are a small up-front cost at the compile stage, plus an indeterminate cost at the development stage in possible more complex reasoning about the code, which then pays off over the lifetime of the application. While not free, that's definitely a trade-off that seems to make sense to me.
Or are you talking about some other cost? Can you explain where I'm not understanding your point?
I agree we could do better than C. Everyone here talks about it every time this happens, but does nothing. I'll take battle-hardened C over written-last-week Rust, too. Especially crypto code.
Apparently 40 years are not enough to acquire that level of discipline.
We've enjoyed orders of magnitude performance and latency improvement over scripted or GC'd languages.
Its certainly possible to write competent C++ with a professional team, and clear reasons for doing so.
C++ does provide the necessary features to write safe code, as long as one stays away from "C with C++ compiler" mentality and adopts those language and library features.
As for the tools, they are largely ignored by the majority of enterprise developers, assuming Herb Sutter's presentation at CppCon 2015 is in any way representative (1% of the audience said they were using them).
Which goes in line with my (now) outdated C++ professional experience up to 2006.
Only one company cared to pay for Insure++ (in 2000) and I was probably the only one that cared to have it installed.
Also except for my time at CERN, I never met many top elite C++ developers.
Most were using it as "C with C++ compiler" and in teams with high attrition, leading to situations where no one was 100% sure what the entire codebase was doing.