Things I learned from OpenSSH about reading very sensitive files
utcc.utoronto.ca
utcc.utoronto.ca
There's no inherent problem with the approach taken (not like OpenSSL reinventing several wheels and causing things like Heartbleed). In fact, the article's conclusion that the OpenSSH team should have rolled their own memory management code rather than use tried, tested and approved libraries doesn't sound like a good recommendation to me.
That's not the conclusion I draw from the article.
I rather read it as warning that even innocent and quite common operations like realloc() or using buffered I/O can have security implications, so one need to be extra extra extra careful when writing code that handles sensitive data. In other words, security can be even harder than security professionals think.
Let ssh spawn an ephemeral ssh-agent just for the duration of the key-exchange and just for this single ssh instance. Kill it once authentication is complete.
I agree with your general statements but disagree with your ending conclusion. Notice that they never mentioned `malloc` in that post, nor `read` and `write`, etc. If you're writing code that needs to be secure, then it's important to be able to recognize what features you can use, which you can't, and thus what pieces of the standard-library are usable.
`malloc` is perfectly fine, because it will never touch the memory you put inside it - If you clear the memory with a `memset` before `free` (And make sure the `memset` isn't removed), then `malloc` and `free` will never be able to copy your memory or leave it laying around. `realloc` breaks this guarantee because it explicitly will make a copy if it has to find a new place to put the buffer, and thus it's not usable if you need to keep the above guarantees about copies of the data.
`fread` and `fwrite` are obviously out for the same reason - There's no guarantee they'll clear their buffers. For such a situation you definitely want unbuffered IO, and the best way to do that is to just directly do `read` and `write`. You could write your own buffer layer on top of it (that is securely cleared), but it's not necessary to use buffers at all - Especially if the file is small enough that you could just read the entire thing into a single buffer anyway. (Note that it may be possible to get those guarantees through some `setvbuf` and similar calls - I'm not completely sure since I've never tried - It's not the defaults regardless).
The big problem with the approach taken is that it doesn't take enough precautions when handling secure data, and uses functions that explicitly do things they don't want.
This is a problem in direct proportion to how leaky the abstractions are. C's abstraction over memory is comically leaky, so indeed, if you're using it, you need to think about the details underneath it.
But the solution isn't for everyone to forget about abstractions, it's to use better-made abstractions.
For memory in particular, it's been understood what it takes to have a safe abstraction for quite a long time now, and numerous widely-available languages provide one. Use one of those instead of C.
As far as I can see, if something like Heartbleed exists, you can't rely on "safe memory abstractions" at a lower level, because Heartbleed let the attacker look at the internal memory of those abstractions. They quit being abstractions; their raw memory is open to the attacker.
Or is your claim that safe abstractions would have prevented Heartbleed (and all other possible attacks that let the attacker read memory)?
In C, the memory abstraction the language gives you has that leak, so it's really hard to build a container on top of it that doesn't have that leak.
In safer languages, the memory abstraction doesn't have that leak, because bounds are checked, and new byte buffers are initialised to zero. You can still build an abstraction which does have that leak - say, by writing your own reusable byte buffer class which doesn't check bounds or zero on reuse - but at least you have the option of getting it right.
In an even safer language, the type system might make it impossible to make that mistake, or the library might supply a safe implementation of that abstraction already.
Remember, Heartbleed is not about the attacker doing some kind of Jedi memory trick to read your secrets: they are relying on your own code to read things it shouldn't. If the abstractions your program is built on don't allow that, then there's no way for it to be used to that end.
>In safer languages, the memory abstraction doesn't have that leak, because bounds are checked, and new byte buffers are initialised to zero.
Most of the implementations of these safer languages are in C, so clearly there's a way to do this correctly in C.
Is it not the implementation of those abstractions that determines whether they are safe? Offhand, I don't recall seeing guarantees of safety in this regard being offered by any language I am familiar with, though I have to admit that it has not been something I have been looking for.
The abstraction i was thinking of was byte arrays, which are provided by the language, rather than anything higher-level. Many languages have byte arrays which guarantee that you can't read past the end, and can't read values that you didn't write. Java would be one example.
That advice doesn't really help the OpenSSH team at all. OpenSSH is a C project and it certainly is not going to be rewritten in a safer language. Their are many reasons for that, first reimplementing OpenSSH would take lot's of effort, and the OpenSSH(OpenBSD) people don't have much man power to spare. Secondly OpenBSD is a C shop for better or worse, like most open source operating systems. Thirdly what language can fulfill all the constraints of OpenSSH and be memory safe? Let me try and explain, OpenSSH is included in many operating systems, like OpenBSD, FreeBSD, most of the linux distros, etc. Most of these groups are not going to want to include languages that require big complicated runtimes into their base operating systems. So that get's rid of any JVM based language, erlang, and any other non compiled language. So that leaves languages like Go or rust, which most certainly are safer languages than C, but can they run on all the platforms listed on OpenSSH portable?[1] And just as importantly will Operating systems add another programming language to their base system, which increases the effort required for maintenance just so they can switch the implementation of a working program. I personally don't think most major operating system would make the switch. Finally OpenSSH has a pretty good track record when it comes to security, especially considering how little funding they receive, it would probably require less effort to continue improving OpenSSH then to ditch it and start over and get everyone to use the new version.
You're right that languages with big runtimes would probably not be worth the cost, and new languages like Rust and Go are still too risky. But how about something like Ada or Modula 2? Both are well-established (virtually ancient) languages with mature compilers, small runtimes, and proven use in systems development.
I thought about Ada too, but the one thing that gave me pause is the shortage of Ada programmers relative to more mainstream languages.
What's the difference between GCC deleting parts of your code, and an attacker hacking into the source code repository and deleting those parts of the code?
The C standards committee addressed the problem in 2009 with memset_s. [0] The GNU developers reject patches and state they hope this feature is never implemented. [1]
The bug fix is to stop using GCC for sensitive code. Use CompCert instead. https://github.com/AbsInt/CompCert
[0] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1381.pdf
[1] https://sourceware.org/ml/libc-alpha/2014-12/threads.html#00...
License
CompCert is not free software. This non-commercial release can only be used for evaluation, research, educational and personal purposes. A commercial version of CompCert, without this restriction and with professional support, can be purchased from AbsInt. See the file LICENSE for more information.
I'm not entirely sure if something like that is good for an unpaid-for open source project...
I moved to yubikey some time ago and don't regret it.
(Also, this still doesn't provide the key itself. That means even with the vulnerable yubikey you can only spoof connections while it's plugged in)
Yubikeys have had bugs critical enough to warrant replacements in the past, and have yet to undergo any external audit.
Proving code correct is far from obvious though. Apart from single instances like sel and python's sort, can you remember any verified software with proofs right now?
seL4 is actually another verified software I know.
That is starting to ring hollow to me.
We can not just trust that these libraries are higher level esoteric magic that no mortal could understand. It is time to shine light on exactly what guarantees different libraries are providing, and how.
You're being pointed straight at the problems discovered by other people. Of course they're obvious to you now. The question is, can you just pick up some OpenSSH code, read it and run it, and find a new problem yourself when nobody is pointing you straight at it? Because I'm sure there's at least one in there for you to find.
(Note I did not ask you if you could find this problem. Too easy to imagine that you can now, too hard to realistically pretend you don't already know about it. I'm asking you about new problems that nobody currently knows about.)
a) I would not have thought of the issue of using a standard library call to load data. b) nobody would have checked my code and fixed it
so in the end, I'd still be vulnerable while openssh is now fixed.
In any case more organization than the number of orgs that would bother to check my code.
The sentiment behind that is to keep out those that don't know how to code. If everybody followed that advice there wouldn't be crypto code.
If you have the ability to create good code and are confident you can understand the maths of cryptography then by all means contribute to the community. If the past year has taught us anything it's that the crypto community needs more good engineers.
E.g. use VPN and password protection and encrypted volumes which are pretty standard. Add in a custom hardware switch that physically disconnects the power/network of a server containing sensitive data which is rarely needed and you will be protected against many exploits.
An easily exploitable buffer overflow or an extremely easily overlooked crypto blunder in a second layer can easily allow a way in.
I agree with the example given, but it's a rule that's wrong more often than it's right.
But if all the lessons the author took are that secure code should be lower level than this, well... I can not agree with it.
Remember RowHammer? https://github.com/IAIK/rowhammerjs
I don't think rowhammer lets you read memory, though; does it? There are certainly some cache timing attacks that do, though.
- Your input parser can compromise many more things than what you'll think about erasing.
- You won't be able to clean all secretive data anyway. You can't win here.
- You can be sure to a huge degree that your input parsing code won't get outside of its bounds.