Auditing Rust Crypto: The First Hours
research.kudelskisecurity.com
research.kudelskisecurity.com
A zeroing Drop is theater unless the sensitive data is Pinned too. Rust likes to copy things around.
Of course not, but for efficiency you might have an [u8;32][0] instead and that will get memcpy'd when it's moved around.
[0] 32 bytes on the stack isn't much, a single Vec or String is 24
> Since Rust itself has no notion of immovable types, and will consider moves to always be safe, this trait cannot prevent types from moving by itself.
From the documentation on std::pin:
> In order to prevent objects from moving, they must be pinned by wrapping a pointer to the data in the Pin type. A pointer wrapped in a Pin is otherwise equivalent to its normal version, e.g. Pin<Box<T>> and Box<T> work the same way except that the first is pinning the value of T in place.
> First of all, these are pointer types because pinned data mustn't be passed around by value (that would change its location in memory). …
Looking at Pin itself, it appears that it works by only implementing DerefMut if its target is Unpin, thus preventing you from mutating (and therefore swapping/replacing) the data unless the data supports Unpin. But beyond that, it also only implements Deref and DerefMut if its wrapped type itself implements Deref/DerefMut, which means using Pin around a non-pointer basically prevents you from accessing the wrapped data.
Crypto Implementation Is Hard. Especially when you need it to be secret. It is lots easier if you don't care.
In other words, why is a zeroing Drop of a type more secure if it uses Pin too?
Afaict, Rust memcopys moved objects if they're larger than some size when passed as function arguments. Also take into account implicit copies such as those left laying around when a vector outgrows it's heap segment.
[0] though I guess the compiler might be smart enough to implicitly pass big-enough values by pointer instead
Optimizers consider code that zeroes a dying object to be dead code. (This is also a problem in C++ destructors.) It can be arbitrarily difficult to keep such code from being elided in various passes. Even inline asm instructions can be removed by a peephole pass.
Then, the CPU might decide not to flush all those zeroes from its writeback cache until some random time later when it needs that bit of cache for something else.
This is all wonderful for performance, usually, but it means you it takes a lot of dog-work to disable it.
If you have not verified that memory actually has zeroes in it, as seen by another core or DMA engine, it doesn't.
"Transactional Memory" seems to have flopped for lock-free synchronization, but it is a nice way to keep writes from ever going out to RAM or beyond. It reserves a bit of cache for temporary storage, and automatically wipes it if anything happens, like an interrupt -- even on a guest OS.
That is better than zeroing, and the toolchain won't fight you. But it might not be available everywhere you want it. Also, it wasn't designed for crypto, so there are probably vulnerabilities to the host OS.
Everything around security is hard. It gets easier if you don't care whether it really is secure, which is common.
OpenBSD provides
void explicit_bzero(void *s, size_t n);
In libc, which is used for this purpose. Recently, glibc added function, presumably because it's a good idea.But it still has limitations due to lack of compiler support. I think the problem can only be solved at the compiler/OS level.
> The explicit_bzero() function addresses a problem that security-conscious applications may run into when using bzero(): if the compiler can deduce that the location to zeroed will never again be touched by a correct program, then it may remove the bzero() call altogether. This is a problem if the intent of the bzero() call was to erase sensitive data (e.g., passwords) to prevent the possibility that the data was leaked by an incorrect or compromised program. Calls to explicit_bzero() are never optimized away by the compiler.
> The explicit_bzero() function does not solve all problems associated with erasing sensitive data:
> 1. The explicit_bzero() function does not guarantee that sensitive data is completely erased from memory. (The same is true of bzero().) For example, there may be copies of the sensitive data in a register and in "scratch" stack areas. The explicit_bzero() function is not aware of these copies, and can't erase them.
> 2. In some circumstances, explicit_bzero() can decrease security. If the compiler determined that the variable containing the sensitive data could be optimized to be stored in a register (because it is small enough to fit in a register, and no operation other than the explicit_bzero() call would need to take the address of the variable), then the explicit_bzero() call will force the data to be copied from the register to a location in RAM that is then immediately erased (while the copy in the register remains unaffected). The problem here is that data in RAM is more likely to be exposed by a bug than data in a register, and thus the explicit_bzero() call creates a brief time window where the sensitive data is more vulnerable than it would otherwise have been if no attempt had been made to erase the data.
There's the problem. That is a lot of levels. There is support at certain levels; nowadays you can force writebacks from certain cache lines, and OSes let you keep pages from being swapped. The problem is that it needs coordinated support at all levels, but there is no money behind it, unlike features that may improve performance, and hardly anybody knows there is even a problem.
Spectre and Meltdown have raised awareness of security threats in infrastructure, but are so big they also absorb all the budget for it, and will continue indefinitely.
I've once read something about Remote Direct Memory Access, it allows one to completely bypass the kernel, CPU, cache, and stream data directly to the network adapter of another computer - you break a lot of essential abstractions of modern computer system, meanwhile still be able to provide a consistent and usable infrastructure.
Its security equivalent (something allows low-level control of the protected data) is unlikely to happen, except for DRM, I guess.
One way to combat these side channels is to use constant time multiplication, loop conditions (including functions that might return early, like memcmp), etc.
Is it possible for Rust to have any new or surprising mechanisms that cause loss of constant time execution?
(Of course also other aspects can leak than the CPU instruction execution latency itself. For example cache effects might be measurable.)
See for example this page: https://bearssl.org/ctmul.html
> could you elaborate on how it applied to symmetric encryption?
Non-constant time multiplication is mostly (only?) an issue with asymmetric crypto.
Similar attacks occur when your encryption uses lookup tables depending on key bits, uses different amounts of power, etc.
It's a surprisingly effective attack surface, and why the industry is moving away from older style schemes like AES and towards ARX based constructions like ChaCha.
I believe it's certainly possible. And the precise consequences of writing this code in Rust have yet to be fully explored. The `ring` (Rust port/fork of BoringSSL) library still links to constant-time primitives implemented in C.
The ubiquitous assumption that all code should be executed as fast as possible isn't new or Rust-specific but it will be a continuing source of security surprises for the foreseeable future.
This isn't an exploitable vulnerability in Rust, though? I thought the program would just abort if it overflowed the stack.
let huge_ary: [u8; 2097152 /* 2MiB */] = unsafe { ::std::mem::uninitialized() };
// hopefully huge_ary has bumped the stack pointer past any guard pages
// and any further stack values will overwrite the heap let huge_ary: [u8; 2097152];
if false_value_compiler_cant_prove_is_false {
huge_ary = [0; 2097152];
}
// hopefully huge_ary has reserved space on the stack despite not being initialized
// and so any further stack values would overwrite the heapThen again, 2Mb is nothing on a 64 bit processor, so I don't think the chanse of hitting the heap is very big.
The 2Mib was assuming you've already pushed the stack to its limit and are sitting just above the guard page(s). You could replace that with whatever size you think is appropriate.
error[E0381]: borrow of possibly uninitialized variable: `huge_ary`
--> src/main.rs:6:7
|
6 | bar(&huge_ary);
| ^^^^^^^^^ use of possibly uninitialized `huge_ary`
https://play.rust-lang.org/?version=nightly&mode=debug&editi...What is the preferred way in Rust to say "If at all possible, please do not allow this value/object/whatever to be paged out of RAM?"
For instance if you're running in a VM, all bets are off, beacuse the host can and will page you out guest phys memory.
As you note, you can't just mlock memory handed out by the ordinary allocator, as you never know when it will be free to munlock.
In rust, there's the zeroize crate [1], which uses compiler fences and volitile memory to have a shot at it, but it's still not a sure thing.
That library doesn't help with the paging issue you mention, but what you can do for that is entirely platform dependent (up to the kernel's paging mechanism), and so rust isn't going to make it any easier or harder than just making the libc calls/syscalls yourself.
[0]: http://www.daemonology.net/blog/2014-09-06-zeroing-buffers-i...
[1]:https://github.com/iqlusioninc/crates/tree/zeroize/v0.5.2/ze...
If you write in, say, assembler, you will have precise control in zeroing the buffers and not spilling sensitive register values to the stack, and also zero out XMM registers.
In fact that the very article you linked to outlined a solution in the form of a proposed language extension disagrees with your characterization that this problem cannot be solved in any languages. Clearly, the author thinks that it's still possible to zero buffers correctly in a language, but just that C is not such a language unless it gains his proposed extension.
It has to be solved at all levels; each level gets a chance to sabotage it, and will.
Not really, because micro-coded CPUs have won and you don't have any control over what they do, and as last year has proven there are ways to exploit execution timing in micro-code.
As mlock / VirtualProtect are OS dependent I strongly suspect this is not part of Rust itself. Additionally, I suspect there would be some design problems for example putting this protected thing on the stack (protect the whole stack? Or just parts? Bookkeeping wether a stack lock can be released after a function exits?). How do you move values into / out of such an object?
So I don't think this exists, but I've been pleasantly surprised by Rust before.
My general solution for this case is to either mlockall the whole process if it is small / short-living or to have no swap in the machine, for example long running crypto servers. Who runs servers with swap anyways?
For instance see Raymond Chen's explanation here
> Is VirtualLock safe to use for cryptographic purposes – that is, to avoid sensitive memory being written to disk?
> ...
> [There does not appear to be any guarantee that the memory won't be written to disk while locked. As you noted, the machine may be hibernated, or it may be running in a VM that gets snapshotted. -Raymond]
VirtualLock and the equivalents only guarantee that the canonical copy of that page stays in RAM, not that other copies won't make their way out to disk.
Edit: add link https://blogs.msdn.microsoft.com/oldnewthing/20140207-00/?p=...
As for where it comes up, it's usually because you're handing ownership of the memory to something outside of Rust's normal memory management. In most cases you'll be using a higher-level API like Box::into_raw(), but if you look at the implementation of that you'll see it uses mem::forget() internally.
With the precondition you're running on bare metal (no virtualization), I think you'd need a kernel mode driver for that to cover all scenarios, including system hibernation and suspend.
You can reliably zero out the secret data before hibernation or even suspend by implementing power management calls – mlock or VirtualProtect aren't going to protect you from that leak.
But PBT_APMSUSPEND is not necessarily. There's a 2 second timeout before Windows suspends or hibernates regardless. It's possible for the event handling thread not to run until after the timeout has passed, that is, after the system has returned from suspend or hibernate.
That's important considering the adversary might induce this kind of scenario to steal the secret.
Kernel mode driver does not have that limitation, because it can't be pre-empted.
See mlock(2) for details [1].
> Since Linux 2.6.9, no limits are placed on the amount of memory that a privileged process can lock and the RLIMIT_MEMLOCK soft resource limit instead defines a limit on how much memory an unprivileged process may lock.
I had the simple task[2] of encrypting bytes with AES using the library. But (as detailed in the post), they didn't expose a simple, "give us plaintext bytes, give us key bytes, give us options, we'll spit back the ciphertext bytes". Instead, you have to delve into the irrelevant details of constructing mutable buffers (which require special care in Rust), so it can read from and write to them, and without even any good examples in the docs for how to construct those buffers.
It seems designed from the perspective of "I have extremely limited memory" -- which is fine, but why not have a simple wrapper on top of that for the common, non-embedded use case, and some example code?
[1] https://news.ycombinator.com/item?id=18943940
[2] for the Matasano/cryptopals challenge. Interestingly, I found that the hardest parts weren't figuring out the vulnerabilities, but in implementing a spec (every writeup was poor) or getting "off-the-shelf" software to behave as expected so I can have a utility function for it and forget about it from then on!
https://en.wikipedia.org/wiki/Mersenne_Twister#Pseudocode
(I use the plural because I'm lumping that in with the difficulty of getting openssl to yield the expected results for AES encryption.)
The `ring` library is a bit more modern, but still requires a bit of a run around. Here's an example of encrypting a cookie: https://github.com/alexcrichton/cookie-rs/blob/master/src/se...
I'm reminded of the Onion article about "Stating current year still the leading argument for reform" [2] (e.g. "It's 2019, people!")
I guess the corollary is, "That was 3 years ago" still leading excuse for shoddy work :-p
[1] https://news.ycombinator.com/item?id=18943056
[2] https://www.theonion.com/report-stating-current-year-still-l...
When your immediate reply is to dismiss the dominant rust crypto library in 2016 as being understandably bad because the development ended in 2016, you're suggesting that the year (and before) was some kind of dark age for rust.
>Static, perfect code is rare.
My criticism isn't "hey, there's a lingering bug that no one fixed and which broke my use case". My criticism is, "The design of the exposed interface to the core library functionality forces me to care about irrelevant implementation details and enable attributes (mutability) that my use case doesn't need."
When you say that a project developed in 2016 can be expected to fail on that front, what you're saying is, 2016 was too early to expect developers of a major project (in a hot language followed by a lot of smart people) to understand the concept of separation of concerns and abstraction of irrelevant details.
Do you see why I might be skeptical of the relevance of the year of development?
There's a similar point to be made about (the lack of) clear examples, a concept that didn't spring into existence in 2017.
I was there and it actually was.
Rust 1.0 was just released and the ecosystem was mostly maturing at that point.
You're talking about version 0.2.36 of a library that had been in development for less than two years during a quite tempestuous time in Rust. I'm sure they were more concerned with making the code work, making it secure, etc.
Lack of clear examples is not an edge case bug, it's failing to think about your user.
Breaking separation of concerns is not an edge-case bug.
I'm old, so I can assure you that, as of 2016, developers were well aware of the concepts of separation of concerns and abstraction. It being 2016 does not address the criticism I made. See also my longer reply.
I don't doubt that 2016 had some great libraries in rust, of which a lot are still used today. But most of them are updated to use new language features like `?` instead of `try!`, new macros system, async libraries, etc..
Take a look at a library like `serde`, which has been around since 2015 (https://crates.io/crates/serde/versions) and is awesome.
That is a surprisingly uncommon API for a crypto library. It's far more common that they make you do a lot of the setup that really should be under the hood and that if you get wrong your security will be ruined. Plus the docs won't explain what any of the values mean because obviously you wouldn't be using crypto unless you already took a college course in it.
It's far too rare that the docs even explain what an IV is, or what you should set it to, or if it needs to be kept secret, or even what the initials mean. I think it is on purpose to scare off developers that have not studied the field extensively, but in reality they have a job to do so they set it to 0 and then ship 84 million appliances where this code is baked into the firmware update system.
Why do you think rust developers weren’t aware of separation of concerns or good examples in 2016?
The author seems to have forgotten about Cargo.lock....
I consider a library insisting on locking it's second level dependencies for no good reason to be obnoxious behaviour. As producer of an app, I want the ability to upgrade your dependencies to patch security problems.
In my experience, excessive locking can make security problems worse not better.
Also I'm not sure you can get any constant time guarantees today without calling into C or asm.