> However, suppose we flush our cache before executing the code, and arrange a, b, c, and d so that v is zero. Now, the speculative load in the third cycle:
> v, y_ = u+d, user_mem[x_]
> will read from either address 0x000 or address 0x100 depending on the eighth bit of the result of the illegal read. Because v is zero, the results of the speculative instructions will be discarded, and execution will continue. If we time a subsequent access to one of those addresses, we can determine which address is in the cache. Congratulations: you’ve just read a single bit from the kernel’s address space!
To my understanding it is that saying that by...
1) ...flushing the cache so you have a 'clean' state, you can get...
2) ...the speculative execution to 'pull in' to cache the address user_mem[x_] but...
3) ...the particular address that's pulled into cache, 0x000 or 0x100, is determined by whether...
4) ...the illegal read of kern_mem[address] 8th bit was a 1 or 0...
5) ...which you can then subsequently determine the value of by...
6) ...timing how long it takes to access that user_mem[x] address once again and...
7) ...thereby leaking the value of kern_mem[address]...
So you still have to perform some logic on the result of the speed of the access to the secondary address read right?
If read of 0x000 is slow you know kern_mem[address] was a 1 and if fast kern_mem[address] a 0, and if 0x100 is slow you know kern_mem[address] was a 0 and if fast that kern_mem[address] was a 1?
Is that correct?
If it is it seems that timing is the key right, and actually the clever leap of creativity in completing the exploit, at least to my untrained mind.
Please do correct anything I've got wrong, I'm not an engineer/developer!
Consider an Olympic 100 metre sprinter. Today we time this event very accurately, I think it's to one hundredth of a second, using sophisticated technology.
But even if the judges used a much less accurate mechanical stopwatch, Usain Bolt wouldn't actually be slower, we'd just be less confident of how ridiculously fast he is.
In some special cases, timing things very accurately might be essential to a use of Javascript, but I can't think of any examples off the top of my head.
What does that do, besides turn the exfiltration problem from an immediate one into a statistical one?
Maybe there will end up being a new Jumping Around Kernal Address Space System (JAKASS - a cousin of Linux's FUKWIT patch) that periodically resets kernel ASRL to make it fully impossible.
The paper they linked to references this one: https://www.usenix.org/system/files/conference/usenixsecurit...
I think this is what all sandboxes have to do: set the TSC disable flag, restrict system timer precision (make it configurable per sandbox: web servers generally don't need more than 1ms precision), make system timer report fuzzy (randomized) time. Heck, why not also make the CPU run at randomized frequency to mess with busy loop timers.
It's also quite fun to think that the little Pi I have chugging away in a tiny corner doing a variety of background tasks, which was already the most trouble-free machine I own, may also be the safest (OK I know that's an oversimplification, but I'm feeling affectionate towards it).
Instead, the read-ahead/speculative logic causes one of two addresses in user space to be read, and thus placed in the cache. So, by reading both of them, and checking the time it took, the exploit can indirectly determine one bit (0 or 1) of kernel memory. Scary!