V8: A Year with Spectre
v8.dev
v8.dev
We should all be more angry at chip makers for this. Intel isn't willing to admit that they broke things because then they'd have to fix it, but we shouldn't accept that kind of approach.
the crazy thing is that nobody saw this until recently.
Mere causation can’t get you there. E.g. when a car hits a pedestrian, the driver and pedestrian equally “caused” the accident. It is only by way of characterizing their behaviors in one of the ways above that we can identify wrongdoing. Perhaps the driver wasn’t paying attention (negligent or reckless) and ran a red light. Or perhaps the pedestrian was intentionally throwing themselves in front of traffic. Etc.
[0] Products liability law on its surface does eschew the moral-wrongdoing requirement in favor of strict liability for some kinds of product defects. But that has to do with economic incentives, practical ability to prove claims, etc.
They could of course just reject the necessity of need Y, but if the majority of their clients actually do have need Y, can it really be said that the chip is successful at being general purpose?
Correct me if I'm wrong, but speculative execution attacks (or at least the possibility) were known for several years before Spectre.
https://randomascii.wordpress.com/2018/01/07/finding-a-cpu-d...
chip consumers missing a communication is very different from intel actively developing this to cut corners for raw performance (which is the only reason they cornered the market) and forcing all other manufacturers to follow up or die.
And this feature was used to support not only many programs running at once, but also TSO (Time Sharing Option) which by definition is mutually untrusted code.
ARM for instance had gone so far as to make JS specific instructions (FJCVTZS, Floating-point Javascript Convert to Signed fixed-point, rounding toward Zero), and before that they had instructions to help a JIT (ThumbEE).
Not sure why we're buying the chip companies' shtick of "oh, poor us, we never knew people would use our chips like that, we just specifically optimized for it and provided support instructions"
Hardware and software is codesigned, and yes the onus is on the chip manufacturers to release chips that let you continue to use them securely.
It's like all the whining people do about GCC doing unexpected things when faced with code that relies on undefined behavior. That's not GCC's fault.
The issue isn't that there's a bug in their VM implementation, it's that with current hardware general VMs and same process isolation are mutually exclusive.
This isn't like GCC, where the C standards bodies officially got together and said "don't do this".
Hardware moves slowly and can't be fixed easily. Browsers should just be doing this (which they are with site isolation) instead of hardware trying to hack in patches to cover up bad software architecture.
And you can't fix a bug you didn't know about anyway, so why would you have expected ARM to fix something nobody knew was real in the first place?
Prior to web browsers when was such a thing ever widespread? And if you eliminate web browsers from the picture, how many usages are even left?
> Java got big right when the first consumer OoOE chips came out in the 90s.
Java doesn't do this, so how is that relevant?
BPF has existed since the mid 90s for one example.
> Java doesn't do this, so how is that relevant?
Java and the client web were next to inseparable concepts at the time. Java ran their VM as a shared library in the browser process for applets.
And that is a different discussion. The article, and the discussion here, is regarding side channels attacks within a single process. I'm pretty sure everyone agrees that the hardware (or some conspiracy of the hardware and kernel) must provide process isolation.
Spectre is simply a subtle oversight in the way different pieces of the system interact.
[1]: https://www.hpl.hp.com/techreports/Compaq-DEC/WRL-87-2.pdf
> This access control mechanism does not in itself protect against malicious or erroneous processes attempting to divert packets; it only works when processes play by the rules. In the research environment for which the packet filter was developed, this has not been a problem, especially since there are many other ways to eavesdrop on an Ethernet.
And even on modern systems, you need to be root to install a packet filter, which is typically also sufficient privilege to simply open /dev/mem and read kernel memory directly. (Or you did, then people started using bpf for everything, but that came many years after the 386 and 486.)
I don't think BPF is a good example of running untrusted code, at least not the early versions, since it wasn't untrusted.
Yes in the sense that I assume chip makers didn't foresee it and don't like its ramifications, but it's also one that's essential to how the chips currently give their users the performance they want.
(There is something in Meltdown's case: not flushing the TLB on every context switch is relatively new. But it's directly encouraged by CPU manufacturers via the ASID feature. And Meltdown is a narrower bug anyway, something that's easy to fix in a new hardware design, not like Spectre which is more fundamental to the concept of speculative execution.)
Sometimes I think that humanity is not ready for computers, and it's time to go full-on Butlerian jihad.
> Our research reached the conclusion that, in principle, untrusted code can read a process’s entire address space using Spectre and side channels. Software mitigations reduce the effectiveness of many potential gadgets, but are not efficient or comprehensive. The only effective mitigation is to move sensitive data out of the process’s address space.
Probably the most effective hardware mitigation would be to shift the isolation to a page-level granularity, so that you could say that speculation is disabled for memory in specific pages.
Either way, my point was that it's absolutely possible to fully mitigate Spectre in software. A particular group of people just found it too hard to do for the things they work on. Doesn't mean the same applies to anything else.
And something like this already existed, 32bit chrome would play games with the LDT on OSes that allowed that to build a better sandbox. It's just that the LDT more or less disappeared on x86_64.
What I want is a new ring-4. Ring-4 is like ring-3, except that cache attacks are mitigated (disable speculative execution is the obvious option, but there might be something else that I don't know of). Ring-3 code can put parts of itself in ring-4 mode as required.
Why I say that bringing back rings one and two is a good idea is that it gives another context to user mode _and_ kernel mode, so that the kernel can protect itself from BPF style programs at some point.
Ring 1 and 2 already exists on x86, we can't just reassign what they do, which is why I introduced ring-4. Ring-1 and Ring-2 are too powerful for user programs, the web browser should not use them.
Failing that maybe it is time to bring back segments.
https://github.com/monocasa/remu-playground
Because you have to round trip through the host kernel to do anything that leaves that context, it's a bit slower and doesn't really get you the perf gains overall you might think. : /
But yeah, segments coming back would be really neat.
Intel also support memory protection keys. I wonder if they would be enough. Although they might be affected by meltdown like issues.
Does AMD support them?
This would outright prevent code generated by the JIT to do out of bounds speculative loads directly, it would just load a rewritten address from within the restricted region. Instead you'd have to trick an API call (which would be executed without the restriction) in to doing it for you.
We already ship H.264 decoders on everything. And vendors are sort of settling on, Chrome/WebKit rules all. Why not just standardize the model in the hardware?
On the other hand, image and video codecs are actually chock full of security issues nobody bothers to exploit.
What's wrong with the hardware mitigation of just using multiple processes? It already exists, it already works, and it already has decades of tooling & infrastructure built around it.
Processes? They don't need to be expensive, as proven by Linux. Trivially cheap enough to do one per site origin in a browser, anyway.
> Site isolation in Chrome explodes memory pressure on the system.
How do you figure? Code pages are shared, after all. Only duplicate heap would be an issue, but shared memory exists and can mitigate that if there's read-only data to be shared.
So what memory pressure is "exploded"?
> Language enforced confidentiality is a lot less resource intensive.
Not at all clear-cut or self-supporting. What resource(s) is it less intensive on, and what are you using to support such a claim?
CPU time is a resource, too, after all. All this software-injected mitigations and maskings aren't free.
Problems worth thinking about:
Sharing jitted code across processes (including runtime-shared things like standard library APIs) - lots of effort has gone into optimizing this for V8.
Startup time due to unsharable data. Again lots of effort goes into optimizing this.
Cost of page tables per process. (This is bad on every OS I know of even if it's cheaper on some OSes).
Cost of setting up process state like page tables (perhaps small, but still not free)
Cost of context switches. For browsers with aggressive process isolation this can be a lot.
Cost of fetching content from disk cache into per-process in-memory cache. This used to be very significant in Chrome, they did something recent to optimize it. We're talking 10-40 ms per request from context switches and RPC.
Most importantly the risk of having processes OOM killed is significant and goes up the more processes you have. This is especially bad on Android and iOS but can be an issue on Linux too.
ASLR and other security mitigations also mean you're touching some pages to do relocation at startup, aren't you? You're paying that for dozens of processes now.
As for the actual problems, many of those are very solvable. Startup time, for example, can be nearly entirely eliminated on OS's with fork() (and those that don't have a fork need to hurry up and get one) - a trick Android leverages heavily.
And a round-trip IPC is not 10-40ms, it's more like 20-50us ( https://chromium.googlesource.com/chromium/src/+/master/mojo... )
> Most importantly the risk of having processes OOM killed is significant and goes up the more processes you have. This is especially bad on Android
There's no significant risk here on Android. Bound services will have the same OOM adj as the binding process.
Yes, but various things that require relocations may not be. That can include code, but definitely includes data like C++ vtables, as a simple example. Just to put a number to this, for Firefox that is several megabytes per process for vtables, after some work aimed at reducing the number.
There are ways to deal with that by using embryo processes and forking (hence after relocations) instead of starting the process directly; you end up with slightly less effective ASLR, since the forking happens after ASLR.
> So what memory pressure is "exploded"?
Caches, say. Again as a concrete example, on Mac the OS font library (CoreText) has a multi-megabyte glyph cache that is per-process. Assuming you do your text painting directly in the web renderer process (which is an assumption that is getting revisited as a result of this problem), you now end up with multiple copies of this glyph cache. And since it's in a system library, you can't easily share it (even if you ignore the complication about it not being readonly data, since it's a cache).
Just to make the numbers clear, the number of distinct origins on a "typical" web page is in the dozens because of all the ads. So a 3MB per-process overhead corresponds to something like an extra 100MB of RAM usage...
Is that really due to multiple processes though? Isn't the overhead for a new process reasonably low, like <2MB?
Chromium already had a lot of this infrastructure built up, and it was still a significant task to split processes up further by origin, on top of the splitting by tab chromium always had.
https://en.wikipedia.org/wiki/Singularity_(operating_system)
I can't find a link at the moment, but I recall a paper showing even a very granular clock will suffice for Spectre exploits, albeit with lower bandwidth. Also, something else would need to be done about multithreading, as an application could always just spin up another thread counting as fast as it can to make a poor man's timer.
The linked article mentions this and the linked paper gives some references: https://gruss.cc/files/fantastictimers.pdf
Any kind of networking access can get it (with enough samples, you can get some crazy precision over even the most inconsistent networks), and really any kind of I/O could be abused when combined with another exploits.
And if your permissions system doesn't allow I/O, is there really a lot that your program can do?
https://www.microsoft.com/en-us/research/publication/deconst...
By design, Singularity didn't support dynamic code loading so untrusted code would run in another software isolated process (SIP) and separated by a channel boundary (IPC). With Spectre, you'd need to rethink what happens with IPC to and fro the untrusted processes. The core of the system wouldn't need this though.
Singularity also looked to proof carrying code as a way of building reliable systems. Unfortunately, it'd be hard to prove there isn't a Spectre style attack lurking in a piece of code.
HIP is the full address space change, but any mitigation steps can be introduced in the IPC hand-off between SIPs (or at the kernel ABI) compiled into the untrusted process. This code is compiler controlled. The system is in ring-0 so IPC code gets full access to available instructions (such as mitigation is possible there).
Of course doing this makes channel communication in both directions slow for untrusted processes, but that is the cost of doing business with Spectre. And if someone wrote a browser that didn't use channels for talking to the JS engine, then all bets would be off.
Yep, you could say I'm talking about "crippling" web pages, but I'd rather phrase it as "pruning" or "removing the worst abuses".
Those who read carefully might have noticed I didn't say "kill Javascript with fire" or anything like that, and I think we could get far just by having the browsers limiting (by default) Javascript run time to 3 seconds: start at 20 seconds this summer and aim for 3 seconds next summer, because who seriously thinks web pages should need to run scripts for minutes after they have loaded?
I realize there are a couple of problems with the simplified approach above, so let me try to defuse the ones I see right away:
- This will break a number of sites: Yep. But if we really wanted and we got all browser vendors behind it we it wouldn't take long before all mainstream sites would be optimizing their js like crazy to make sure it loaded withing 20, then 19 seconds etc. That said: Good luck getting all browser vendors on board with this. (And, in my defense: I didn't say it would be easy, or even possible, only that I think it would be a good idea to do.)
- Some pages need Javascript because they load content using JS: We can reset the counter to something reasonable for certain types of user input.
- Some web pages needs Javascript for <reasons>: let them have a popup like the ones they get for location sharing etc.
Because that's what you'll end up with.
I'm quite clear in the comment you replied to that I don't want to remove all possibilities for JavaScript.
I want to reduce attack surface significantly and as a nice bonus create a better browsing experience for everyone.
Even I realize Javascript has its place for now at least:
- autocomplete
- client side validation
- complex browser side apps, including games
What I want to curb is websites that keep exercising my CPU for no good reason (for me as a user) because they:
- are poorly written
- are busily loading and reloading ads and trackers
- etc
I don't see how you go from removing the worst abuses i.e. to "everything on the server again like the early 2000s"
Nobody's going to accept a non-interactive web, or even one with crippled interactivity[1]. They only tolerated it in the 90s due to hardware and bandwidth limitations, not to mention that the web was new technology back then. That ship has long since sailed.
Even back then, technologies such as Flash arose because people - initially website owners/creators and then, of course, users - wanted more interactivity than HTML 3.2 and 4.x, and slow/limited JavaScript, could deliver.
[1] With very few exceptions, the GP being one.
I don't talk about that: I talk about limiting web pages from using unreasonable amounts of client side resources.
We're not talking about a lack of interactivity. We're talking about webpages
- loading the page, then leaving the CPU alone
- not running crypto miners (at least not without asking)
- not having all the time in the world to run timing attacks
Something like Flash did make a comeback and it's just called JavaScript now. Pretty much exactly the same attack surface, but at least now its a hopelessly complicated "standard"!
I know, right? I almost miss Flash: it was certainly a lot simpler to work with. I'm not even sure about the "almost" in that sentence.
Flash exploits were a lot more commonplace and frankly embarrassing than what you see in JS engines these days, so I have trouble seeing it as the same attacking surface.
Suppose we devise a protocol that would allow browsers to directly do what we use JS for today? In other words, it would send a request to the server for every interaction with the page, but the response would be a DOM diff in some well-specified standard format, which the browser would apply to "refresh" the page.
At that point it feels like the main difference would be in latency for small page updates. How common are those, and how bad is it on a typical Internet connection? I'd expect there to be enough to break stuff like custom scrollbars and other such widgets, but many people would say good riddance to those.
Or, maybe a simpler question: What web applications do you think of that routinely need to run lots of JavaScript in the background without user interaction?
Right now I can only come up with games and simulations.
As a result of our work, we now believe that speculative vulnerabilities on today’s hardware defeat all language-enforced confidentiality with no known comprehensive software mitigations, as we have discovered that untrusted code can construct a universal read gadget to read all memory in the same address space through side-channels. In the face of this reality, we have shifted the security model of the Chrome web browser and V8 to process isolation.
Processes are pretty heavyweight as a way to perform this sort of isolation. I can't help thinking that something like sthreads and tagged memory from Andrea Bittau's wedge system would be great OS primitives to have right now.
http://www0.cs.ucl.ac.uk/staff/M.Handley/papers/wedge.pdf
RIP Andrea.
There is a known simple mitigation. Don't JIT random JavaScript, back to interpretation. Of course that means there is no use for V8, which is why it's not in this paper.
But even an interpreted language can still be vulnerable to Spectre attacks.
It doesn't matter if your timing source is high noise with lots of jitter, the attacker can repeat the mesurements over and over again and filter everything out.
Even if the attacker can only pull out a few bytes per second, that might be enough to leak something critical like and encryption key or an ASLR offset.
What about a pure functional language, in which a program computes a result that is a function of the input only.
In this case the only timing information usable for side-channel info leakage would be in the input to the program.
I guess the problems in this case become twofold:
* How can we determine we are not leaking any such info in the input to the program?
and
* Is a pure, functional language sufficient to do the stuff we want, or are they too limited, e.g. for use as a browser scripting language.
But spectre is even worse than that as it attacks at the hardware level. If the CPU does any sort of speculative optimization, through things like caches or branch prediction buffers, then it likely can be the target of a spectre attack. You can try to add spectre mitigations to your language's compiler but, as the article discusses, this approach is an uphill battle.
Pure algorithms are a useful thing in programming. However pure algorithms are not useful on their own, you always need something impure to get the output out.
How do you figure? Processes on some OS's, such as Linux, are pretty cheap. So what are you considering heavy? And how do you imagine a "process switching in userspace" type thing to not just have the same weight as real processes? What's the expensive thing you're trying to eliminate?
I'm really thinking that, we don't really need all the features offered by processes. My current understanding is that we just need a different address space. Why can't we, for example, switch from one set of pages to another when we switch from running trusted browser code to running JIT-ed untrusted code? (Leaving of course a small piece of trampoline code mapped, like KPTI.) On a simple level, this could just be calling mprotect() at certain key locations that result in flipping a few bits in the kernel-maintained page tables. With some good design, perhaps the address for untrusted code and data can be so far away from trusted code and data that maybe just one bit flip is needed in a PML4E.
So V8 never was and never will be a silver bullet for running JS/WASM without isolation. I wonder what Edge CDN such as Cloudflare[0] and Flastly[1] are doing to isolate their functions.
[0]https://blog.cloudflare.com/cloud-computing-without-containe...
[1]https://www.fastly.com/blog/announcing-lucet-fastly-native-w...
Yeah, but it's the atypical workloads that get you, and in every place I've worked there's always been at least the odd atypical workload regardless of system, product, platform, technology, target market (including internal and external).
I doubt they will resort to that though. They can do other tricks, since they control the infrastructure, like turning on process isolation automatically for suspiciously behaving code.
In-order isn't a magic bullet against Spectre, they still do spectulative execution after predicting branches and they can still be vunerable. ARM have listed at least one of their in-order cores as vunerable.
To be free of all speculation you have to go back to the 486, which didn't even have branch prediction.
Besides, if you are making custom CPUs there are other options to avoid Spectre that don't require eliminating all spectulative execution.
I haven't fully ingested this paper, and I'm still getting caught up on all the details of things, but https://www.infoq.com/presentations/cloudflare-v8 talks a bit about this. I gotta run right now so that's all I can say at the moment.
As an aside, the notion that it's very risky to rely only on the type system to enforce security boundaries is not new. In 2003, Govindavajhala and Appel [1] showed that you can break the security of the JVM through memory errors. Basically, an untrusted attacker fills memory with a specially formatted data structure such that in case of a random bit flip, with high probability, you get an integer field and a pointer field aliasing each other, allowing you to do pointer arithmetic. By contrast, it's extremely unlikely that a random bit flip allows (say) an unprivileged Unix process to get root access.
Compare this to the various Intel Management Engine exploits from the past couple years which received no media attention and thus no industry attention.
>Extract the hidden state to recover the inaccessible data. For this, the attacker needs a clock of sufficient precision. (Surprisingly low-resolution clocks can be sufficient, especially with techniques such as edge thresholding.)
Might anyone have some good resources or links they could share on this "edge thresholding" technique?
Is there a way to determine whether site isolation is available/enabled on my chrome platform?
Because we're stuck with the hardware we have, we can make software fixes - however they are going to require either slowing down or removing certain features until we get a hardware fix for the issue.
With meltdown, the issue was definitely a failure of the hardware vendors obligation -> Kernel memory should definitely not be exposed to the process and yet it was.
I would say such a justification is somewhat of a cop-out by the hardware vendors.
Spectre not so much.
Spectre is different, it allows reading pages that a process would already have hardware permission to read[1] but actual permission is enforced at the software level (i.e. software as opposed to hardware bound checking).
[1] It is possible to construct spectre v1 attacks against other processes in some cases, but are much harder, low bandwidth and I do not think are yet shown to be practical.
Okay, but the hardware issue is pretty much exactly the same... the bug in hardware that enabled spectre is also what enabled meltdown.
The only reason we managed to fix this is by moving most of the kernel task memory out of the user page table - a software fix for something that the hardware should have been doing.
I'm not an hardware designer, but likely both instances of speculation use the same snapshot and rollback logic, but that's about it.
But either way nothing was actually "crippled", and if anything the reverse is true. They are un-crippling aspects of JavaScript (like re-introducing SharedArrayBuffer).
It isn't like a simple branch with two possible values. The cache line could be one of 256 (if scanning a byte at a time).
Edit: that I know, at least.
Ultimately I don't think many speculative instructions actually pull in cache lines, so we can either clean up when they do, or stop speculating when a line would need to be loaded.
The CPU keeps on executing whatever instructions are ready (i.e. have data dependencies met), irrespective of the in-flight branches. When a mispredicted branch is retired (~0.1-1% of the time), the CPU can throw away the entire reorder buffer and start over. A CPU can also abort-on-execute, throwing away only work related to the given mispredicted branch. So probably 99% of all cycles are spent with at least one unexecuted branch somewhere in the reorder buffer--the CPU is essentially always speculating.
Shrug. Not sure why.
As mentioned elsewhere, this is incorrect and does not describe what CPU designers have done. That probably explains the downvotes.
As noted elsewhere OP completely misunderstood the fact that his CPU has received some of Intel's Spectre mitigation firmware patches and came to the silly conclusion that this meant his CPU was no longer doing speculative execution of any sort.
HN is still a (vaguely) computer-savvy forum. Being extremely wrong about the fundamentals of how a modern CPU works is going to get you downvoted.
This had been the most overt blown security topic I've ever seen. All the POCs required higher help from the code being attacked: eg, known memory locations, nothing else running on the system.
There still has not been a single real world exploit developed for Spectre: eg somebody attack a running server. But yet all this time and effort has gone into a theoretical toy. It is so incredibly difficult to pull this off, that is should be one of the lowest priorities. I guess it gets ink because it pushes the security issues past all the software foul ups and lets people hate on Intel.
Developers gotta develop I guess.
Edit: down to -3 already for expressing an opinion, and nobody really responded except for a variation of the precautionary principle.
I wish more people understood that spectre isn't the bogeyman that we need to give up 5 to 15% of our CPU for (and meltdown was more of a bug and when easily patched is no worse than Spectre).
Maybe I should put an encryption key in a sandboxed system and play $1000 if anybody can use Spectre to recover it? Even then it probably would shut everybody up.
I find statements like yours difficult. Can we know that something did not happen just because nobody talked about it?
A vulnerability as deep and far reaching as Spectre is extremely interesting to nation-state level hacking groups as a means of warfare, we should not assume that we know about everything these actors are up to.
[1] https://twitter.com/davywtf/status/1119783380734836737
[2] https://www.usenix.org/system/files/1401_08-12_mickens.pdf