Plundervolt: Software-Based Fault Injection Attacks Against Intel SGX [pdf]
plundervolt.com
plundervolt.com
There are already simpler solutions for local execution.
Surely someone involved would have known about that. I wonder what chain of events led to the creation of a secure enclave with such well understood flaws.
Don't you think that most engineers will just laugh the idea off, and do just enough to make their PMs happy?
Such nebulous concepts as "secure enclave" only exist in imagination of MBA types. Even the name "enclave" sounds like it was picked by a seasoned talkshop person.
> adversary abuses an undocumented Intel Core voltage scaling interface
Instead of having to buy hardware to mess with the power rail in a controlled manner, you can now just do it with a code snippet!
Which means the model is just as broken even without it being a software attack.
How does such a random event get turned into an exploit?
If I can't undervolt anymore, then my laptop will be unable to run the CPU at 100% without thermal throttling. Not Intel's fault that Dell used a shitty cooling system and a too-high default voltage, but it's yet another penalty that I have to take because Intel failed.
[1] https://www.theverge.com/circuitbreaker/2019/8/28/20837336/a... [2]https://zombieloadattack.com/
Power analysis attack and power glitch attacks are well-known in cryptography and electronics. The classic technique is to monitor the Vcc voltage and current on an oscilloscope and try deducing the internal operation of the chip and extract the secret, or to inject a glitch by on the Vcc rail to induce a fault. A classic side-channel attack, but these attacks required complete physical control over the hardware, and you need to put a dozen of probes on the motherboard in a lab with an elaborate setup, so it was typically not a concern unless it's a crypto key or something (if so, it would be done on a separate security chip with physical defense internally).
But this attack showed that, since every CPU and SoC now has builtin dynamic voltage scaling and power management, by using these features, you can use the CPU to launch a power analysis attack against itself, and you don't even need to touch even a single trace on the PCB, the attack can be launched remotely, and all you need is root access!
This is frightening. Who knows what is going to be the next.
> In 2017, Tang et al. [65] discovered a software-based fault attack, dubbed CLK screw. They discovered that ARM processors allow configuration of the dynamic frequency scaling feature, i.e., overclocking, by system software. Tang et al. show that overclocking features may be abused to jeopardize the integrity of computations for privileged adversaries in a Trusted Execution Environment (TEE). Based on this observation, they were able to attack cryptographic code running in TrustZone. They used their attack to extract cryptographic keys from a custom AES software implementation and to overcome RSA signature checks and subsequently execute their own program in the TrustZone of the System-on-Chip (SoC) on a Nexus 6 device.
it further notes though:
> However, their attack is specific to TrustZone on a certain ARM SoC and not directly applicable to SGX on Intel processors.
(ARM TrustZone is what AMD has in their chips)
Well, I'll be damned: "The PSP itself is an ARM core inserted on the main CPU.[6]" (https://en.wikipedia.org/wiki/AMD_Platform_Security_Processo...)
I'd imagine that the ARM core on the AMD CPU doesn't support changing the clock rate or anything much though.
Seems like you should be able to just avoid applying the bios patch. That said, you won't be able to keep your BIOS up to date.
Also it will be interesting to see whether directly talking to the voltage regulation controller on the mobo will circumvent the microcode protection.
IIRC, reading memory out of the enclave has already been accomplished with other side-channel attacks, so that aspect has been broken. What would be novel is if you could read the per-CPU secret key[1], but that doesn't appear to be what they've done.
[1] Which only Intel knows, which is why remote attestation requires a service contract with Intel for online verification of attestation responses.
Is it even possible to do that without physical access?
Can you confirm that? Do modern motherboards expose an interface that allows directly control over the voltage regulation in software without physical access?! If so, it's terrifying...
There is no way to get the most performance out of your Intel chip without undervolting - especially on mobile, they run really hot under constant load and often throttle. Manufacturers using barely capable heatsinks doesn't help.
Time to switch to AMD when the new Zen mobile chips arrive.
Thanks Intel, you really tried and succeeded.
The very same DFVS is also possible to exploit for side channel attacks. Say, one branch makes the processor kick in in a higher gear, and over millions of branches, you can reliably deduce branch result for operation behind the MCU barrier.
Such a proof can never be final, because the laws of physics aren't either. But just because it's not perfect doesn't mean it wouldn't be good.
Not to mention - these sort of fault injection attacks are way past issues in high-level RTL, and instead are in the post-synethesis, analog/metastable domain specific to a given target process. Simulating those effects sounds prohibitively expensive, and I'm not sure there is any existing formal verification suite that can even do that.
Unless someone modeled at a higher-level that modifying an MSR can cause undervolting, and then this undervolting can in turn cause bitflips to occur, this wouldn't have been caught. And if they did model it, they would have probably thought of this attack anyway :).
Agreed. It's still possible to verify smaller components of the design, though.
> [...] These sort of fault injection attacks are way past issues in high-level RTL. [...] Simulating those effects sounds prohibitively expensive, and I'm not sure there is any existing formal verification suite that can even do that.
It's possible to have a "high-level RTL" design that inherently resists some types of fault injection attacks. TMR is a trivial example: https://en.wikipedia.org/wiki/Triple_modular_redundancy
In fact, there is a wide body of literature studying the application of formal verification to side-channel and fault-injection analysis. Some systems can even synthesize a fault-injection resistant design. Unfortunately it is not realistic to be resistant to "all types of [fault injection] attacks according to the current laws of physics". We can, however, make different models of fault attacks and then prove (or synthesize) that some design is resistant to attacks in that model.
If you're interested, look for publications in CHES, TCAD, FMCAD, DAC, DATE, etc. with keywords like "DFA", "DFIA", "SAT", "fault injection", etc.
I think this is the kind of thing that might happen in 2050 but for now it just sounds infeasible.
- SGX is disabled by default, it has to be enabled for this exploit to be relevant
- POC requires privileged execution, at which point you can safely assume all is already lost
Anyone who has spent time around digital logic circuits will know that messing with voltages will cause errors. If the power lines are too low some transistors will not be able to switch their load. Or too high and you will cause parasitic losses or capacitance in unexpected places. This is actually a really nice attack to show off to people with an interest in computer/electrical engineering because it demonstrates how a basic design constraint can cascade in unexpected ways.
This is an attack on SGX. If you are not using it, it is irrelevant regardless of whether it is enabled.
> POC requires privileged execution, at which point you can safely assume all is already lost
For SGX, this is different. The threat model behind SGX is that anything outside SGX (including OS, BIOS, motherboard, etc.) is untrusted. The whole motivation behind SGX is to create a trusted environment in an untrusted host.
First, this is explicitly a vulnerability compromising SGX itself, so saying SGX has to be enabled is tautological. It's not a vulnerability attacking Intel processors generally, and it's not being billed that way either.
Second, securing memory against untrusted privileged execution is a defining characteristic of SGX, so it's likewise unsurprising that it would be required. That is quite literally the intended purpose of SGX.
The design thesis of SGX is to prevent "all hope is lost" from being true in the context of privileged execution.
In this attack, the researchers found that by adjusting the voltage and clock frequencies of the CPU, they are able to generate errors inside SGX enclave and recover the secret keys inside.
How long have these attacks been around? I only started hearing about them maybe like... in the past week or so? Nothing remotely close to explaining why these aren't used in production.
Also, the domains that have the most to benefit from SGX (heavily regulated ones like healthcare) tend to be very slow adopters of new technology.
I guess there is definitely some level of concern for sidechannel attacks, given Intel's track record, but I don't think that's whats been holding back adoption.
* No outside code can gain access to any part of it - not even the host OS
* Applications run precisely as executed (binaries need to be chryptographically signed)
You could generate cryptographic keys inside the enclave and ensure that the private key is never seen in cleartext outside it.Let's say you're building a SaaS product to host sensitive data and expose an API. Using SGX, the remote endpoint can be ensured of the above.
Root keys are signed by Intel (and recently delegated to cloud providers) - so there is some element of trust there.
This attack uses power management interfaces to the hardware to induce faults in computations in the enclave, leading to data leaks. In the example above, the machine operator (or someone who gained privileged access on their machine) could abuse this to get access to sensitive data that's supposedly only inside the enclave.
In any case and even in the absence of these threat vectors, SGX and similar solutions should be used as a layered security approach, not something to rely solely upon.
More on faults-- A "fault attack" (I like to call them active attacks) is an attack where someone actively tampers with a device to "force" it into an "unreachable" state. An example of an "unreachable" state might be, for a single cycle, cutting off the voltage to a flip-flop to prevent a bit from being written. Contrived example, but enough to get the idea of "fault" in your head hopefully.
Intel SGX is heavily advertised as being fault-resistant (although in recent years, time and time again it has shown to not be the case!). This is especially crucial because some crypto algorithms (e.g. RSA-CRT and AES in general) heavily rely on zero/minimal faults due to the nature of how these algorithms are constructed. RSA-CRT, for example, only needs a single fault in its algorithm for the private key to be compromised.
CPU design today will also typically employ frequency or voltage scaling based on workload as a power saving metric-- say, if the CPU is idle. They will also make this ability software-visible, as it's quite handy for general uses as a whole.
The researchers here exploit this ability to "cut the voltage" from software in the CPU to directly affect SGX as well, resulting in faults in the underlying operations.
While fault attacks (especially timing ones) on SGX are nothing new, the shocker frankly is how they did the fault attack. Typically, in my experience, whenever someone says "power fault attack" you might envision someone with physical access to the chip "glitching" the clock or VCC to bypass some verification step, maybe in a lab or using a ChipWhisperer. This is obviously not the case here-- this is the first case of software power glitching I've seen (note that software clock glitching was done a few years back, but not on Intel machines I believe).
I feel the need to point out that this isn't a fault (pun intended) of Intel's design-- OK maybe it is a little bit since this does fall under the SGX-type threat model, but I feel like this would more be an industry wide issue for any system with an actual hardware enclave (looking at you Apple), so Intel is unlikely alone in the blame, in the same way that speculative execution was (and still is, in my opinion!).