Ginseng: Keeping secrets in registers when you distrust the operating system
blog.acolyer.org
blog.acolyer.org
It arranges things such that this sensitive data is only ever in the clear in registers (never memory),
and is saved in an encrypted memory region called the secure stack on context switchesFrom the article
For data confidentiality, in addition to the call stack management described previously, Ginseng must also intercept all exceptions to save sensitive registers to the callstack before the exception can be handled by the OS. Ginseng intercepts exceptions using dynamic trapping. A NOP instruction is inserted at the beginning of the exception vector code in the kernel, and replaced by a call to the secure monitor at runtime when sensitive data enters registers. (Once the registers are clear of sensitive data, the NOP is restored). Once the OS serves the exception, control is handed back to the app. GService manages the return address to ensure we resume at the correct point in the sensitive function.
Ouch, my performance!
.
Write to kernel code pages?!?
Ouch, my cache! Oh, and my security!
And yes, you have to trust Ginseng more than you trust the OS, it's the whole point and an explicit initial premise.
Basically the threat model seems to be “malicious kernel”, but if the kernel is actually malicious then it can do whatever it wants before the target has an opportunity to “protect” itself
From Intel x86 manual
To aid in handling exceptions and interrupts, each architecturally defined exception and each interrupt condition requiring special handling by the processor is assigned a unique identification number, called a vector number. The processor uses the vector number assigned to an exception or interrupt as an index into the interrupt descriptor table (IDT). The table provides the entry point to an exception or interrupt handler (see Section 6.10, “Interrupt Descriptor Table (IDT)”).
The allowable range for vector numbers is 0 to 255. Vector numbers in the range 0 through 31 are reserved by the Intel 64 and IA-32 architectures for architecture-defined exceptions and interrupts. Not all of the vector numbers in this range have a currently defined function. The unassigned vector numbers in this range are reserved. Do not use the reserved vector numbers.
Vector numbers in the range 32 to 255 are designated as user-defined interrupts and are not reserved by the Intel 64 and IA-32 architecture. These interrupts are generally assigned to external I/O devices to enable those devices to send interrupts to the processor through one of the external hardware interrupt mechanisms (see Section 6.3, “Sources of Interrupts”).
But it's still not relevant, as the kernel can arbitrarily halt execution of any process at any point, and therefore can read the content of all registers whenever it wants them.
Again, if your attack model is a malicious kernel nothing you're doing in user mode is going to protect you. If you're using kernel APIs to install protections against that, you're still dealing with a malicious kernel that can ignore or wrap whatever you do.
If you're trying to mitigate/protect against kernel bugs that's a different threat model, but the same general problems exist only by accident so are less likely to leak useful information.
Not if the kernel can't handle interruptions? The only way for the kernel to "steal control" from the process is through an interruption (software or hardware). If all interrupts are first handled by the secure monitor, which handles saving the registers securely, then it's OK.
In fact, if I was a malicious kernel, I would do precisely that
This whole idea will not work unless your hypervisor is a real complete hypervisor. This attempt at a halfway-hypervisor-lite is doomed to failure for this and many other reasons.
We now describe Ginseng’s runtime protection against such accesses. The runtime protection heavily relies on GService, a passive, app-independent piece of software in the Secure world. GService ensures the code integrity, data confidentiality and control-flow integrity (CFI). It does so only for sensitive functions to minimize overhead. It also modifies the kernel at three points, when booting, when modifying the kernel page table, and when handling an exception. Since we do not trust the OS, the kernel may overwrite the modifications. However, when any of these modifications is disabled, the kernel will infinitely trigger data aborts trying to modify readonly memory, thus ensuring sensitive data are always safe.
I would have look at the source code to find more but the github repository has been deleted.
[1] https://en.wikipedia.org/wiki/Trusted_execution_environment
https://en.wikipedia.org/wiki/ARM_architecture#Security_exte...
This sort of thing is exactly why TEE exists. It cannot be done half in userspace half in TEE
1. the os in charge of the hypervisor
2. the os in charge of the virtualised system.
We're saying that we don't trust 2. so we're going to get the kernel from 1. to intercept interrupts. But that means we already have a trustworthy kernel. The one running the hypervisor, so why aren't we just using that?
Whether this is worthwhile, or even works without holes, is sort of an open question. I agree it's sounds a little heavy on the serpent fat, but the technical promise is definitely achievable.
Here are some likely holes:
They prevent the kernel from mapping “sensitive” code into its own address space, but they don’t seem to prevent the kernel from mapping it into a user address space with write permission, which is just as bad. (Also, the kernel already maps most memory writably in the direct map, and they don’t mention what they do about this, so I would guess that they have a bug.)
They don’t mention protecting sensitive code from DMA.
The kernel can corrupt use code execution in many ways, such as corrupting non-“sensitive” registers. They don’t seem to have a rigorous model to defend against this.
They protect the IDT, but I don’t see anything about protecting the SYSCALL MSRs. The kernel could redirect SYSCALL to skip the magic hook. This might not matter if there are no syscalls in sensitive regions.
I don't understand academia's obsession with these security theater devices -- it would make sense to see it in industry as a hyped buzzword, but I think it makes for weak scholarship.
edit: as some note below, it is a TEE not a TPM or a SE. I don't think the distinction should distract from the point, so I have amended above.
Also where does it say TPM? The paper references Intel SGX and ARM TrustZone, which to my knowledge are both on-CPU, whereas TPM's generally sit separately on the motherboard.
I think they're theater because they do something complicated with ambiguous security benefits. And even if they are used correctly, flaws in the designs like spectre/meltdown/foreshadow/rowhammer/etc etc compromise these use cases.
Here's a sketch of an attack against this:
the registers don't actually get overwritten they get renamed in modern CPUs. That means the old data is still there, just not logically accessible by non-pipelined instructions. It's possible the data is still sitting in the registers, with an old epoch name. The predictors will be predicting branches and other things based on those registers. So by issuing the right instruction and measuring delay, you might be able to create an oracle to see if you guess a byte correctly from the stale data.
I also don't think the secrets are actually out of memory, they are out of your memory, where your is the kernel and the user space but not necessarily the TEE Secure World memory. This is, of course, the same RAM but protected by a page table.
""" To ensure code integrity the kernel page table is made read-only at boot time. The kernel is modified to send a request to GService whenever it needs to modify the page table, GService honours this request only when doing so would not result in mapping the code pages of a sensitive function. The kernel is also prevented from overwriting its page table base register so that it can’t swap the table with a compromised one. """
Of course, this sounds like the perfect kind of attack for rowhammer to break. Just overwrite the page table by doing reads and then overwrite a sensitive function after it's been invoked once and approved, and now you can leak secrets out that way.
etc
TPM is effectively a smart card hanging off a system bus, usually on a separate package.
SGX and TrustZone are CPU features enabling a "secure" run time environment separate from the main run time env.
You have to trust the OS to set everything up, do IO etc. I don't see how its tenable to not trust the OS.
https://www.cs.utexas.edu/~shmat/courses/cs380s/overshadow.p...
EDIT: I see Overshadow[14] was indeed one of the cited references.
Also, direct link to the actual paper is here: https://www.ndss-symposium.org/wp-content/uploads/2019/02/nd...
Good.
If you want as much security as you can get in an unsafe environment, you're going to need to write the entire routine yourself, sometimes on a level down to the bare metal if you have to.
You can also just write it in a separate assembler file and then the C compiler does not see it.
That leaves just the linker, and most optimizing linkers will treat code outside of the purview of the compiler as a "black-box" otherwise you wouldn't be able to link with code that makes system-calls.
TL;DR: It doesn't affect it.
...On the other hand, the design itself seems like a pretty massive hack. The goal is to turn parts of a userland process into the equivalent of a TEE component, without having to manually separate the codebase into two pieces and set up IPC between them. But although that kind of "automagic" approach is easier to use, it also makes it really easy to write security flaws.
For instance, in the example code:
void hmac_sha1(sensitive long key_top,
sensitive long key_bottom,
const uint8_t *data,
uint8_t *result) {
sensitive long tmp_key_top, tmp_key_bottom;
/* all other variables are insensitive */
/* HMAC_SHA1 implementation */
}
It's quite dangerous to say that all other variables are insensitive! It's hard to say for sure without seeing the actual implementation, but SHA-1 requires first expanding the message into a state of 80 32-bit words, before performing 80 rounds of hashing on them. If the state is treated as insensitive, another core could read it out before it actually goes through hashing, in which case the key could be easily recovered. This design might be secure if the SHA-1 function is separate and itself marks all state as sensitive, as long as the key never leaks into memory in between, but that's not how SHA-1 implementations usually work, so I'm pretty suspicious.I tried to find the actual code to determine whether it's actually vulnerable, but failed: it's supposed to be released as open source [1], but the instructions involve downloading from a GitHub repo [2] which is currently marked as private, I guess by mistake.
...I don't really understand why Ginseng doesn't just mark all variables in a sensitive function as sensitive; it's not like memory for the secure stack is particularly scarce. That still leaves other attack vectors, though.
[1] http://www.ruf.rice.edu/~mobile/ginseng.html
[2] https://github.com/susienme/ndss2019_ginseng_arm-trusted-fir...
Is ARM seen as a mess? Are x86's days numbered as ARM catches up and becomes more widespread for personal computing devices?
Encryption, keystrokes?
Web has made the need for desktop applications often unnecessary.
Pick any old version of Linux and browse the local privilege escalation attacks. It seems likely that recent versions of linux have as-yet undiscovered attacks. The existence of even one means that the OS is not trustworthy in the presence of non-trustworthy applications.
As long as Ginseng is more amenable to verification than the Linux kernel, this isn't just pushing the issue around, but rather reducing the work needed.
Reflections On Trusting Trust, http://cm.bell-labs.com/who/ken/trust.html