DEF Con 32 – AMD Sinkclose Universal Ring-2 Privilege Escalation (Not Redacted) [pdf]
media.defcon.org
media.defcon.org
(Actual Ring 2 is very rarely seen, so perhaps I should have known!)
On X86 we have ring 0 and 3, with 1 and 2 never used and removed in newer CPUs. ARM has 3 or 4 privilege layers, but they're named differently.
They probably just called it ring -2 because it's a couple layers below ring 0.
Arm (Aarch64) Exception Level 1 corresponds to Ring 0 of x86.
Arm (Aarch64) Exception Level 2 corresponds to the Hypervisor level a.k.a. Ring -1 of x86.
Arm (Aarch64) Exception Level 3 corresponds to the System Management Mode a.k.a. Ring -2 of x86.
Fortunately, in Arm EL3 the same instruction set is used as in any other level, unlike in x86, where SMM uses the obsolete 16-bit 8086 ISA, so for compiling programs that will be executed in SMM you have to use a special tool set.
Unfortunately, both the Arm EL3 and the x86 SMM allow the manufacturers of computing devices to do things that are either stupid or in direct contradiction with the interests of the owners of the devices and the owners may not be able to do anything to correct this, unless they can exploit vulnerabilities like the one that has now been patched by AMD.
There are no valid arguments for the existence of SMM and EL3 and the fact that they are not forbidden by law is a disgrace for the computing industry.
Arm EL3 has been created as an imitation of the Intel SMM. The Intel SMM has been created because Microsoft was too lazy to introduce the required power management functions in the Windows and MS-DOS operating systems, so they passed the task to the motherboard or laptop manufacturers, for which Intel has provided SMM, to enable this.
https://medium.com/swlh/negative-rings-in-intel-architecture...
There was a time when people thought if we could put the secure code in a lower ring, then with it we could protect the rest of the system. With virtualization, the hypervisor is in ring -1, which is technically not a ring, but rather a mode called VMX root operation, post-VMXON. This enables things like the blue pill attack, where the hypervisor is itself presented with a false image of the underlying physical hardware, by a malicious layer. You can find the same pattern in ARM TrustZone, where the secure code is repeatedly broken.
"if only we had ring -N"
No, not 100% (it was originally), 104.5%. Why? Because you don't go back and change all your rules and documentation following subsequent developments in the field, that causes unnecessary confusion and errors down the road.
[1]: https://en.wikipedia.org/wiki/RS-25#Engine_throttle/output
Unless running hardware also used by powerful hosting providers (some of which care for security), these mitigation will not reach many systems. Checked a few "client" samples, seems like MSI has provided updated binary blobs, ASRock has provided some, Gigabyte has provided broken ones first and then backdated the new ones, ASUS (ROG/RUF/CSM) and Biostar customers are still waiting.
https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...
Originally on x86 systems memory was in VERY short supply - SMM mode memory was the DRAM that the VGA window in low memory (0xa0000) overlaid - normal code couldn't access it because the video card claimed memory accesses to that range of addresses - so the north bridge when the CPU was in SMM mode switched data and instruction accesses to that range to go to DRAM rather than the VGA card .... that's great except remember that SMM mode was used for special setup stuff for laptops .... sometimes they need to be able to display on the screen .... that's what this special mode was originally for: so that SMM mode code can display on the screen (it's also likely why SMM mode graphics were so primitive, you're switching in and out of this mode for every pixel you write)
That said, the SMM can probably be a little less intrusive if it needs to be. Like it doesn't have to freeze the cores if all it is doing is reading your bitcoin addresses and passphrase out of memory, just stalling the memory bus for a moment or two.
pKVM's primary goal is to protect guest pages from a compromised host by enforcing access control restrictions using stage-2 page-tables. Sadly, this cannot prevent TrustZone from accessing non-secure memory, and a compromised host could, for example, perform a 'confused deputy' attack by asking TrustZone to use pages that have been donated to protected guests. This would effectively allow the host to have TrustZone exfiltrate guest secrets on its behalf, hence breaking the isolation that pKVM intends to provide..
FF-A provides (among other things) a set of memory management APIs allowing the Normal World to share, donate or lend pages with Secure. By monitoring these SMCs, pKVM can ensure that the pages that are shared, lent or donated to Secure by the host kernel are only pages that it owns.. the robustness of this approach relies on having all Secure Software on the device use the FF-A protocol for memory management transactions with the normal world, and not use vendor-specific SMCs that pKVM is unable to parse.
On x86, SMM attestation was introduced by Intel (PPAM / Hardware Shield, 11+ gen) and AMD, https://www.microsoft.com/en-us/security/blog/2020/11/12/sys...> Because of its traditionally unfettered access to memory and device resources, SMM is a known vector of attack for gaining access to the OS and hardware.. One could have perfect code in SMM and still be affected by behavior like trampolining into secure kernel code.. Isolating SMM is implemented in three parts: OEMs implement a policy that states what they require access to; the chip vendor enforces this policy on SMIs; and the chip vendor reports compliance to this policy to the OS.
the same thing happened with the ryzenfall/masterkey exploit, where people were just in utter denial there was an actual exploit there, because root is root! People literally spent more time talking about who released it and their background image than the actual exploit. AMD obvious cannot have exploits, that's only an intel thing. /s
"alleged" flaws" (rolls eyes) https://old.reddit.com/r/Amd/comments/845w8e/alleged_amd_zen...
assassination attempt* https://old.reddit.com/r/hardware/comments/849paz/assassinat...
doxxing the researchers: https://old.reddit.com/r/hardware/comments/845xks/some_backg...
https://old.reddit.com/r/Amd/comments/84tftt/clarification_a...
https://old.reddit.com/r/Amd/comments/8589t2/cts_labs_clarif...
HN discussions were not much better, although tpacek is cool.
https://news.ycombinator.com/item?id=16576342
And, like, the fact that AMD released an urgent patch for it should kind of speak to the severity of the issue in the first place. AMD doesn't patch "sudo lets you do root things", obviously, so it necessarily must have been more than that, and this was obvious even at the time. But we have to go through this dance with literally every single AMD exploit.
AMD has a unpatched exploit in all Zen3 and below processors that leaks data from kernel at a faster rate than meltdown did. It was discovered by the same researchers that discovered meltdown. AMD has chosen to leave that unpatched, and put out a weaselly deflection about "it doesn't cross address boundaries" but they also still refuse to turn KPTI on by default because it would hurt their benchmarks. And without KPTI there is no address boundary to cross, that's the weaselly part. AMD very craftily made it sound like they're saying there's not an issue, but in fact they are fully confirming the finding from the researcher, including the suggested mitigation (enabling KPTI), they just don't recommend that you do it. The statement is deliberately short to avoid inclusion of too many details that might dispel these misleading impressions.
https://www.usenix.org/system/files/sec22summer_lipp.pdf
https://www.amd.com/en/resources/product-security/bulletin/a...
proof of concept: https://github.com/amdprefetch/amd-prefetch-attacks/blob/mas...
tested a few months ago and it's still working: https://news.ycombinator.com/item?id=40850526
This follows that same researcher (who previously discovered meltdown) uncovering a prior series of vulnerabilities in the cache ways predictor that also nullify KASLR... which AMD refused to patch because it "didn't leak actual data, only metadata"... the metadata being the page-table layouts. That one is still unpatched too - as the researchers note, AMD never actually mitigated this either, just more weasel words.
https://www.tomshardware.com/news/new-amd-side-channel-attac...
https://mlq.me/download/takeaway.pdf
(this one literally doesn't even seem to have a security bulletin page for itself so I guess they have fully shoved this one down the memory hole now, but here's the news item from wayback) http://web.archive.org/web/20200325045817/https://www.amd.co...
After 6+ years of watching the community defend this behavior, downplay exploits from their favorite megacorporation, etc, it just gets old. Not liking how CTS labs did it or whatever is fine. It doesn't mean there's not a serious exploit, and so often that is where people end up with these AMD exploits, they like AMD so much that they argue against the existence or significance of the exploit, attack the researchers or whine about research grants, etc.
"Does this really deserve this CVE score" is a constant refrain in AMD vuln threads and it just gets so old. As tptacek noted... intel ME vulns are frontpage news and have people asking where they can buy a processor without ME in it. Literally nobody cares that AMD has had these vulnerabilities left open and unmitigated for years and years even though they're actually worse (as judged by the researcher who found both these issues and meltdown).
People would have flipped the fuck out if Intel left meltdown unpatched and released misleading statements implying that it wasn't an issue etc. It is wild just how much AMD is playing on story-mode difficulty with the average enthusiast, and honestly most people don't even realize they're doing it. And that drives me nuts - just decide if security issues are a problem or not, and if the answer is "not" then let's just turn all the mitigations off and see how long they remain un-exploited. If we want to have the security version of the drug-assisted olympics then fine, there is value in having dragsters that just do the thing as quickly as possible, right? But the double-standard people apply to anything AMD is crazy. Talk about your "tyranny of low expectations".