Zenbleed
lock.cmpxchg8b.com
lock.cmpxchg8b.com
Remember: just because this one bug gets fixed in microcode doesn't mean there's not another one of these waiting to be discovered. Many (most?) 0-days are known about by black-hats-for-hire well before they're made public.
CPU vulnerabilities found in the past few years:
https://en.wikipedia.org/wiki/Meltdown_(security_vulnerability)
https://en.wikipedia.org/wiki/Spectre_(security_vulnerability)
https://aepicleak.com/
https://en.wikipedia.org/wiki/Software_Guard_Extensions#SGAxe
https://en.wikipedia.org/wiki/Software_Guard_Extensions#LVI
https://en.wikipedia.org/wiki/Software_Guard_Extensions#Plundervolt
https://en.wikipedia.org/wiki/Software_Guard_Extensions#MicroScope_replay_attack
https://en.wikipedia.org/wiki/Software_Guard_Extensions#Enclave_attack
https://en.wikipedia.org/wiki/Software_Guard_Extensions#Prime+Probe_attack
https://www.vusec.net/projects/crosstalk/
https://en.wikipedia.org/wiki/Hertzbleed
https://www.securityweek.com/amd-processors-expose-sensitive-data-new-squip-attack/It strikes me that it's either that branch prediction is so inherently complex enough it's always going to be vulnerable to this and/or it just so defies the way most of us intuitively think about code paths / instruction execution that it's hard to conceive of the edge cases until too late?
At what point does the complexity of CPU architectures become so difficult to reason about that we just accept the performance penalty of keeping it simpler?
Never for branch prediction. It just gets you too much performance. If it becomes too much of a problem, the solution is greater isolation of workloads.
It even says so in the document:
Note that it is not sufficient to disable SMT.
Apple's chips don't have this vulnerability, but it's not because they don't have SMT. They just didn't write this particular defect into their CPU implementation.I think we may be seeing an industry-wide shift away from SMT because the performance penalty is small and the complexity cost is high, if so that fits parent's speculation about the trend. In a narrow sense Zenbleed isn't related to SMT but OP's question seems perfectly relevant to me. I come from a security background and on average more complicated == less secure because engineering resources are finite and it's just harder and more work to make complicated things correct.
Speculation is hard, it's sort of akin to the idea of introducing multithreading into a program, you are explicitly choosing to tilt at the windmill of pure technical correctness because in a highly concurrent application every error will occur fairly routinely. Speculation is great too, in combination with out-of-order execution it's a multithreading-like boon to overall performance, because now you can resolve several chunks of code in parallel instead of one at a time. It's just also a minefield of correctness issues, but the alternative would be losing something like the equivalent of 10 years of performance gains (going back to like ARM A53 performance).
The recent thing is that "observably correct" needs to include timings. If you can just guess at what the data might be, and the program runs faster if you're correct, that's basically the same thing as reading the data by another means. It's a timing oracle attack.
(in this case AMD just fucked up though, there's no timing attack, this is just implemented wrong and this instruction can speculate against changes that haven't propagated to other parts of the pipeline yet)
The cache is the other problem, modern processors are built with every tenant sharing this single big L3 cache and it turns out that it also needs to be proof against timing attacks for data present in the cache too.
Basically never for anything that's at all CPU-bound, that growth in complexity is really the only thing that's been powering single-threaded CPU performance improvements since Dennard scaling stopped in about 2006 (and by that time they were already plenty complex: by the late 90s and early 2000's x86 CPUs were firmly superscalar, out-of-order, branch-predicting and speculative executing devices). If your workload can be made fast without needing that stuff (i.e. no branches and easily parallelised), you're probably using a GPU instead nowadays.
For various reasons, all these infras eventually lost out in the market due to market pressure (and cost/watt/IPC, I guess).
With SMT allowing twice the cores on a CPU for most workloads, disabling it would double the cost for most providers!
There are VPS providers that will let you rent dedicated CPU cores, but they often cost 4-5x more than a normal virtual CPU. Overprovisioning is how virtual servers are available for cheap!
Providers usually won't disable SMT completely, they'd run a scheduler which only allows 1 VM to use both SMT threads of a core. Ultra cheap VPS providers may still find that not worth the pennies though as if you sell a majority of single core VPS then the majority of your SMT threads are still unavailable even with the scheduler approach.
Fully dedicated cores aren't necessarily required because in the timesliced case the registers are unloaded and reloaded when different VMs are shuffled on and off the core. That said, they definitely prevent the cross-vm-data-leak case of this bug.
Registers are unloaded and reloaded when different processes / threads are scheduled within a running VM too. That should protect the register contents, but because of this issue, it doesn't, so I don't see why it would if it's a hypervisor switching VMs instead of an OS switching processes. If you're running a vulnerable processor on a vulnerable microcode, it seems like you can potentially read things put into the vulnerable registers by anything else running on the same physical core, regardless of context.
On the other hand, "context switching" for VMs is done via hardware commands like VMSAVE/VMLOAD or VMWRITE/VMREAD which do save/load the entire guest register context, including the hidden context not accessible by software which this CVE is relying on. Not that it isn't impossible for this to be broken as well, but it's a completely different procedure and one the hardware is actually responsible for completely clearing instead of "supposed to be reset by software".
So while the CVE still affects processes inside of VMs the loading/unloading behavior inter VM should actually behave as a working sandbox and protect against cross-VM leaks, barring the note by lieg on SMT still possibly being a problem (I don't know enough about how the hardware maintains the register table between SMT threads of different VMs to say for sure but I'm willing to guess it's still vulnerable on register remappings).
There may well be other reasons I'm completely mistaken here but they'd have to explain why the inter-VM context restore is broken not why it works for inter-process restore. The article already explains why the latter happens, but it doesn't make a claim about the former.
If we assume the register file isn't saved, just the visible registers, what's happening is the visible registers are restored, but the speculative dance causes one of the other values in the register file to become visible. If that's one of the restored registers, no big deal, but if it was someone else's value, there's the exploit.
If you look at the exploit example, the trick is that when the register rename happens, you are re-using a register file entry, but the upper bits aren't cleared, they're just using a flag to indicate the bits are cleared; then when rolling back the mispredicted vzeroupper unsets the flag, the upper bits of the register file entry are revealed.
For example, let's imagine a toy architecture with two registers: r0 and r1. We can create a little assembly snippet using them: "r0 = load(addr1); r1 = load(addr2); r0 = r0 + r1; store(addr3, r0)". Pretty simple.
Now, what happens if we want to do that twice? Well, we get something like "r0 = load(addr1); r1 = load(addr2); r0 = r0 + r1; store(addr3, r0); r0 = load(addr4); r1 = load(addr5); r0 = r0 + r1; store(addr6, r0)". Because there is no overlap between the accessed memory sections, they are completely independent. In theory they could even execute at the same time - but that is impossible because they use the same registers.
This can be solved by adding more physical registers to the CPU, let's call them R0-R6. During execution the CPU can now analyze and rewrite the original assembly into "R1 = load(addr1); R4 = load(addr4); R2 = load(addr2); R5 = load(addr5); R3 = R1 + R2; R6 = R4 + R5; store(addr3, R3); store(addr6, R6)". This means we can now start the loads for the second addition before the first addition is done, which means we have to wait less time for the data to arrive when we finally want to actually do the second addition. To the user nothing has changed and the results are identical!
The issue here is that when entering/exiting a VM you can definitely clear the logical registers r0&r1, but there is no guarantee that you are actually clearing the physical registers. On a hardware level, "clearing a register" now means "mark logical register as empty". The CPU makes sure that any future use of that logical register results in it behaving as if it has been clear, but there is no need to touch the content of the physical register. It just gets marked as "free for use". The only way that physical register becomes available again is after a write, after all, and that write would by definition overwrite the stale content - so clearing it would be pointless. Unless your CPU misbehaves and you run into this new bug, of course.
As usual, cost is the biggest hindrance to security.
Leaks between different EC2 instances would be far more serious, but I suppose that wouldn't happen unless two tenants / EC2 instances shared SMT cores, or the contents of the microarchitectural register file was persisted across VM context switches in an exploitable manner.
Still, I think that if your company is handling user data it's worth seriously considering dedicated instances for any service that encounters plaintext user information.
IBM's VM was and is a hypervisor. It dates to the mid 1960s, in the form of CP-40, and it didn't run opcodes in software, but in hardware.
https://en.wikipedia.org/wiki/IBM_CP-40
p-code machines, which interpret bytecode, date back almost as far, such as the O-code machine for BCPL.
https://en.wikipedia.org/wiki/BCPL
Getting people to distinguish between these concepts is probably a lost cause.
"PC" virtualization's getting closer to big iron virtualization, but likely will never quite get there.
Also -- I was running virtual machines on a 5150 PC when it was a big fast machine -- the UCSD P System ran a p-code virtual machine to run p-code binaries which would run equally well on an apple 2. In theory.
I think people here of all places should distinguish between these concepts.
There are big performance and security implications of the two approaches.
Just how many times is the average operating system workload (with or without a virtual machine also running a second average operating system workload) context switching a second?
Like... unless I'm wrong... the kernel is the main process, and then it slices up processes/threads, and each time those run, they have their own EAX/EBX/ECX/ESP/EBP/EIP/etc. (I know it's RAX, etc. for 64-bit now)
How many cycles is a thread/process given before it context switches to the next one? How is it managing all of the pushfd/popfd, etc. between them? Is this not how modern operating systems work, am I misunderstanding?
For anyone interested in seeing this on the nearest Linux box:
vmstat -S M 1
Watch the 'cs' column go wildDepends on a lot of things. If it's a compute heavy task, and there's no I/O interrupts, the task gets one "timeslice", timeslices vary, but typical times are somewhere in the neighborhood of 1 ms to 100 ms. If it's an I/O heavy task, chances are the task returns from a syscall with new data to read (or because a write finished), does a little bit of work, then does another syscall with I/O. Lots of context switches in network heavy code (io_uring seems promising).
> How is it managing all of the pushfd/popfd, etc. between them?
The basic plan is when the kernel takes an interrupt (or gets a syscall, which is an interrupt on some systems and other mechanisms on others), the kernel (or the cpu) loads the kernel stack pointer for the current thread, then it pushes all the (relevant) cpu registers onto the stack, then the kernel business it taken care of, the scheduler decides which userspace thread to return to (which might be the same one that was interrupted or not), the destination thread's kernel stack is switched to, registers are popped, then the thread's userspace stack is switched to, then userspace execution resumes.
A nice way of thinking about it is the kernel visualizes the CPU among multiple programs.
Great reading material on all this OS stuff: https://pages.cs.wisc.edu/~remzi/OSTEP/
I'd like to be educated here why a big switch statement wouldn't necessarily protect us from these CPU vulnerabilities? Anyone willing to help?
So, yes, the switch statemement might be safe, but you would need to prove that your switch statement doesn't use those instructions. You don't get to claim that for free just because you are using a switch-statement.
Conversely, even if you execute bare metal instructions for the user of the VM, you could also deny those instructions to the user. Eg by not allowing self-modifying code, and statically making sure that the relevant code doesn't contain those instructions.
So the switch statement by itself does not do anything for your security.
But yes, unless you solve the halting problem, anything that bans all bad programs will also have false positives. It's the same with type systems in programming languages.
Wouldn't we be able to avoid the "big payoffs" of no-breakout exploits if we had specialized hardware handle the secrets?
- `2023-05-09` A component of our CPU validation pipeline generates an anomalous result.
- `2023-05-12` We successfully isolate and reproduce the issue. Investigation continues.
- `2023-05-14` We are now aware of the scope and severity of the issue.
- `2023-05-15` We draft a brief status report and share our findings with AMD PSIRT.
- `2023-05-17` AMD acknowledge our report and confirm they can reproduce the issue.
- `2023-05-17` We complete development of a reliable PoC and share it with AMD.
- `2023-05-19` We begin to notify major kernel and hypervisor vendors.
- `2023-05-23` We receive a beta microcode update for Rome from AMD.
- `2023-05-24` We confirm the update fixes the issue and notify AMD.
- `2023-05-30` AMD inform us they have sent a SN (security notice) to partners.
- `2023-06-12` Meeting with AMD to discuss status and details.
- `2023-07-20` AMD unexpectedly publish patches, earlier than an agreed embargo date.
- `2023-07-21` As the fix is now public, we propose privately notifying major distributions that they should begin preparing updated firmware packages.
- `2023-07-24` Public disclosure.
> As the fix is now public, we propose privately notifying major distributions that they should begin preparing updated firmware packages.
AMD had to drop the ball somewhere didn't it.
Not that I think it's realistic to develop an exploit and gain real value in three days, but theoretically, if all parties had taken more than three days to distribute and apply the patches?
amd-ucode 20230625.ee91452d-5
last updated 2023-07-25 11:48 UTC
Contains the microcode update that addresses this?
https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin... says that the fixed version is 2023-07-18, but the amd-ucode version in Arch is 20230625.. but it was last updated in 2023-07-25..
My guess is that this is still getting the 20230625 firmware, per the PKGBUILD at https://gitlab.archlinux.org/archlinux/packaging/packages/li...
Which contains those lines
_tag=20230625
source=("git+https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...")
I suppose that it isn't up to date and thus Arch Linux is still vulnerable, right?
edit:
but actually there's two commits in the _backports array (which contains cherry-picked commits) that was last edited 20 hours ago
https://gitlab.archlinux.org/archlinux/packaging/packages/li...
Which is 0bc3126c9cfa0b8c761483215c25382f831a7c6f and b250b32ab1d044953af2dc5e790819a7703b7ee6
And b250b32ab1d044953af2dc5e790819a7703b7ee6 appears to be the commit I linked ealier at git.kernel.org so hopefully up-to-date Arch is not vulnerable to zenbleed
Either way, as noted elsewhere in the comments, only the Rome CPU series has received updated microcode with fixes. All other Zen 2 users need the fix that was released as part of Linux 6.4.6: https://lwn.net/Articles/939102/
(which has been built and packaged for Arch)
Thankfully the exploit is highly dependent on a specific asm routine so exploiting it from JS or WASM in a browser should be extremely difficult. Otherwise a nefarious tab left open for hours in the background could exfiltrate without an issue.
I'm eagerly waiting for Fedora maintainers to push the new microcode so the kernel can update it during the boot process.
At least one commentor here claims to be able to reproduce this with javascript: https://news.ycombinator.com/item?id=36849767 .
I assume that once/if a method is found it will be applicable broadly though. At the same time, hopefully software patches in V8 and SpiderMonkey will be able to mitigate this further and sooner.
But a JS exploit would require some way to exfiltrate data and presumably doing that would be quite difficult to hide entirely.
The latest amd firmware version is 20230625.
Apart from that, it's necessary to "sudo emaint sync -A && sudo emerge -av sys-kernel/linux-firmware", while checking that the correct files are included in the savedconfig file if using it. After that, rebuild the kernel or the initramfs and reboot.
Really what ELI5 is, is a technique to allow the asker to not have to look anything up. From the parent comment, you can look up "patch", "AMD", "microcode"; or you can demand "ELI5!" and have someone else type up long, careful definitions that don't reference context or words that a 5 year old doesn't know.
Regarding what microcode is, here is a good explanation of the differences between microcode and firmware:
https://superuser.com/questions/1283788/what-exactly-is-micr...
Appreciate the link! I'm not OP but that's exactly what I was looking for.
Sounds like cope being outprogrammed by a kindergartner i Roblox
Microcode updates are always discussed when talking about microarchitectural security vulnerabilities (and other scary CPU errata like https://lkml.org/lkml/2023/3/8/976).
Microcode is always mentioned when discussing CPU design evolution.
Just because something is familiar to you, or even large swaths of a given population, doesn't mean everyone should be expected to know it.
I love learning new things. I love discovering topics I know nothing about, and I love picking the brains of those passionate about them. But the condescension from a certain type of tech nerd sucks all the fun out of learning. I've certainly been guilty of this in the past.
you're not going to convince others that microcode is some kind of foreign concept to CPUs just because you yourself were unfamiliar.
Yes, it can be a downer to discover that you're more naive in a subject than you had previously thought you were more familiar.
>Also curious the Wikipedia article for CPU design doesn't mention it, since it's "always" referenced.
microcode is something that is implemented by CPUs that are too big and expensive to replace -- it's not something that is fundamental to processor designs. It's something we now live with to prevent things like the 'pentium bug' from costing Intel many-many dollars after a consumer-products forced recall/replacement.
At this point in history I think that if someone wants to consider themselves to be well-versed or knowledgeable about consumer CPUs then learning about microcode is a hard requirement. It's a false metaphor now to consider a CPU to be an unchanging entity, and that's important to at least be aware of -- it's literally one of the only ways that t
Since wikipedia is the source du joure, here : https://en.wikipedia.org/wiki/Microcode
p.s. : I think it's a strange as you that the processor wiki page doesn't at least mention microcode, I guess they're trying to keep it 'pure'.
> At this point in history I think that if someone wants to consider themselves to be well-versed or knowledgeable about consumer CPUs then learning about microcode is a hard requirement.
This statement strikes me as hyperbolic. A CPU/hardware engineer, or even security-conscious software engineer, sure. But I can't understand why there is a reason for a consumer to care.
A modern generalist CPU is made of many smaller, simpler, specialized CPUs : there's a whole orchestra inside.
Amongst those smaller CPUs, there's a master : it'll see to decoding of instruction, sending jobs to the various CPU units, and fetching the results of said jobs. That master is running a program, executing ... microcode ! And of course, if there is a program, there are bugs. CPUs have bugs since CPUs were invented.
Microcode itself was present in early CPUs, (say, the Z80), but hardcoded. Nowadays, microcode can be uploaded to a CPU to fix bugs.
I don't think that is correct. AMD has released a microcode update[0] for family 17h models 0x31 and 0xa0, which corresponds to Rome, Castle Peak and Mendocino as per WikiChip [1].
So far, there seems to be no microcode update for Renoir, Grey Hawk, Lucienne, Matisse and Van Gogh. Fortunately, the newly released kernels can and do simply set the chicken bit for those. [2]
[0] https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...
[1] https://en.wikichip.org/wiki/amd/cpuid#Family_23_.2817h.29
[2] https://github.com/torvalds/linux/commit/522b1d69219d8f08317...
`good_revs` as per the kernel: https://github.com/torvalds/linux/commit/522b1d69219d8f08317...
Currently published revs ("Patch") (git HEAD):
https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...
As of this writing, only two of the five `good_rev`s have been published.
That's the same codename Intel used for Celerons 24 years ago, the ones famous for 50% overclocks:
https://ark.intel.com/content/www/us/en/ark/products/codenam...
This technique is CVE-2023-20593 and it works on all Zen 2 class processors, which includes at least the following products:
AMD Ryzen 3000 Series Processors
AMD Ryzen PRO 3000 Series Processors
AMD Ryzen Threadripper 3000 Series Processors
AMD Ryzen 4000 Series Processors with Radeon Graphics
AMD Ryzen PRO 4000 Series Processors
AMD Ryzen 5000 Series Processors with Radeon Graphics
AMD Ryzen 7020 Series Processors with Radeon Graphics
AMD EPYC “Rome” ProcessorsThe above are desktop. If they meant APUs, it would list "Ryzen 3000 Series Processors with Radeon Graphics."
Is it likely that this same technique (or similar) also works on earlier (Zen/Zen+) or later (Zen3) cores, but they just haven't been able to demonstrate it yet?
I have an "AMD Ryzen 9 5950x Desktop Processor" which appears to be Zen 3. I think I'm good?
(Not that I'm running untrusted workloads, but yknow, fortune favors the prepared)
But yes, you are fine, 5950x is Zen3.
Anecdotally, I tried to reproduce on my 5600g but couldn't. Which is surprising because they claim it works on 5700u...
Edit: just discovered that while my 5600g is Zen3, the 5700u is Zen2. Lol.
I would worry about cross site leakage. From my understanding that would be unavoidable as soon as you have more tabs open than cores, which feels like an unworkable restriction.
Imagine opening a 9th tab and bring told you need to upgrade your 3700X to a 3900X.
But that means essentially reserving a core for the browser only. I don't think that would be shippable by default.
and also xbox and that thing from valve?
0 - https://blog.playstation.com/2020/03/18/unveiling-new-detail...
I don't know if that mechanism persists into the PS4/PS5.
As for the value of being able to do 'hero attacks' on game consoles, let me point out that once you have a cleartext dump of a game, you've already done most of the work. The Xbox 360 was actually very well secured, to the point where it was easier to hack a disc drive to inject fake authentication data into a normal DVD-R than to actually hack a 360's CPU to run copied games. That's why we didn't have widely-accessible homebrew on that platform for the longest time. Furthermore, you can make emulators that just don't care about authenticating media (because why would they) and run cleartext games on those.
It's a single-core 128 MB VPS, which seemed fine for my boring static html articles. I guess I underestimated the interest.
HTTP/1.1 200 OK
Date: Mon, 24 Jul 2023 17:05:06 GMT
Server: ApacheI'll have to debug when things cool down.
I'm aware 128M is ludicrous in 2023... "a fun challenge", I thought to myself. I can be a dummy.
Cache control headers will help with return traffic
More cpu cores
If using nginx ensure sendfile is enabled and workers are set to auto or tuned for your setup
Check ulimit file handle limits
Offload static assets to cdn
Since it’s a static html site, you could even host on s3, netlify, etc
Unless your CPU is burning due to additional system calls being made.
I wouldn't be so sure, given that without zlib HTTP connections take longer, thereby increasing the size of the wait queue and the number of parallel connections.
According to AMD's security bulletin, firmware updates for non-EPYC CPUs won't be released until the end of the year. What should users do until then, disable the chicken bit and take the performance hit?
It previously read:
> The attack can even be carried out remotely through JavaScript on a website, meaning that the attacker need not have physical access to the computer or server.
Now it reads:
> Currently the attack can only be executed by an attacker with an ability to execute native code on the affected machine. While there might be a possibility to execute this attack via the browser on the remote machine it hasn’t been yet demonstrated.
I don’t really understand how CPU microcode updates work. If I’m keeping Ubuntu up to date, will this just happen automatically?
microcode changes are provided to the CPU at boot time and are only valid early in the boot process. the machine UEFI/BIOS must apply them.
> microcode: microcode updated early to revision 0xa6, date = 2022-06-28
every time I think I'm right, I'm wrong, and every time I think I'm wrong, I'm right.
except here. I'm always wrong, here.
Sort of weirds me out that my OS can just silently update my CPU - I didn’t realize I was giving it that level of control… I guess it’s good vs the alternative of no-one actually updating for exploits like his though.
They mention perf issues for the workaround but they're notably absent from the microcode commentary.
I wonder what this is going to do to the new AMD hardware AWS is trying to roll out, which is supposed to be a substantial performance bump over the previous generation.
They've proven Zen 2 has this problem. They haven't proven no other AMD processors have it. A bunch of people looking to make names for themselves are probably busily testing every other AMD processor for a similar exploit.
I am OOTL on this one, do you have some information you could share?
> It was challenging to get the details right, but I used this to teach my fuzzer to find interesting instruction sequences. This allowed me to discover features like merge optimization automatically, without any input from me!
What happens when AI can start fuzzing software? Seems like a golden opportunity for opsec folks.
Edit: just realized it must have been that the initramfs image is not updated with the manually updated firmware in /lib/firmware.
Edit2: Updated the initramfs and even if the benchmark.sh fails, ./zenbleed -v2 still picks out and prints strings which doesn't happen with the wrmsr solution.
The fixed Renoir microcode should have revision >= 0x0860010b as per the kernel: https://github.com/torvalds/linux/commit/522b1d69219d8f08317...
Updated microcode shows 0x08600106 revision so I guess that explains it.
But a simple example is `vzeroupper` followed by anything that writes a secret to the same register file entry would be leaked on a subsequent flush.
For example, a failed speculation of vzeroupper could result in it erroneously claiming a register by clearing the zero flag on the wrong register - which would mean that the previous data of that register is now suddenly available. If that register has not been touched since a context switch, it could leak data from another process.
The linked article has an animation which suggests that it clears the zero flag on the previously-used register - which indeed requires the victim to reuse the register in the small amount of time between it being marked as zero and the zero being cleared again.
However, the linked Github repo states:
> The undefined portion of our ymm register will contain random data from the register file. [..] Note that this is not a timing attack or a side channel, the full values can simply be read as fast as you can access them.
This suggests that it does indeed do something akin to clearing the zero flag of a random register.
So even with SMT disabled, each core will execute sequentially many threads, switching every few milliseconds from one thread to another, and each context switch does not modify the hidden registers, it just restores the architecturally visible registers.
This is done frequently for high-performance applications.
In Linux it's possible to stipulate that, for instance, core 7 can only be used by super secret process PID 1234. If you have 400 other threads, that means the other threads will have to compete for cores 0-6. And if super secret PID 1234 is idle and there are 12 threads that are marked for scheduling, then they get to just wait for cores 0-6 to become available while core 7 stands idle.
I watched a talk several years ago about a HFT firm that ... abused? this principle. They had a big ass monster of a machine. Four sockets, four CPUs with gobs of cores and gobs of cache on each one. But the only thing they cared about was the latency on their HFT trade sniping process. If they could reduce the latency of receiving interesting information to executing a trade on that interesting information from (making up numbers) 1.1ms to 0.9ms, that was potentially thousands, millions of dollars in profit.
So if CPU socket 0 has cores 0-15, CPU socket 1 has cores 16-31, CPU socket 2 has cores 32-47, CPU socket 3 has cores 48-63, they marked cores 17-31,33-47,49-63 to be usable by nothing. Those cores are permanently and forever idle. They will never execute a single instruction. Ever. Core 16 can be used by PID 12345 and only by PID 12345, core 32 can be used by PID 7362 and only PID 7362, and core 48 can be used by PID 8765 and only PID 8765. This ensured that all data and all instructions used by their super high priority HFT process can never, ever be evicted from the cache.
Apparently it made a notable improvement in latency and therefore profit.
However, one thing that bothers me is that the author claims it's possible to retrieve private keys or root passwords by triggering a faulty revert from the instruction that resets the upper bits of a register. Where is the demo results? All I see is a small-enough gif that looks like the Matrix terminal text scrolling through. Is there any way (other than running the exploit program myself) to check the results and see that it actually leaked the root password and other information?
One thing I don't get though, if amd64 is affected, shouldn't the complementary instructions in X64/intel also be affected? Does intel move around RAT referenced values instead of just setting/unsetting the z-bit?
They don't persist across a reboot, so you can't break anything. You can undo what you just did without a reboot, just use `... & ~(1 << 9)` instead (unset the bit instead of set it).
Does this mean Ryzen CPUs without integrated graphics are fine?
https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Al...
Both of them are missing the newer 7000-family products with Zen2 like 7520U etc.
https://www.amd.com/en/products/apu/amd-ryzen-5-7520u
The Athlon is missing, though.
Wait... now there's also APU's under the AMD Athlon brand? I know that people are happy when AMD's product offerings are on-par or outperforming Intel, but they didn't have to outdo Intel in the consumer confusion arena as well.
https://www.techpowerup.com/cpu-specs/athlon-200ge.c2073
Intel also used the Pentium branding for low-end processors (below i3 and in the Atom lineup), and followed it up with the rather perplexing move of using their company name as the sole branding for their worst products ("Intel Processor").
So I'm guessing it shouldn't affect any of the more recent and very popular Zen3 cpus like the 5600, 5700 etc. I personally own a 5600, which are a great bang for buck.
It's in rather a sweet-spot as far as performance-power-area, so this isn't entirely a bad thing. Zen3's main innovation was unifying the CCXs/caches, but if you only have a 4C, or you want to be able to power-gate a CCX (and its attendant IF links/caches) down entirely, Zen2 does that better, and it's slightly smaller. We'll be seeing Zen2 products for years to come, most likely.
They're not planning to fix this until fucking December...
As far as I can see there is no easy way to use MSR on windows.
bcdedit /set xsavedisable 1
and then reboot.
At least it fails the zenbleed POC.
[0]: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
That and these days you can run so much on a single machine its really hard to understand the thinking that colo isn't and wasn't the best option. AWS isn't some mythical devops dream that never fails. It's a highly convoluted (in price as well as function) affair.
You can see this in glibc's implementation, which checks for crossing page boundaries: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/x86... (line ~68)
> If you can’t apply the update for some reason, there is a software workaround: you can set the chicken bit DE_CFG[9].
It reminds me of the compiler switches which can alter the way code at different levels (global, procedure, routine) can access variables declared at different levels and the change in scope that ensues.
Maybe some of this HW caching should be left to the coders.
Schemes for the blind are news to me though
Yes, I love flashing BIOS...
edit nvm, Microcode can get updated via system updates.
Put the file on a USB drive, plug it in, restart and go into the bios, look for the flashing utility, select the file, done. As long as the machine is on a UPS in case of disaster, everything's accounted for.
Not just for convenience, but safety. You don't want to be caught out when something goes wrong, even without flashing the bios.
A lot of boards these days have 7-segment displays. They're not great, but they're a good step up. Don't need to spend a lot, I think they show up on $300-ish boards. Mine definitely does.
A good UPS does more than just protect from outages. It also protects from surges and low-voltage situations that can both damage the equipment severely.
A UPS doesn't cost much and will last many years. Buying a new motherboard and GPU because they got fried is much more expensive.
The substation is mandated to trip in that anomaly, btw. Otherwise some very common types of motors (AC induction, single and 3-phase) would burn from the excessive current they draw to compensate for the reduced voltage.
Even EFI updates rarely are very intrusive or dangerous, and can also be handled by the Operating System via an update.
But of course BIOS updates have many downsides and often stop after a few years.
> If you remove the first word from the string "hello world", what should
> the result be? This is the story of how we discovered that the answer
> could be your root password!
Can you please expand on your question?If you don't believe the text then how would you believe the video? Anything can be done in devtools beforehand and I can think of a million different ways to fake the video.
Personally, if I didn't trust the text then an easily faked video wouldn't placate me either.
You've even been quoted elsewhere in this thread about this topic.
From the article:
We now know that basic operations like strlen, memcpy and strcmp will use the vector registers -
so we can effectively spy on those operations happening anywhere on the system! It doesn’t matter
if they’re happening in other virtual machines, sandboxes, containers, processes, whatever!
All you need to do is write some JavaScript that will "trigger something called the XMM Register Merge Optimization2, followed by a register rename and a mispredicted vzeroupper". It's up to the hacker to determine how to do this explicitly in JS, but it's theoretically possible by literally any application at any time on any operating system. Even if some language or interpreter claims to prevent it, it's possible to find an exploit in that particular language/interpreter/etc to get it to happen.This is how exploit development works; if you can't go straight ahead, go sideways. I guarantee you that someone will find a way, if they haven't yet.
> The attack can even be carried out remotely through JavaScript on a website, meaning that the attacker need not have physical access to the computer or server.
https://web.archive.org/web/20230725020052/https://blog.clou...
To:
> Currently the attack can only be executed by an attacker with an ability to execute native code on the affected machine. While there might be a possibility to execute this attack via the browser on the remote machine it hasn’t been yet demonstrated.
https://web.archive.org/web/20230726204030/https://blog.clou...
I think I am vindicated -- they did just make that up, likely from the claim posted here.
There's tons of million dollar/month businesses on ~$20/month accounts on shared machines.
>We now know that basic operations like strlen, memcpy and strcmp will use the vector registers - so we can effectively spy on those operations happening anywhere on the system! It doesn’t matter if they’re happening in other virtual machines, sandboxes, containers, processes, whatever!
>This works because the register file is shared by everything on the same physical core. In fact, two hyperthreads even share the same physical register file.
>It turns out that mispredicting on purpose is difficult to optimize! It took a bit of work, but I found a variant that can leak about 30 kb per core, per second.
>This is fast enough to monitor encryption keys and passwords as users login!
TLDR: The vector registers this bug affects are used for string functions like strcmp, so anything could get loaded into them, including passwords.
It's stochastic; the attacker randomly gets data from whatever happens to be using the XMM/YMM/ZMM registers at the time. So if the attacker could eavesdrop in the background constantly, they might eventually see a password. Or they might be able to trigger some system code that processes your password, then eavesdrop for the next few milliseconds.
The attacker needs to run code on your machine. Unclear if running code in a web browser is sufficient or not. It requires an unusual sequence of machine instructions, which isn't necessarily possible in JS/WASM, but 'sounds' says they did it: https://news.ycombinator.com/item?id=36849767
The blog post discusses the discovery of a vulnerability called "Zenbleed" in certain AMD Zen 2 processors. It revolves around the use of AVX2 and the vzeroupper instruction, which zeroes upper bits in vector registers (YMM) to avoid dependencies and stalls during superscalar execution.
The vulnerability arises from a misprediction involving vzeroupper, which can be exploited with precise scheduling and triggering the XMM Register Merge Optimization, leading to a use-after-free-like situation. This allows attackers to monitor operations using vector registers, potentially leaking sensitive information like encryption keys and passwords.
The author found the bug through fuzzing and developed a new approach called Oracle Serialization to detect CPU execution errors during testing. The vulnerability (CVE-2023-20593) affects various AMD Zen 2 processors, but AMD released a microcode update to address the issue. For systems unable to apply the update, a software workaround exists by setting the DE_CFG[9] "chicken bit."
The post concludes with acknowledgments to individuals who contributed to the discovery and analysis of the Zenbleed vulnerability.