Reptar
lock.cmpxchg8b.com
lock.cmpxchg8b.com
(via https://news.ycombinator.com/item?id=38268043, but we merged the comments hither)
> Prefixes allow you to change how instructions behave by enabling or disabling features
Why do we need “prefixes” to disable or enable features? Is this for dynamically toggling feature so you don’t have to go into BIOS?
[1] https://en.wikipedia.org/wiki/Intel_8086#The_first_x86_desig...
It's interesting that ASCII is transparently just a bunch of control codes for a physical printer/typewriter, combining things like "advance the paper one line", "advance the paper one inch", "reset the carriage position", and "strike an F at the carriage position", all of which are different mechanical actions that you might want a typewriter to do.
But now we have Unicode, which is dedicated to the purpose of assigning ID numbers to visual glyphs, and ASCII has been interpreted as a bunch of glyph references instead of a bunch of machine instructions, and there are the control codes with no visual representation, sitting in Unicode, being inappropriate in every possible way.
It's kind of like if Unicode were to incorporate "start microwave" as part of a set with "1", "2", "3", etc.
The endless CR/LF/CRLF line ending problem would have been solved if the RS (Record Separator) ASCII code was used instead of the physical CR = carriage return, ie move print head back to start of line, and LF = line feed, ie rotate paper up one line.
But Unix decided on LF, Apple used CR, Windows used CRLF, and even today, I had to get a guy to stop setting his system to "Windows" because he was screwing up a git repo with extraneous CRs.
The REP prefixes are the most common; they just let you perform the same instruction a variable number of times. It looks in the CX register for the count. This makes many common loops really, really short, especially for moving objects around in memory. The memcpy function is often inlined as a single REP MOVS instruction, possibly with an instruction to copy the count into CX if it isn’t already there.
I suppose the REX (operand size) prefix is pretty common too, since 64–bit programs will want to operate on 64–bit values and addresses pretty frequently.
None of the prefixes toggle things that can be set globally, by the BIOS or otherwise. They all just specify things that the next instruction needs to do.
After all of those, REP is not that uncommon of a prefix to run into, although many people prefer SIMD memcpy/memset to REP MOVSB/REP STOSB. It is slightly unusual.
TL;DR Weeee! Intel machine language is crazy!
mov BYTE PTR [rdi+rbx*4],0x4
SIB is determined by the register indices of rdi, rbx, and 4, all right there in the instruction. Likewise, Mod R/M encodes the addressing mode, which is clear from the operands in the assembler listing. Though x86 is such as mess that there are cases where you can encode the same instruction in either a Mod R/M form or a shorter form, eg PUSH/POP.
REX is a prefix, but it is a bit special as it must be the last one, and repeats are undefined. It is not elided because of commonality but because its presence and value is usually implied from the operands, it is therefore redundant to list it.
For instance, PUSH R12 must use a REX prefix (REX.B with the one byte encoding).
REP prefixes are pretty rare. Depending on compiler they are used rarely, for a few specific operations (like rep movsd for memcpy) or usually never.
Most common prefixes are by far REX prefixes in 6h4 64bit assembly (don't believe me? Look at the 64bit code in vim and see all those `H` letters around. It's REX.W). Segment override prefixes are another class of prefixes that are used in handwritten assembly (in bootloader or special runtime functions) but almost never used by compilers.
In older code, most common prefixes are 0x66h (it doesn't even has opcode, there's no way to emit it directly) and maybe 0x67h.
They are all used to modify some aspect of the next executing instruction, for example the "default" register size is 32bit, but you can change it with a prefix. I think an example will help.
$ echo 31c0 | xxd -r -ps | ndisasm -b64 /dev/stdin 00000000 31C0 xor eax,eax $ echo 6631c0 | xxd -r -ps | ndisasm -b64 /dev/stdin 00000000 6631C0 xor ax,ax $ echo 4831c0 | xxd -r -ps | ndisasm -b64 /dev/stdin 00000000 4831C0 xor rax,rax
And there are also the extensions of base, SIB and Mod/RM fields, mentioned by the article (to the extended set of high registers):
$ echo 4431c0 | xxd -r -ps | ndisasm -b64 /dev/stdin 00000000 4431C0 xor eax,r8d
$ echo 4131c0 | xxd -r -ps | ndisasm -b64 /dev/stdin 00000000 4131C0 xor r8d,eax
So rarely-used addressing modes get a "segment prefix" that causes them to use a segment other than DS. Or x86_64 added a "REX" prefix that added more bits to the register fields allowing for 16 GPRs. Likewise the "LOCK" prefix (though poorly specified originally) causes (some!) memory operations to be atomic with respect to the rest of the system (c.f. "LOCK CMPXCHG" to effect a compare-and-set).
All these things are operations other CPU architectures represent too, though they tend to pack them into the existing instruction space, requiring more bits to represent every instruction.
Notably the "REP" prefix in question turns out to be the one exception. This is a microcoded repeat prefix left over from the ancient days. But it represents operations (c.f. memset/memmove) that are performance-sensitive even today, so it's worthwhile for CPU vendors to continue to optimize them. Which is how the bug in question seems to have happened.
The latter is probably more suited to you unless you wanna go on a dive into computer architecture itself. There's older editions available for free (authorized by the authors) on the web.
I first read the 3rd edition of Computer Architecture and besides being one of the most clear textbooks I've ever read it vastly improved my understanding of what's going on in there in relation to OoO speculative execution, etc.
I think for such meaningless titles that the poster should add a small description
HN goes for a middle ground that promotes intellectual curiosity and link clicking. If you refuse to click the link for obscure titles at least you’re stuck replying to those who did click the link and that’s still better than what we have on the rest of the internet.
Submissions that don’t have the payoff to justify more obscure, whimsical titles fall off the first page unlike this one.
Does anyone know if AMD CPUs are affected?
Actually, come to think of it, my hunch is that the uOP decoder (presumably in hardware) is actually fine and that the microcoded optimized copy routine is trying to infer things about the uOP stream that just aren't true --- "Oh, this is a rep mov, so of course I need to go backward two uOPs to loop" or something.
I expect Intel's CPU team isn't going to divulge the details though. :-)
Are these just CPUID flags that tell you that you can use a rep movsb for maximum performance instead of optimized SSE memcpy implementations? Or is it a special encoding/prefix for rep movsb to make it faster? In case of the later, why would that be necessary? How does one make use of fsrm?
Seems like ERMS was a cheaper replacement for AVX and FSRM was a better version, for shorter blocks.
> Cheapest versions of later processors - Kaby Lake Celeron and Pentium, released in 2017, don't have AVX that could have been used for fast memory copy, but still have the Enhanced REP MOVSB. And some of Intel's mobile and low-power architectures released in 2018 and onwards, which were not based on SkyLake, copy about twice more bytes per CPU cycle with REP MOVSB than previous generations of microarchitectures.
> Enhanced REP MOVSB (ERMSB) before the Ice Lake microarchitecture with Fast Short REP MOV (FSRM) was only faster than AVX copy or general-use register copy if the block size is at least 256 bytes. For the blocks below 64 bytes, it was much slower, because there is a high internal startup in ERMSB - about 35 cycles. The FSRM feature intended blocks before 128 bytes also be quick.
[1] https://stackoverflow.com/a/43837564
[2] http://www.intel.com/content/dam/www/public/us/en/documents/...
Choosing an optimal instruction choice and scheduling can be done statically during compile time or dynamically (via chosing one of several library functions at runtime, or jitting).
In order to be able to detect which is the optimal instruction scheduling at runtime you need to know the actual CPU. You could have a table of all cpu models or you could just ask your OS whether the CPU you run on has that optimization implemented.
Linux had to be patched so that it can _report_ that a CPU does implement that optimization.
Intel would like to thank Intel employees:[...] for finding this issue internally.
Intel would like to thank Google Employees: [...] for also reporting this issue.
[1] https://www.intel.com/content/www/us/en/security-center/advi...
So as described, this isn't a "valuable" bug.
The space of theoretically-very-bad attacks is much larger than practical ones people will pay for, c.f. rowhammer.
Intel knows exactly how their ROB works.
Therefore Intel knows the possible consequences of this bug and how to trigger them.
If there is a privilege execution path from this, Intel knows. And anyone Intel chose to share it with knew.
Thankfully, since it's public now, the value of that decreases and customers can begin to mitigate.
No, or at least not yet. I mean, I've written plenty of bugs. More than I can count. How many of them were genuine security vulnerabilities if properly exploited? Probably not zero. But... I don't know. And I wrote the code!
Particularly since the MCEs triggered could prevent an automatic reboot. Would depend what the hardware management system did - do machines presenting MCEs get pulled?
So you go and find yourself a thousand cheap / free tier accounts, spin up an instance in a few regions each, and boom, you've taken out 10k physical hosts. And run it in a lambda at the same time, and see how well the security mechanisms identify and isolate you.
Causing a near simultaneous reboot of enough hosts is likely to take other parts of the infrastructure down.
e.g. the GRU
Stolen credit cards are a dime a dozen, and nation state actors can just use their domestic banks or agents in the banks of other countries in a pinch to deflect blame or lay false trails.
If I were Russia or China, I'd invest a lot of money into researching all kinds of avenues on how to take out the large three public cloud providers if need be: take out AWS, Google, Microsoft and on the CDN side Cloudflare and Akamai and suddenly the entire Western economy grinds to a halt.
The only ones who will not be affected are the US government cloud services in AWS, as this runs separate from other AWS regions - that is, unless the attacker gets access to credentials that allow them executions on the GovCloud regions...
But I don't believe it. People are not that stupid.
Why target the management plane? Fire off payloads to take down the physical VM hosts and suddenly any cloud provider has a serious issue because the entire compute capacity drops.
This subthread started with "is this issue a valuable exploit". Needless to say, if you need to invoke superpower-scale cyber warfare to find an application, the answer is "no". Russia and China have plenty of options to "take out" western infrastructure if they're willing to blow things up[1] at that scale.
[1] Figuratively and literally
See, e.g., https://madsciblog.tradoc.army.mil/156-what-is-the-threshold...
> responses are usually proportional to and in the same domain as the provocation
Which is both good and bad at the same time. Cyber warfare has been significantly impacting our economies and our citizens - anything from scam callcenters over ransomware to industrial espionage - to the tune of many dozens of billions of dollars a year. And yet, no Western government has ever held the bad actors publicly accountable, which means that they will continue to be a drain on our resources at best and a threat to national security at worst (e.g. the Chinese F-35 hack).
I mean, I'm not calling for nuking Bejing, that would be disproportionate - but even after all that's happened, Russia and China are still connected to the global Internet, no sanctions, nothing.
some bored kid with a couple of hundred stolen credit cards can bring down a significant chunk of AWS/GCP/...
DoS isn't as lucrative as other things; I assume that most state actors would far prefer to find a way to turn this into a privilege escalation. But being able to possibly take out a cloud provider for a while is still monetizable.
It’s anything but a minor bug and anyone that says so clearly hasn’t worked with CPUs
> Sequence of processor instructions leads to unexpected behavior for some Intel(R) Processors may allow an authenticated user to potentially enable escalation of privilege and/or information disclosure and/or denial of service via local access.
Not sure if fully related, but possibly.
> This fact is sometimes useful; compilers can use redundant prefixes to pad a single instruction to a desirable alignment boundary.
so I imagine that could happen under the right optimization mode.
Is this still the case or modern intel CPUs optimize out the nop in the frontend decoder?
"Looking in the old AMD optimisation guide for the then-current K8 processor microarchitecture (the first implementation of 64bit x86!), there is effectively mention of a “Two-Byte Near-Return ret Instruction”.
The text goes on to explain in advice 6.2 that “A two-byte ret has a rep instruction inserted before the ret, which produces the functional equivalent of the single-byte near-return ret instruction”.
It says that this form is preferred to the simple ret either when it is the target of any kind of branch, conditional (jne/je/...) or unconditional (jmp/call/...), or when it directly follows a conditional branch.
Basically, when the next instruction after a branch is a ret, whether the branch was taken or not, it should have a rep prefix.
Why? Because “The processor is unable to apply a branch prediction to the single-byte near-return form (opcode C3h) of the ret instruction.” Thus, “Use of a two-byte near-return can improve performance”, because it is not affected by this shortcoming."
...
" If a ret is at an odd offset and follows another branch, they will share a branch selector and will therefore be mispredicted (only when the branch was taken at least once, else it would not take up any branch indicator %2B selector). Otherwise, if it is the target of a branch, and if it is at an even offset but not 16-byte aligned, as all branch indicators are at odd offsets except at byte 0, it will have no branch indicator, thus no branch selector, and will be mispredicted.
Looking back at the gcc mailing list message introducing repz ret, we understand that previously, gcc generated: nop, ret
But decoding two instructions is more expensive than the equivalent repz ret.
The optimization guide for the following AMD CPU generation, the K10, has an interesting modification in the advice 6.2: instead of the two byte repz ret, the three-byte ret 0 is recommended
Continuing in the following generation of AMD CPUs, Bulldozer, we see that any advice regarding ret has disappeared from the optimization guide."
TLDR: Blame AMD K8! First x64 CPU. This GCC optimization is outdated and should only be used when specifically optimizing for K8.
Maybe I should start hacking while watching Teenage Mutant Ninja Turtles.
When a nop truly is necessary you will see compilers and performance engineers add prefixes to the nop to make it the desired size.
The more proximate cause is that some instructions with multiple redundant prefixes (which is legal, but pointless) have their length miscalculated by some Intel CPUs, which results in wrong outcomes.
ARM 64 gets this right, with fixed length 32 bit instructions.
At the expense of code density, yet RISC-V is easy to decode, with implementations going up to 12-way decode (Veyron V2) despite variable length.
ARM64 hardly "gets it right".
Higher code density is valuable. E.g.:
- The decoders can see more by looking at a window of code of the same size, or we can have a narrowed window.
- We can have less cache and save area and power. We can also clock the cache higher, enabled by it being smaller, lowering latency cycles.
- Smaller binaries or rom image.
Soon to be available (2024) large, high performance implementations will demonstrate RISC-V advantages well.
If you read the manual, Intel encourages minor variations of the `nop` instructions that can be lengthened into different number of bytes (like `nop dword ptr [eax]` or `nop dword ptr [eax + eax*1 + 00000000h]`).
It is never recommended anywhere in my knowledge to rely on redundant prefixes of random non-nop instructions.
It's a pretty old and well known technique:
https://stackoverflow.com/questions/48046814/what-methods-ca...
Note that this technique is really only legitimate where the used prefix already has defined behavior with the given instruction ("Use of repeat prefixes and/or undefined opcodes with other Intel 64 or IA-32 instructions is reserved; such use may cause unpredictable behavior."), and of course the REX prefix has special limitations. The key is redundant, not spurious. It is not a good idea to be doing rep add for example. But otherwise, there is no issue.
Using specialized prefixes wastes encoding space for no real gain. You realize on most common processors NOP itself is a pseudo-instruction? Even the apparently meme-worthy (see sibling comment) RISC-V, it's ADDI x0, x0, 0.
> Moving a register to itself is functionally a nop, but the processor overloads it to signal information about priority.
https://devblogs.microsoft.com/oldnewthing/20180809-00/?p=99...
What does this even mean? How can a program do this when thread priority is an OS thing? It's seems just weird.
The issue here is their verification of possible internal CPU states didn't account for this one.
(There is, perhaps, an argument to be made that the x86 architecture has become so complex that the emulator between its embarrassingly stupid PDP-11-style single-thread codeflow and the embarrassingly parallel computation it does under the hood to give the user more performance than a really fast PDP-11 cannot be reliably tested to exhaustion, so perhaps something needs to give on the design or the cost of the chips).
Itanic would like to object! Unfortunately it can’t get through the door.
Or at least put a practical bound on how many bits per second at most you can from any such side channel (the reasoning being, if you can get at most a bit for each million years, you probably don't have an attack)
Then you verify if a given design meets this constraint
Theoretically you can do this from software down to (idealized) gates, but in practice the effort is so great that it's only been done in extremely limited systems.
Do you have some link about the designed CPU?
Is the future leads to a swarm of disconnected A55
cores each running a single application?
don't you dare tease me like thatYes, of course. But we'd have to put actual effort in, and realistically people wouldn't pay enough extra to make it worthwhile.
Most likely dedicated resources on demand will be the future. Some companies already offer it.
(I think there is a limit on microcode, they seem conservative to release new ones - I don't remember the details)
That said, I'm sure there's some verification framework like SPARK for VHDL, and this feels like exactly the kind of thing it should catch.
[1] https://www.academia.edu/60937699/The_IMS_T_800_Transputer
...our validation pipeline produced an interesting assertion...
What is a validation pipeline?> I’ve written previously about a processor validation technique called Oracle Serialization that we’ve been using. The idea is to generate two forms of the same randomly generated program and verify their final state is identical.
They also could be just discarding any where it runs for longer than X time, or a bunch of other possibilities.
Intel added two features to modern x86 chips that detects rep movsb and accelerates it to be as fast as those other ways. However, those features have a bug. You see, because rep is a prefix byte, you can just keep adding more prefix bytes to the instruction (up to a maximum of 16 AFAIK). x86 has other prefix bytes too, such as rex (used to access registers 8-16), vex, evex, etc. The part of the processor that recognizes a rep movsb does NOT account for these other prefix bytes, which makes the processor get confused in ways that are difficult to understand. The processor can start executing garbage, take the wrong branch in if statements, and so on.
Most disturbingly, when multiple physical cores are executing these "rep rep rep rep movsb" instructions at the same time, they will start generating machine check exceptions, which can at worst force a physical machine reboot. This is very bad for Google because they rent out compute time to different companies and they all need to be able to share the same machine. They don't want some prankster running these instructions and killing someone else's compute jobs. We call this a "Denial of Service" vulnerability because, while I can't read someone else's computations or change them, I can keep them from completing, which is just as bad.
Do they ? As these issues keep piling up, it just seems that it's not worth the hassle, and they should instead never do sharing like this...
If you ever download untrustworthy code and run it in a VM to protect your main set of data, that's another case.
The success of cloud computing is from the idea that multiple people can share the same computer. You only need one core, but CPUs come with 128, but with the cloud you can buy just that one core and share 1/128th of the power supply, rack space, motherboard, ethernet cable, sysadmin time, etc. and that reduces your costs. That assumption is all based on virtualization working, though; nobody wants 1/128th of someone else's computer, they want their own computer that's 1/128th as fast. Bugs like these demonstrate that you're just sharing a computer with someone, which is bad for the business of cloud providers.
That said, machines are really monstrously huge these days, and it can be hard to put them to good use. You also miss out on cost savings like burstable instances, which rely on someone else using the capacity for the 16 hours a day when you don't need it. It's a balance, but I'd say "just buy a computer" would be my starting point for most application deployments.
[0] Based on the cost to rent an EC2 Dedicated Host (a1 family). See https://aws.amazon.com/ec2/dedicated-hosts/pricing/
# Google’s commitment to collaboration and hardware security
## As Reptar, Zenbleed, and Downfall suggest, computing hardware and processors remain susceptible to these types of vulnerabilities. This trend will only continue as hardware becomes increasingly complex. This is why Google continues to invest heavily in CPU and vulnerability research. Work like this, done in close collaboration with our industry partners, allows us to keep users safe and is critical to finding and mitigating vulnerabilities before they can be exploited.
There's a tension between the NSA wanting backdoors and service providers (CPU designers + Cloud hosting) wanting secure platforms. It's possible that by employing CPU and security researchers, Google can tip the scales a bit further in their favor.
Acid Burn: What oil spills?
Lord Nikon: Yo, brain dead, today's the 13th.
Cereal Killer: Whoa, this hasn't happened yet!
Credit where credit is due: Google has some of the best codenames.
> This bug was independently discovered by multiple research teams within Google, including the silifuzz team and Google Information Security Engineering.
>A potential security vulnerability in some Intel® Processors may allow escalation of privilege and/or information disclosure and/or denial of service via local access.
https://www.intel.com/content/www/us/en/security-center/advi...
(As of this writing, this post has more votes, the other has more comments)
"We know something strange is happening, but how microcode works in modern systems is a closely guarded secret."
My question: How likely is it that this is an intentional bug door that was added into the microcode by Intel and its government partners?
I don't know enough about microcode and CPU's to be able to answer this myself, so backed-up opinions welcome!
This isn't how anyone would backdoor a CPU. An actual backdoor would be done via some instruction sequence that is basically impossible to trigger by accident and hard to detect even when triggered.
One is to make the condition for the backdoor trigger based on multiple (unlikely) instructions in sequence. This bug was triggered by a single instruction, so it would have been a pretty easy case for fuzzing. If you need a sequence of 10 specific instructions in a specific sequence, with no kind of observable side-effects for getting just the first 9 right so that nobody can do a guided search? That's not going to be found just by random chance. It doesn't matter what those instructions are, as long as they're not something that would get generated by real compilers on real programs.
The other is to make it dependent on the data rather than just the static instructions. Like, what if you had the SHA1 acceleration instructions trigger a backdoor iff the output of the hash is a certain value? You could probably even arrange for the backdoor to get triggered from managed and sandboxed runtimes like Javascript, rather than needing to get the victim to run native code. And somebody triggering this by accident would be equivalent to a SHA1 preimage collision.
In a duopoly market there seems to be no real competition. And yes I know that some (not all) bugs also happen for AMD.
Some of these novel side-channel attacks actually even apply in completely unrelated architectures such as ARM [1] or RISC-V [2].
I think the problem is not (just) a lack of competition (although you're right that the duopoly in desktop/laptop/non-cloud servers for x86 brings its own serious issues, I've written and ranted more often than I can count [3]), it rather is that modern CPUs and SoCs have simply become so utterly complex and loaded with decades worth of backwards-compatibility baggage that it is impossible for any single human, even a small team of the best experts you can bring together, to fully grasp every tiny bit of them.
[1] https://www.zdnet.com/article/arm-cpus-impacted-by-rare-side...
[2] https://www.sciencedirect.com/science/article/pii/S004579062...
[3] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
For now, AI lacks the contextual depth - but an AI that can actually design a CPU from scratch (and not just rehashing prior-art VHDL it has ... learned? somehow), if that happens we'll be at a Cambrian Explosion-style event anyway, and all we can do is stand on the sides, munch popcorn and remember this tiny quote from Star Wars [1].
Possible? Yes. But far less likely.
Complexity carries over and breeds bugs. RISC-V is an order of magnitude simpler than ARM64, which in turn is an order of magnitude simpler than x86.
And it is so w/o disadvantage[0], positioning itself as the better ISA.
This seems like a "simple" bug of the type that people write every day, not deep architectural problems like Spectre and the like, which also affected AMD (in roughly equal measure if I recall correctly).
The reason why Meltdown has a more dramatic name than Spectre, despite being the same vulnerability, is that hardware privilege boundaries are the only defensible boundary against timing attacks. We already expect context switches to be expensive, so we're allowed to make them a little more expensive. It'd be prohibitively expensive to avoid leaking timing from, say, one executable library to a block of JIT compiled JavaScript code within the same browser content process.
[0] https://randomascii.wordpress.com/2018/01/07/finding-a-cpu-d...