I found a bug in Intel Skylake processors
gallium.inria.fr
gallium.inria.fr
Why seek the expert if not for his advice? It brings to mind people disregarding doctors who give them inconvenient medical advice.
I was tired of this problem, didn't know how to report those things (Intel doesn't have a public issue tracker like the rest of us), and suspected it was a problem with the specific machines at SIOU (e.g. a batch of flaky chips that got put in the wrong speed bin by accident).
were a doctor, he'd be guilty of malpractice. This bug went unreported eight months longer than it needed to. Am I misreading all this somehow?
The doctor metaphor isn't perfect; what I was going for is, when you are seeking out an expert's advice and you ignore it, why do you go to see the expert in the first place?
There's at least one obvious error in your statement: if an inconvenienced user's bug report results in less downtime for other users, it is a "gift" to other users, as well as a "gift" to the vendor.
But it says something about our profession if we regard putting flags down to mark the landmines we find a mere courtesy (a gift!) instead of an obligation. I guess that's a debate for a different time and place.
But just like a user has no obligation to 'mark the landmines' for vendors they also have no such obligation towards other users. They do have a right to receiving bug free software in the first place, alas our industry is utterly incapable of doing so which has lowered our expectations to the point where you feel that we have an actual obligation as users to become part of the debugging process.
That is not going to make our lives better.
What will make our lives better is if software producers accept liability for their crap they put out and if they were unable to opt-out of such liability through their software licenses and other legal trickery.
You're just a small step away from making it an obligation rather than an optional thing for users to report bugs, the only difference is that for you the obligation is a moral one rather than a legal one. I really do not subscribe to that, when I pay for something I expect it to work and I expect the vendor (and definitely not the other users) to work as hard as they can to find and fix bugs before the users do.
But we're 'moving fast and breaking shit' in the name of progress and part of that appears to extend to being in perpetual beta test mode. That's not how software should be built and I refuse to subscribe to this new world order where the end user is also the Guinea pig.
Keep in mind that users have their own work to do, are not on the payroll of the vendors usually have forked over cold hard cash in order to be able to use the code (ok, not in the case of open source) and tend to be less knowledgeable about this stuff than the vendors. They really should not have a role in this other than that they may - at their option - upgrade their software from time to time when told very explicitly what the changes are (and hopefully without pulling in a boatload of things that are good for the vendor but not for them).
I'd argue that a person does have that obligation in some circumstances, yes. And yes, I am thinking in moral rather than legal terms. The legal picture is pretty far outside my expertise, and the professional ethics of software engineering (which would in turn inform the legal picture) seems to be woefully opt-in. As you say, 'moving fast and breaking shit,' perpetual beta test mode, etc. So I'd put the legal stuff aside for now.
For me, the key is that "user" is a deceptive term here. A mere user cannot point to a small piece inside a much larger machine and say "that will blow up occasionally, and I know exactly when." We are talking about engineers. Or at least, I was thinking of the professional obligations of engineers - on the user side of the fence and the vendor side of the fence - and that was informing my comments.
> Keep in mind that users have their own work to do, are not on the payroll of the vendors usually have forked over cold hard cash in order to be able to use the code (ok, not in the case of open source) and tend to be less knowledgeable about this stuff than the vendors.
Yeah, and I don't think I disagree with you in the "user" case. I really think a software engineer finding a CPU bug is a different case. It seems me that if we're in possession of knowledge of something as serious and wide-reaching as a CPU bug, we have a reproducible test case, and we don't do anything with it (I mean, at least a tweet or something, for the love of God) we are part of the problem with our profession.
(1) a payment from the vendor to the reporter compensating them for time and effort spent at getting the bug to be reproduced
and, crucially,
(2) a requirement for all vendors of software and hardware to timely respond to bug reports and to have a standardized reporting process.
In that case I can see how such a shared responsibility would work, but as it is the companies get the benefits and the users get the hardship with a good portion of reported bugs (sometimes including a solution) that go unfixed, that's not a fair situation.
Case in point: I've reported quite a few bugs to vendors over the years but I've stopped doing it because in general vendors simply don't care, most of the time bug reports seem to result in a 'wont fix' or 'here is a paid upgrade for you with your fix in it'.
The difference between something like Google's bug bounty (capped at over $30k, I think) and a hypothetical bounty for Intel is, well, Intel has a lot more at stake. It's honestly strange that they don't have something in place already. Something like Skylake costs on the order of billions to get out there. It's cool that this Skylake bug was fixable via microcode, but the Pentium FPU bug back in the day cost them half a billion dollars. If such a bug exists, that is the kind of thing Intel should want to have reported as soon as humanly possible. Even the reputational damage they take from something milder like the Skylake bug would justify a bounty system with very serious payouts.
As technologists, we develop tools and services to capture bugs (both server side and client) so that we gain more insight into how our product operates. A user that takes time to give a well crafted bug report is rare. Most of the time it tends to be legitimate gripe that a feature isn't working.
Update: After writing I am re-reading your comment and now thinking we are on same page.
Do you have a step-by-step guide for doing that or are you just assuming that there must be some magic way?
2. That's not what he wrote
> Doctors are not obligated to write up case studies or submit them to medical publications.
That's true, although I think we'd all agree that a person who has the knowledge to create a lifesaving treatment for a disease and doesn't bother writing it down because, well, writing is boring, is behaving rather unethically.
But this is merely computer science, not medicine.
> The errata refers to the problem showing up on short loops of less than 64 instructions that use AH, BH, CH or DH.
> Looking at the Skylake microarch, the instruction decode queue is 128 uOps thread, 2*64 uOps when threaded. The Loop Stream Detector "can stream the same sequence of µOPs directly from the IDQ continuously without any additional fetching, decoding, or utilizing additional caches or resources." ... "capable of detecting loops up to 64 µOPs per thread". https://en.wikichip.org/wiki/intel/microarchitectures/skylak...
> So maybe the microcode update just shuts off the loopback detector.
https://groups.google.com/d/msg/comp.arch/UkO4Z2FT18c/7YlC0a...
So if the bug is in the loop-detector, and the patch possibly disables it rather than fixes it, then does anyone have any before-and-after performance stats?
No. Skylake does not have 128 μops with HT disabled. Skylake indeed was a big jump from Broadwell where the loopback buffer has 56 entries, 28 per hyperthread or 56 with HT off. Skylake has 64 μops per thread, HT on or off.
64 μops is a lot.
rr can often be a time-saver in situations by providing deterministic replays up to the point of a crash, whereas coredump analysis is a single retrospective snapshot.
Also a pretty good read and recently discussed here: https://news.ycombinator.com/item?id=14661473
My 2016 Skylake MBP used to crash very regularly when waking it up from suspend (sometimes multiple times a day). When I first heard about this issue a week ago, I used XCode Instruments to disable Hypterthreading.
I have not observed a single crash since.
It's of course not wrong, but using AH when you're dealing with RAX is a weird anachronism
Clang does the obvious, correct thing.
obvious, correct != obviously correctMore bluntly, before now wouldn't most people have said it was “obvious” that Intel would support their own documented features?
Most people here are not familiar with x86 assembly and its caveats it seems. Reading the Intel and AMD optimization manuals might be a good start
(and yes, the bug is not in GCC it's on the Intel processor)
Please don't make unsupported assertions that everyone but you is speaking out ignorance. It doesn't add anything to the conversation, especially when dealing with older codebases unless you can prove that this is and never has been the correct way to write that code. Otherwise it's just another way to say “CPU optimizations change over time and an open-source project doesn't have a team of experts tracking microbenchmarks to decide when to switch”.
> “CPU optimizations change over time and an open-source project doesn't have a team of experts tracking microbenchmarks to decide when to switch”
Yes, that's what's happening, but the problem comes from the P6 architecture (though Netburst doesn't have those problems), this problem has been known for around 20 years
Why would you expect GCC to optimize for Pentium 4 (Netburst) in this day and age? (Especially given that the article is talking about Skylake.)
> The quoted delay of 5 - 6 clocks is much better on later microarchitectures. For example from Sandy Bridge and Ivy Bridge, The Ivy Bridge inserts an extra μop only in the case where a high 8-bit register (AH, BH, CH, DH) has been modified
And you see that even on later architectures the point of avoiding AH makes sense (which is the opposite of what that GCC code does)
Making use of the "partial registers" (I see them more as separate smaller registers that can be grouped together) effectively can avoid many more instructions.
Optimizing for size is good, but what GCC did there made sense in the 32-bit days, but not that much today
Some code snippets use AH/AL as 2 separate registers hence the processor might rename them to different internal registers. But then when reading EAX the processor needs to update EAX accordingly as well.
Can we all stop shitting on GCC all the time? Thanks.
There is zero guarantee that the infamous Sufficiently Sophisticated Attacker couldn't predict it. Hardware is largely deterministic, even when it doesn't behave in the documented way. I wouldn't interpret this as literally unpredictable, it's just a generic slogan they always use in their errata. And they aren't going to say anything more for obvious reasons.
Patch this damn microcode.
Oh compilers. Like VC++6.0 initializing uninitialized memory to 0xCDCDCDCD in DEBUG.
They used to use, uh, more obvious patterns but the PC brigade called them on it so they settled on 0xcd.
Yea, I knew it was on purpose, but it had the unintended consequence of masking bugs that would only show up in release builds.
http://www.mathemainzel.info/files/x86asmref.html#int
Before Intel processors had execute protection, this was a good way to catch bugs in your buggy C bugs. I mean programs.
With VS .NET (VC7.0), the C++ compiler no longer defaulted uninitialized memory in debug builds.
Our software didn't just run on x86 (windows/linux), but also SGI MIPS (Irix), VAX (VMS), DEC Alpha (Tru64), AND PowerPC (VxWorks), so across all those compilers, we ended up having portable and robust code.
Really?
I've even seen, when developing standard cell libraries for a new fabrication process, bugs that occur because of unforeseen interactions between different semiconductor doping concentrations that occur when (due to pure statistics in fabrication) they overlap in the wrong way.
Reminded me of a developer for Crash Bandicoot who had seemingly random crashes: http://www.gamasutra.com/blogs/DaveBaggett/20131031/203788/M...
There is no other reason why a 64bit multi-core CPU developed in 2015, that makes heavy use of pipelining and other advanced and complicated code execution strategies, would need to support instructions that address the second-to-last byte of a register (eg. %ah) while keeping the rest of the register 'unchanged', which of course means making a complete mess of the code execution path.
The only reason this crap still exists is to keep Windows users' ability to run random EXE and DLL files from the 90s, if not random COM files from the 80s, at the expense of CPU cost, stability, and correctness for everyone else (such as the OCaml developers and users who ran into this bug.)
It's lazy to the point of dishonesty to act as if Microsoft is the only one with decades of accumulated code.
Why mess up a stable API when you don't have to?
int f(int x) { return x | 256; }
Gcc generates the following code: movl %edi, %eax
orb $1, %ah
ret
In contrast, clang uses an orl instead of orb. The advantage of using orb is that it generates shorter code with an immediate operand. This is on an x86_64 architecture and does not involve 32-bit code.During the move to AMD64, AMD decided to remove many features of the old architecture (for example, removing many one-byte instructions that were taking up valuable opcode space). However, they decided not to touch the AH/BH/CH/DH register access modes - yet it was entirely in their power to do so because they were developing a new architecture with no existing code. So, if you're going to blame someone for this, blame AMD for wanting to keep too many legacy features in what was designed to be a brand new architecture.
For examples of a 64-bit architecture that aren't simply extensions of the 32-bit versions, look at AArch64 (compared with the classic ARM instruction set), or Itanium (Intel's ultimately-unsuccessful attempt to build a brand-new 64-bit architecture cleanly broken from the legacy 32-bit set).
If you think stability and correctness are bad now, try compiling random codebases from the 90s with a motley melange of random people using random compiler versions - some prerelease, some stable, some unpatched and broken - many of which have been just as if not more aggressive than CPU vendors with regards to things like "undefined behavior" in C and C++ codebases.
Yes, microcode bugs suck. No, getting rid of x86 won't eliminate them. No, getting rid of QAed, tested, and sometimes disassembly-verified binaries (for such things as verifying fixed duration crypto operations) in favor of compiling with any and every C compiler under the sun isn't going to improve stability and correctness. No, I really don't want to go through the bother of installing configuring and fixing your build chain either.