The Pentium FDIV bug, reverse-engineered
oldbytes.space
oldbytes.space
Except that he kept getting wrong answers on his homework.
And then he realized that when he did it on one of the unix machines, he got correct answers.
And then a few months later he realized why ....
Welcome to engineering. That's what we do. It's all we do.
"You know so much about computers, it's easy for you to figure it out!"
Anyway, as far as “learn the tools they want to teach the students how to use” goes, I dunno, hard to say. I wouldn’t be that surprised to hear that some department head got the wrong-headed idea that students needed to learn more practical tools and skills, shit rolled downhill, and some professor got stuck with the job.
Usually professors aren’t expert tool users after all, they are there for theory stuff. Hopefully nobody is expecting to become a vscode expert by listening to some professor who uses eMacs or vim like a proper grey-beard.
Reminds me of when I was able to climb aboard an old steam locomotive and spent a wonderful hour tracing all the pipes and tubes and controls and figuring out what they did.
A good engineering college doesn't actually teach engineering. They teach how to learn engineering. Then you teach yourself.
That aside, yes it's interesting that an old program from 1998 can still serve us quite well.
Never heard about it. BLAS, Octave, Scilab
I look forward to the promised proper write up that should be out soon.
How many Intel engineers does it take to change a light bulb? 0.99999999
Incidentally, Pentium M to Intel Core through 16th gen Lunarrow Lake all identify as P6 ("Family 6") for 686 because they are all based off of the Pentium 3.
> Intel used the Pentium name instead of 586, because in 1991, it had lost a trademark dispute over the "386" trademark, when a judge ruled that the number was generic.
I wonder how many yet to be discover silicone bugs are out there on modern chips?
I'm not sure when Intel started supporting microcode updates but I think it was much later.
- https://thechipletter.substack.com/p/intel-vs-nec-the-case-o...
- https://www.upi.com/Archives/1994/03/10/Jury-backs-AMD-in-di...
I would totally believe the FDIV bug is why Intel went to a patchable microcode architecture however. See “Intel P6 Microcode Can Be Patched — Intel Discloses Details of Download Mechanism for Fixing CPU Bugs (1997)” https://news.ycombinator.com/item?id=35934367
My guess -- and I hope you can confirm it at some point in the future -- is that more modern CPUs can patch other data structures as well. Perhaps the TLB walker state machine, perhaps some tables involved in computation (like the FDIV table), almost certainly some of the decoder machinery.
How does one make a patchable parallel multi-stage decoder is what I'd really like to know!
Related article: https://news.ycombinator.com/item?id=16058920
But if they had understood possible aftermath of non-tested block they would have implemented two blocks, and switch to older one if misworking was detected.
The microcode update would need to disable the entire FDIV instruction and re-implement it without using any floating point hardware at all, at least for the problematic devisors. It would be as slow as the software workarounds for the FDIV bug (average penalty for random divisors was apparently 50 cycles).
The main advantage of a microcode update is that all FDIVs are automatically intercepted system-wide, while the software workarounds needed to somehow find and replace all FDIVs in the target software. Some did it by recompiling, others scanned for FDIV instructions in machine code and replaced them; Both approaches were problematic and self-modifying code would be hard to catch.
A microcode update "might" have allowed Intel to argue their way out of an extensive recall. But 50 cycles on average is a massive performance hit, FDIV takes 19 cycles for single-precision. BTW, this microcode update would have killed performance in quake, which famously depended on floating point instructions (especially the expensive FDIV) running in parallel with integer instructions.
... until a new CPU support an instruction extension that use the same bit sequence.
Part of the reason why existing functionality is limited and nobody uses them to implement custom instructions, is that trapping is actually quite expensive.
The switch to kernel space and back is super expensive, but even if a CPU did implemente fast userspace trap handlers, it wouldn't be fast.
They're not fast, though. You probably don't want to run trapped fake instructions in tight loops, but you don't want that no matter how they're implemented.