Intel's $475M error: the silicon behind the Pentium division bug
righto.com
righto.com
My Mastodon thread about the bug was on HN a few weeks ago, so this might seem familiar, but now I've finished a detailed blog post. The previous HN post has a bunch of comments: https://news.ycombinator.com/item?id=42391079
For the Pentium, I'm curious about the FPU value stack (or whatever the correct term is) rework they did. It's been a long time, but didn't they do some kind of early "register renaming" thing that had you had to manually manage doing careful fxchg's?
IIRC fadd and fmul were both 3/1 (three cycles latency, one cycle throughput), so you'd start an operation, use the free fxch to get something else to the top, and then do two other operations while you were waiting for the operation to finish. That way, you could get long strings of FPU operations at effectively 1 op/cycle if you planned things well.
IIRC, MSVC did a pretty good job of it, too. GCC didn't, really (and thus Pentium GCC was born).
Compilers got pretty good at optimizing straight line math but were not as good at cases where variables needed to be kept in the stack during a loop, like a running sum. You had to get the order of exchanges just right to preserve stack order across loop iterations. The compilers at the time often had to spill to memory or use multiple FXCHs at the end of the loop.
Huh, are you sure? Do you have any documentation that clarifies the rules for this? I was under the impression that something like `FMUL st, st(2) ; FXCH st(1), FMUL st, st(2)` would kick off two muls in two cycles, with no stall.
You can immediately overlap with a FADD.
How hard is it to "dump" the microcode into a bitstream? Could it be done programatically from high resolution die photographs? Of course, I appreciate that's probably the easy part in comparison to reverse engineering what the bitstream means.
> By carefully examining the PLA under a microscope
Do you do this stuff at home? What kind of equipment do you have in your lab? How did you develop the skills to do all this?
I do this stuff at home. I have an AmScope metallurgical microscope; a metallurgical microscope shines light down through the lens, rather than shining the light from underneath like a biological microscope. Thus, the metallurgical microscope works for opaque chips. The Pentium is reaching the limits of my microscope, since the feature size is about the wavelength of light. I don't have any training in this; I learned through reading and experimentation.
I suppose you might be able to get slightly better resolution using a shorter wavelength, but at that point, it requires a lot of technical skill and environmental conditions and time and money, Just getting to the point you've reached (and knowing what the limitations are) can be satisfying in itself.
I never realised this is how floating point division can be implemented. Actually funny how I didn't realise that multiple integer division steps are required to implement floating point division :-)
In hindsight one could wonder why the unused parts of the lookup table were not filled with 2 and -2 in the first place.
To contrast, I’ve been thinking a lot about the Amazon Colorsoft launch, which had a yellow band graphics issue on some devices (mine included). Amazon waited a bit before acknowledging it (maybe a day or two, presumably to get the facts right). Then they simply quietly replace all of them. No recall. They just send you a new one if you ask for it (mine replacement comes Friday, hopefully it will fix it). My takeaway is that it’s pretty clear that having an incredibly robust return/support apparatus has a lot of benefits when launches don’t go quite right. Certainly more than you’d expect from analysis.
Similarly I haven’t seen too many recent reports about the Apple AirPod Pros crackle issue that happened a couple years ago (my AirPods had to be replaced twice), but Apple also just quietly replaced them and the support competence really seemed something powerful that isn’t always noticed.
Colorsoft: https://www.tomsguide.com/tablets/e-readers/amazon-kindle-co...
AirPods Pro: https://support.apple.com/airpods-pro-service-program-sound-...
On the Apple side the iPhone 4 antennagate is a better comparison since the equivalent fix there would have involved free replacements for a flagship and revenue-critical product which Apple did not offer.
Intel on the other hand did eventually offer free replacements for anybody who asked and took a major financial hit.
Anecdata ofc, but everyone I know already held phones in fingers back then, rather than hugging it as a brick.
The media coverage and the fact that "computer can't divide" is something that the public could wrap their heads around is what made the recall unavoidable.
Intel's own marketing hype around the Pentium has played into it too. It would have been a smaller deal during the 486 era.
https://www.latimes.com/archives/la-xpm-1994-12-14-ls-8729-s...
> Why didn’t Intel call the Pentium the 586? Because they added 486 and 100 on the first Pentium and got 585.999983605 .”
Before anyone well actually’s me, yes they did come out with a separate CDMA iPhone 4 for Verizon where they changed the antenna design
I really respected Apple’s commitment to standing behind their product in that way.
I've been in some consumer Apple "shadow warranty" situations, so I know what you are talking about, but IMO very different than the "IT crisis" that intel was facing. "IBM said so" had a ton of IT weight back then.
> However, IBM performed their own analysis,29 suggesting that the problem could hit customers every few days.
I bet these aren’t as far off as they seem. Intel seems to be considering a single user, while I suspect IBM is thinking in terms of support calls.
This is a problem I’ve had at work. When you process a 100 million requests a day the one in a billion problem is hitting you a few times a month. If it’s something a customer or worse a manager notices, they ignore the denominator and suspect you all of incompetence. Four times a month can translate into “all the time” in the manner humans bias their experiences. If you get two statistical clusters of three in a week someone will lose their shit.
The other failure mode that occurred to me is that if a spread sheet is involved you could keep running the same calc on a bad input for months or even years when aggregating intermediate values over units of time. A problem that happens every time you run a calculation is very different from one that happens at random. Better in some ways and worse in others.
E.g.: Adding a row doesn't invalidate calculations for previous rows in typical spreadsheet usage. The bug is deterministic, so repeating successful calculations over and over with the same numbers won't ever trigger the bug.
I recall a study done years ago where students were supplied calculators for their math class. The calculators had been doctored to produce incorrect results. The researchers wanted to know how wrong the calculators had to be before the students noticed something was amiss.
It was a factor of 2.
Noticing the error, and being affected by the error, are two entirely different things.
I.e. how many people check to see if the computer's output is correct? I'd say very, very, very few. Not me, either, except in one case - when I was doing engineering computations at Boeing, I'd run the equations backwards to verify the outputs matched the inputs.
Only somewhat true. Take any consumer usage here for example. If you're playing a game and it hits this incorrect output but you don't notice anything as a result, were you actually affected?
How much usage of FDIV on a Pentium was for numerically significant output instead of just multimedia?
But if you're doing financial work, scientific work, or engineering work, the results matter. An awful lot of people used Excel.
BTW, telling a customer that a bug doesn't matter doesn't work out very well.
Which is to say, it will depend a lot on the context and the understanding of the person doing the calculation.
I.e. Intel's problem became my problem, grrrr
I AM PENTIUM OF BORG.
DIVISION IS FUTILE.
YOU WILL BE APPROXIMATED.I hate how this sentence makes me feel.
It's possible some other result, likely aligned to an easy binary multiple would still produce a square block of 2, and that allowing the far edges to float to some other value could yield a slightly more compact logic array. Back-filling the entire side to the clamped upper value doesn't cost that much more though, and is known to solve the issue. As pointed out elsewhere, that sort of solution would also be faster for engineering time, fit within the planned space budget, and best of all reduces conative load. It's obviously correct when looking at the bug.
The person generating the table didn't realize filling the out-of-bounds with two would make for a simpler PLA. And the person squishing the table into the PLA didn't realize the zeros were "don't care" and assumed they needed to be preserved.
It's also possible they simply stopped optimizing as soon as they felt the PLA was small enough for their needs. If they had already done the floorplanning, making the PLA even smaller wasn't going to make the chip any smaller, and their engineering time would be better spent elsewhere.
The other thing that's hard for me to believe is there wasn't an extensive and mostly automated QA process that would test absolutely every little feature of this CPU.
This sounds utterly insane. You are making a CPU, if any calculations are wrong it needs to be fixed ?? I supposed this only came to light very late into testing and it was very impractical to bin every cpu, so they rolled the dice.
I believe this is because for any adder you always want 1 bit extra to detect overflow! This is why 9 bit adders are a common component in MCUs
I wonder why they didn't do this in the first place.
Look at it again later, someone asks why not just fill everything in instead and everyone feels a bit silly XD.
At the time, Intel was one of the leading ARM SoC providers, their custom XScale ARM cores were faster than anything from ARM Inc themselves. It was the perfect line of chips for smartphones.
The MBA types at Intel ran some sales projects and decided that such a chip wasn't likely to be profitable. There was apparently debate within Intel, the engineering types wanted to develop the product line anyway, and others wanting to win good-will from Apple. But the MBA types won. Not only did they reject Apple's request for an iPhone SoC, but they immediately sold off their entire XScale division to marvel (who did nothing with it) so they wouldn't even be able to change their mind later even if they wanted.
With hindsight, I think we can safely say Intel's projections for iPhone sales were very wrong. They would have easily made their money back on just the sales from the first-gen iPhone, and Apple would probably gone back to intel for at least a few generations. Even if Apple dumped them, Intel would have a great product to sell to the rapidly market of Android smartphones in the early 2010s.
-----------
But I think it's actually far worse than just Intel missing out on the mobile market.
In 2008, Apple acquired P.A. Semi, and started work on their own custom ARM processors (and ARM SoCs). The ARM processors which Apple eventually used to replace Intel as suppler in laptops and desktops too.
Maybe Apple would have gone down that path anyway, but I really suspect Intel's reluctance to work with Apple to produce the chips Apple wanted (especially the iPhone chip) was a huge motivating factor that drove Apple down the path of developing their own CPUs.
Remember, this is 2006. Intel had only just switched to Intel in January because IBM had continually failed to deliver Apple the laptop-class powerpc chips they needed [1]. And while at that time, Intel had a good roadmap for laptop-class chips, it would have looked to Apple as if history was at risk of repeating itself, especially as they moved into the mobile market where low power consumption was even more important.
[1] TBH, IBM were failing to provide desktop-class CPUs too. But the laptop cpus were the more pressing issue. Fun fact: IBM actually tried to sell the PowerPC core they were developing for the xbox 360 and PS3 to Apple as a low-power laptop core. It was sold to Microsoft/Sony as a low-power core too, but if you look at the launch versions of both consoles, they run extremely hot, even when paired with comically large (for the era) cooling solutions.
This isn’t strictly true. Tony Fadell and one of t- the creator of the iPod and considered co-creator of the iPhone - said in an interview with Ben Thompson (Stratechery) that Intel was never seriously in the running for iPhone chips.
Jobs wanted it. But the technical people at Apple pushed back.
Besides, especially in 2006 less than a year before the iPhone was introduced, chip decisions had already been made.
You can see today with modern Ryzen laptop chips that aren’t that much worse than ARMs fabbed with the same node on perf/watt.
The Apple cores themselves do not have great performance for array operations, but when considering the CPU cores together with the shared SME/AMX accelerator, the aggregate might have a good performance per area and per power consumption, but that cannot be known with certainty, because Apple does not provide information usable for comparison purposes.
The comparison is easy only with the cores designed by Arm Holdings. For array operations, the best performance among the Arm-designed cores is obtained by Cortex-X4 a.k.a. Neoverse V3. Cortex-A720 and Cortex-A725 have half of the number of SIMD pipelines but more than half of the area, while Cortex-X925 has only 50% more SIMD pipelines but a double area. Intel's Skymont a.k.a. Darkmont have the same area and the same number of SIMD pipelines as Cortex-X4, so like Cortex-X4 they are also more efficient than the much bigger core Lion Cove, which is faster on average for non-optimized programs but it has the same maximum throughput for optimized programs.
When compared with Cortex-X4/Neoverse V3, a Zen 5 compact core has a throughput for array operations that can be up to double, while the area of a Zen 5 compact core is less than double the area of an Arm Cortex-X4. A high-clock frequency Zen 5 core has more than double the area of a Cortex-X4, but due to the high clock frequency it still has a better performance per area, even if it no longer has also a better performance per power consumption, like the Zen 5 compact cores.
So the advantage in ISA of Aarch64, which results in a simpler and smaller CPU core frontend, is not enough to ensure better performance per area and per power consumption when the backend, i.e. the execution units, does not have itself a good enough performance per area and per power consumption.
The area of Arm Cortex-X4 and of the very similar Intel Skymont core is about 1.7 square mm in a "3 nm" TSMC process (both including 1 MB of L2 cache memory). The area of a Zen 5 compact core in a "4 nm" TSMC process (with 1 MB of L2) is about 3 square mm (in Strix Point). The area of a Zen 5 compact core with full SIMD pipelines must be greater, but not by much, perhaps by 10%, and if it were done in the same "3 nm" process like Cortex-X4 and Skymont, the area would shrink , perhaps by 20% to 25% (depending on the fraction of the area occupied by SRAM). In any case there is little doubt that the area in the same fabrication process of a Zen 5 compact with full 512-bit SIMD pipelines would be less than 3.4 square mm (= double Cortex-X4), leading to a better performance per area and per power consumption than for either Cortex-X4 or Skymont (this considers only the maximum throughput for optimized programs, but for non-optimized programs the advantage could be even greater for Zen 5, which has a higher IPC on average).
Cores like Arm Cortex-X4/Neoverse V3 (also Intel Skymont/Darkmont) are optimal from the POV of performance per area and power consumption only for applications that are dominated by irregular integer and pointer operations, which cannot be accelerated using array operations (e.g. for the compilation of software projects). Until now, with the exception of the Fujitsu custom cores, which are inaccessible for most computer users, no Arm-based CPU core has been suitable for scientific/technical computing, because none has had enough performance per area and per power consumption, when performing array operations. For a given socket, both the total die area inside the package and the total power consumption are limited, so the performance per area and per power consumption of a CPU core determines the performance per socket that can be achieved.
I am sure if IBM had more of a market than the minuscule Mac market for laptop class PPC chips back in 2005, they could have poured money into making that work.
Even today, I doubt it would be worth Apple’s money to design and manufacture its own M class desktop chips just for around 25 million Macs + iPads if they weren’t reusing a lot of the R&D
They just sat on it, their marketing dept made fancy boxes for high end CPUs and their HR department innovated DEI strategies.
It’s amazing that the “take responsibility”, “pull yourself up by your bootstraps crowd” has now become the “we can’t get ahead because of minorities crowd”
The best people were clearly not staying at Intel and they have been winning hard at AMD, Tesla, NVIDIA, Apple, Qualcomm, and TSMC, in case you have not been paying attention. They could not stop winning and getting ahead in the past 5-10 years, in fact. So much semiconductor innovation happened.
Yes, if you start promoting the wrong people, very quickly the best ones leave. No one likes to report to their stupid peer who just got promoted or the idiot they hire from the outside when there are more qualified people they could promote from within.
--
And re marketing boxes, just check out where Intel chose to innovate:
https://www.reddit.com/r/intel/comments/15dx55m/which_i9_box...
It wasn’t because of “DI&E” initiatives and a refusal to hire white people
That's scam. If you fail to profit, you should admit it, not fake it.
The issue with Intel though is that they needed the money to invest in R&D.
But Intel gave up and sold it off, right as smartphones were reaching mainstream.
Ironically, they have lost the driver advantage in Linux with their latest Arc stuff.
I trust they could have done a lot better, a lot earlier, if they cared to invest in iGPU. Feels like deliberately neglected.
The lunar lake Xe (IE the generation before the current one) is not rock solid on linux - i can get it to crash the gpu consistently just by loading enough things that use GL. Not like 100, like 5.
If i start chrome and signal and something else, it often crashes the gpu after a few minutes.
I've tried latest kernel and firmware and mesa and ....
The GPU should not crash, period.
But it also feels like that's because it crashed so much it was bothering their engineers, so they made the recovery robust.
their ”1 in a billion” (excuse) became $1 billion (cost to them).
of course, the CEOs not only go scot-free, but get to bail out with their golden parachutes, while the shareholders and public take the hit.
This is the employment contract that was negotiated and agreed to by the board / shareholders.
Come eat some chili widdus.
Id'll shore put some hair on yore chest, and grey cells in yore coconut.
# sorry, in a punny mood and too many spaghetti western movies
Nothing like giving piles of cash to a grossly incompetent company (the Pentium math bug, Puma cablemodem issues, their shitty 4G cellular radios, extensive issues with gigabit and 2.5G network interfaces, and now the whole 13th/14th gen processor self-destruction mess.)
I laughed when I read this. It’s hard enough to get support for basic issues, good luck explaining a hardware bug.