Pentium floating-point division bug (1994)
en.wikipedia.org
en.wikipedia.org
To pay my way for the silliness: I will highly recommend “The Pentium Chronicles” for a view into that time and place.
It covers the development of the P6 arch, but there is some inside baseball about the P5. And it’s just one of the best books I’ve ever read about really amazingly well-run engineering efforts.
My favorite from the 90's was, what does a computer and an air conditioner have in common? Both stop working when you open windows.
No Windows, no Gates and Apache inside.
"In a world without walls or fences, who needs Windows or Gates?"
If you're lucky enough to know someone with the right IEEE Xplore subscription: https://ieeexplore.ieee.org/book/5989703
No promises (I’ve moved a lot) but if you email me I’ll try to find it and if I do it’s yours.
Cyrix was a really fascinating company. Extremely short lived but had a huge impact on Intel's grip over the processor space (Intel lost all their lawsuits against them and nearly faced antitrust proceedings because of it).
This bug is a great counterexample to the adage that 'computers don't make mistakes - they do exactly as they're told'.
If spheres have theoretically zero contact area, and they stack in a tube, is there zero contact area between the spheres?
We say that the limits of 1/x and 2/x are equal, but they have different slopes approaching said asymptotic limit
[1] Oral history of Robert P. Colwell (1954- ) Interviewed by Paul N. Edwards, Assoc. Prof., University of Michigan School of Information, at Colwell’s home near Portland, Oregon, on August 24-25, 2009
https://web.archive.org/web/20210726205114/https://www.sigmi...
Dr. Nicely caused quite a bit of excitement at Intel. I was on the p6 architecture team when he discovered the FDIV bug. Our FPU was formally verified and didn't have the same bug. To be nice to Dr Nicely we sent him a pre-release p6 development system to test with his program to demonstrate that his bug was fixed. He was working on a prime number sieve program and came back reporting that the p6 ran at 1/2 the speed of a Pentium for his code. Wow, another blackeye/firestorm caused by Dr. Nicely. He had too much of an audience for him to report to the world this new processor was slower.
So I got to spend a lot of time learning how to sieve works and what is happening. For the most part, it allocates a huge array in memory with each byte representing a number. You walk the array with a stride of known primes setting bytes and whatever is left must be prime. ie. every 3 is not prime, every 5 is not prime, every 7...
So in the steady state, you are writing a single byte to a cache line without reading anything. And every write hits a different cache line.
Now p6 had a write-allocate cache, but the Pentium would only allocate on read, so on the Pentium a write that misses the cache would become a write to memory. On the p6 that write would need to load the cache line from memory into the cache and then the line in the cache was modified. And since every line in the cache was also modified we had to flush some other cache line first to make room. So every 1-byte write would become a 32-byte write to memory followed by a 32-byte read from memory.
Normally write-allocate is a good thing, but in this case, it was a killer. We were stumped.
Then the magic observation: 99% of these writes were marking a space that was already marked. When you get up to walking by large strides most of those were already covered by one of the smaller factors.
So if you change the code from:
array[N] = 1
to: if (!array[N]) array[N] = 1
Now suddenly we are doing a read first, and after that read we skip the write so the data in the cache doesn't become modified and can be discarded in the future.
Also, the p6 was a super-scalar machine that ran multiple iterations of this loop in parallel and could have multiple reads going to memory at the same time. With that small tweak, the program got 4X faster and we went from being 1/2X the speed of a Pentium to being twice the speed. And this was at the same clock frequency. The test hardware ran 100Mhz, we released at 200Mhz and went up from there.I'll never forget the size of the cpus compared to others at the time. Running them in dual-cpu setups was fun also.
And just like that, the infamous Pentium bug hit the headlines globally.
Also, in the memory of Dr. Thomas R. Nicely (1943-2019), who found the bug:
“I am Pentium of Borg. Division is futile.”
just following the example, the result is wrong after the 4th digit. this is absolutely 'could be handled in software with a significant performance loss' territory.
that isn't to say this isn't breaking -- it absolutely is.
what I'm getting at is that my CPU runs 30% slower with mitigations. if I were running multiple servers like the computer I use, and I was near the capacity of my resources, I would need to add some machine(s) -- and people do.
in the same way this is serious, these current issues are the same yet somehow we've lost some stuff along the way and we just don't do recalls or even receive any remedies. where's my check?
You can't catch spectre or similar things where multiple blocks all doing the wrong thing together with UVM or formal nowadays either(too large search space), so I'm eager to see what kind of top-level verif. methodologies people will invent as the designs grow larger and larger.
But it was "fake" advantage (due to ignoring bugs) - now that those bugs come out to light, their processors are slower. But were already sold.
And yes I know that AMD has some similar bugs.
The US government will not go after Intel though due to public safety issues - they want their processors to be manufactured at US soil (different thing is if they are really safe with all those bugs and spying features, also are they even manufactured inside USA anymore?).
Other governments could go aftet Intel though.
So, how does your theory stand?
If we assume that both companies were cutting corners and one was ahead, then?
Meltdown arose because vulnerable designs delay memory protection checks until as late as possible, after illegal accesses could have already affected state. This is a questionable decision, as it is playing fast and lose with the most fundamental building blocks of security. Most companies wisely chose a more cautious approach.
If an investigation found that watermelon is carcinogenic, that would be a problem. If that same investigation found that Andy's Farm sells watermelon that contains benzene, Andy's Farm couldn't defend itself by saying all watermelon is carcinogenic.
way, way back at the beginning of the PC revolution, for example: SRAM vs DRAM? Intel went all in on DRAM. Imagine how different the world could be today if they'd chosen differently an could expect all main RAM to be "cahce latency". DRAM would be a slower dynamic store, possibly looking more like hard drives did.
What if they'd decided that unifying a bus was a good idea instead of tossing out new ones every couple years? Imagine the longevity and variety of the VME bus ecosystem couples with the size of the PC market, in the 80s.
This isn't really how it went. There was never a future with hundreds of megabytes of SRAM as it requires significantly more die area to produce and more power to use, making it significantly more expensive. The entire point of caches was because we couldn't afford to just make everything SRAM. Even today, we are only just getting to the point where you might have a few hundred megabytes of SRAM on the most expensive server CPUs.
https://www.servethehome.com/amd-genoa-x-the-1-1gb-l3-cache-...
It's a little bit mindblowing.
Because SRAM is expensive compared to DRAM. SRAM requires four transistors per bit, but DRAM requires just one. And that one transistor doubles as your capacitor. In addition, routing makes a single SRAM cell a bit bigger than four DRAM ones (at the same process node). So, DRAM can be packed to densities that are not feasible for SRAM. There's a reason AMD (and others) is/are starting to put cache on an entirely separate die (X3D).
The buy back option was based on pricing before the news broke and for a good condition vehicle, there was an adjustment for odometer miles, but the vehicle could be in any condition, just had to run and move on its own power. They were not permitted to resale purchased vehicles anywhere globally unless they were modified to properly comply with emissions controls.
And they were forced into selling EVs and building a charging network. Maybe that will work out for them, maybe not.
https://users.fmi.uni-jena.de/~nez/rechnerarithmetik_5/fdiv_...