SIMD in the 90s: Programming Intel's Pentium MMX
pikuma.com
pikuma.com
https://en.wikipedia.org/wiki/X86-64#Microarchitecture_level...
"Additional XMM (SSE) registers: Similarly, the number of 128-bit XMM registers (used for Streaming SIMD instructions) is also increased from 8 to 16...
"The original AMD64 architecture adopted Intel's SSE and SSE2 as core instructions."
https://en.wikipedia.org/wiki/X86-64
This wansn't v2?
Subsequently, there were additional instructions added in SSE3, SSSE3, SSE4.1 and SSE4.2, which are all incorporated into the v2 ISA level (along with a few other instructions). Then all of these instructions were given 256-bit variants in AVX, and AVX2 adds some more vector instructions; these are incorporated into the v3 ISA level. And then along comes AVX-512 and naming just becomes a podge at that point...
You may be mis-remembering LAHF/SAHF.
If you ever dug into x86 assembler programming... at first they make no sense at all. They only save/restore a tiny part of the available flag registers. The mnemonics themselves make little sense - load/store are not really used in any other base x86 mnemonics (unlike e.g. 6502 mnemonics, which use LD?/ST? instead of MOV like x86).
It only clicked when I read an Intel document about porting assembler code from the 8080 to the 8086.
LAHF/SAHF are basically convenience instructions to make porting easier. Many 8080 instructions did not alter the flags, unlike their 8086 counterparts. Substituting an `INX` instruction with `LAHF; INC; SAHF` made it possible to mechanically translate assembler source code.
And yeah, 8080 mnemonics had LDA and STA like the 6502...
This is a large part of why "floating point to integer conversion is slow" cargo culting came from (probably disappeared today but was often seen back in the day).
Many x86 standard libraries emitted code that first set the floating point rounding mode before the storing conversion, you could use a /QIfist compiler option on MSVC to omit that but if you were unlucky to call some code that fiddled with the rounding mode your code would become unreliable (iirc unfortunally that did include some old D3D or OGL version).
With SSE iirc it's separate instructions for different rounding instead of a FPU mode.
Due to the way the first Pentium 3 CPUs (Katmai) were built, they likewise aliased the x87 (and thus, MMX) registers to the XMM SSE registers, but this was hidden from programs. It wasn't until at least the Coppermine revision that they were separate registers again.
My family got a PIII 533Mhz coppermine, in the early sideways Slot 1 configuration. Blazing fast at the time. Been a long time since I heard anyone say "coppermine".
CPU progress was wild in the 90s, where you could wait two years and your new CPU would be double the old one's speed at the same price point. Today it takes near a decade for CPU speed to double.
I also recall the first game that advertised its exciting use of the new MMX technology: POD.[1]
I remember just after IBM got copper interconnects working, there was a stock market dive, I ranted to my partner about how crazy it is for a company to achieve a sought after goal and have its value decrease. The next day there was a news story about the entire market decline being stalled by the force of IBM buying back it's own shares. My partner told me I should rant less and invest more.
POD was pretty rad :) Thanks for the reminder of it (and all the MMX marketing around it.)
So no, Coppermine had nothing to do with Cu in the process.
There it is :)
Doesn't look like there's any conclusive info out there, so probably someone needs to actually decompile it and look for MMX opcodes.
The remaining references available regarding this factoid are incredibly ancient, too:
- https://groups.google.com/g/comp.sys.ibm.pc.games.misc/c/MRF...
- https://www.mobygames.com/game/644/pod/user-review/2384353/
No idea if any additional support was added in one of the subsequent releases.
During the 00ish years it felt like your system gets killed by a new game.
With the advent of the online update some game makers went way to far, Gothic in Europe was a pre-alpha release that never really got finalized but was really anticipated by gamers here.
I stopped buying games, since this was a rat race.
Everything was evolving and with the utilization of the GForce CPU gaming reached a new hardware design phase where CPU, bus (!), ram, essentially everything had to be finetuned otherwise you would run into massive performance penalties so to say. The term bottleneck wasn’t used back then but describes very clearly the dilemma back then.
In essence everything was getting more expensive while design considerations were critical.
And the systems were so fragile.
Some boards couldn’t handle Intel’s cooling demand for the CPU, it was a crazy time.
I remember the infamous AMD diss that led the company somewhat to the top tier their were then with the Athlon, when a dude filmed what happens during runtime when you remove the heat fan from an Intel vs. AMD.
Weird time, glad it is over. So much hardware and money burned.
(You may be confusing it with XMM vs YMM vs ZMM, i.e. SSE vs AVX vs AVX512.)
But the OP was talking about Intel's P3... so 3DNow! wasn't a factor...
SSE never had aliasing, so OP’s comment is strictly incorrect either way.
It's static text with a few images... that doesn't display until you tell NoScript to enable Javascript from a pile of third-party sites.
https://www.youtube.com/watch?v=5zyjSBSvqPc
https://www.youtube.com/watch?v=cnEuWfFTuxg
I was watching this and thinking "Who's going to even understand what this is about? Why is this playing on primetime TV?" Playing video on a computer wasn't a thing at the family consumer level at this time, and gaming was still a geeky niche.
MMX optimization practically required assembly language. The Pentium MMX was an in-order dual pipe CPU, and while compilers supported MMX intrinsics, their code generation for it was abysmal. Visual C++ 6, for instance, would emit code that was 2/3rds register-to-register moves, with values being unnecessarily moved between two registers between each ALU op. This was also a problem with SSE/SSE2 intrinsics. The worst case I saw was the _mm_set_epi8() intrinsic, which was used to construct a 128-bit vector from 16 inputs. When used with all constants, it should have generated a single 128-bit constant load; instead, Visual Studio 2008 generated ~80 instructions to compute it from byte loads. Microsoft didn't fix it until VS2010.
The latency of MMX instructions combined with the in-order dual pipe architecture also made asm loops messy. Simple ops were single-cycle, but multiplies had 3 cycle latency, stores required data an additional cycle in advance, and computed load/store addresses were also needed a cycle in advance. Simply running 2-4 iterations in parallel wasn't an option as there were only 8 vector registers and you'd still get bottlenecks on functional units. Getting peak performance thus often required interleaving loop iterations with special entry and exit code around the loop.
The issue with EMMS is understated. When the CPU switched to MMX, it marked the entire x87 stack as full. If you forgot the EMMS instruction, it wasn't just some strange floating-point bugs that would happen -- the next few x87 floating point calculations could just outright produce NaNs due to FP stack overflow. Furthermore, as these NaNs propagated, the CPU required microcode assists to handle them. So, even if the program didn't crash, an entire calculation domain would get poisoned and slow down by ~20x.
Ultimately, I don't think SSE2 was what killed MMX, but rather SSE, and specifically floating point. MMX not only didn't support floating point, but was also highly concentrated on 16-bit signed integers and secondarily 8-bit unsigned integers. Support for 32-bit integers was particularly lacking and pack/unpack conversions were a bottleneck. Trying to do 3D was cramped because doing so required fixed-point and MMX didn't have the same affordances as DSPs or NEON for rounding or implicit narrow/widen in operations, or even swizzles. SSE, on the other hand, was just straight floating point with standard automatic IEEE rounding and also had important added operations like swizzles and insert/extract. Thus, when 3D took off, SSE was far more useful than MMX.
MMX, however, still remained useful for a while for image and signal processing. SSE2 being twice as wide didn't help algorithms that couldn't use the greater width, such as 8x8 block motion prediction. Additionally, some CPUs at the time only had a 64-bit data path and had to split SSE2 ops, but because of their 4-1-1 decode template, could only decode one such instruction per cycle. The result was that code using the MMX registers could still run noticeably faster than with the SSE registers. This caused some confusion with the 64-bit version of Windows since Microsoft tried to say that x87/MMX shouldn't be used in long mode, but after queries from video processing companies had to document that the x87/MMX registers were enabled and context switched for user mode code.
I believe that only the first Pentium 3 core, Katmai, did this.
> This caused some confusion with the 64-bit version of Windows since Microsoft tried to say that x87/MMX shouldn't be used in long mode, but after queries from video processing companies had to document that the x87/MMX registers were enabled and context switched for user mode code.
I have some faint memories of hearing somewhere that Long Mode didn't support x87. I wonder if it is related to this early info you mention and it being Microsoft specific.
No, all Pentium 3s as well as the Pentium M. Pentium 4 notably didn't suffer from it, but it of course had many, many, MANY other performance issues.
> I have some faint memories of hearing somewhere that Long Mode didn't support x87. I wonder if it is related to this early info you mention and it being Microsoft specific.
It was VM86 mode that Long Mode didn't support, which was one of the rumored reasons for removing 16-bit NTVDM support (among many). x87 and MMX were always supported in long mode and notably some libraries like OpenBLAS still use x87 instructions. Windows does prohibit use of x87/MMX in kernel mode where the need is negligible.
The whole RAMBUS debacle... OTH DDR chipsets for Tualatin had shown what it was the end game for the P6 arch.
They already had a 16-bit software emulator for running NTVDM on other archs.
Apparently the real reason was they wanted to drop some software compatibility restrictions, like the small max size of HANDLE tables needed for 16-bit compat.
> No, all Pentium 3s as well as the Pentium M. Pentium 4 notably didn't suffer from it, but it of course had many, many, MANY other performance issues.
I had to google a bit to confirm this, and seems like I'm not the only one that understood it the way I did:
https://www.vogons.org/viewtopic.php?p=1360606#p1360606
https://www.vogons.org/viewtopic.php?p=1360620#p1360620
Basically, the info I knew was that Katmai had 128 Bits SSE registers but processed it as 2x64 Bits. That info is well reference pretty much everywhere. What is NOT explicitly mentioned is whenever Coppermine/Tualatin maintained that arrangement or had 128 Bits compute units for SSE, so the wording always made Katmai to look like an exception, as if everything else was 128 Bits.
You see, when two orders are possible, of course Intel had to ship both.
Yeeep, P3 vs P4. Which used which, is left as an exercise to you, dear reader.
No, the Pentium M (Banias / Dothan) also needed multiple cycles for basic SSE2 ops (with a few exceptions like PUNPCKLQDQ or PMOVMSKB), a full four years after Katmai (source: I owned one, and wrote SIMD on it for codec libraries).
Also, the GF4MX was just a GeForce256 in a trenchcoat, and didn't have vertex shaders (at least exposed to software, the T&L engine was still a vertex shader like engine, only the code was all written by nvidia and loaded from ROM).
GF3 was a very short lived product, even more short lived than GF2MX.
And the crucial thing here is what at the time of GF/GF2 you still could buy a computer without a proper[0] 3D and video output acceleration[1], but by the time of GF4MX (which is GF2MX on the booster shoes and VPE - and only a year and half later at the worst) you have no option not to get a full package.
SIS died, 3Dfx died even earlier, S3 became irrelevant and Intel was for the button pushers till i945 - which came only in 2005.
And while at first the price even for GF4MX420 was ~$100 it did two things:
- pushed the price for GF2MX series well below $50 (and quite soon followed the suit) so anything now had DX7 level 3D on-board
- pushed the price for every other 3D accelerator waaay down
- made MPEG and video acceleration the default in everything
[0] emphasis on the 'proper'
[1] though you needed to get out of the way to do so, eg i815, G450 - but you could!
Good lord, 400x300.
The wikipedia page on the Pentium has multiple references to an Intel publication called "Solutions", May/June 1993. It would be interesting to see that, but can't find a copy.
https://en.wikipedia.org/wiki/Pentium_(original)#cite_note-1...
I will keep looking...
An aside, but when SSE came a long that was a real big leap in 3D performance, just as GPU's started to gain some independence. So in about 2010, I tried to fire up Turok 2 just to see how fast it would run on a then modern CPU/GPU setup. It couldn't crack 200fps, however games only a year or two later would fly way past that. Turok 2 came out just before SSE and thus basically ran in purely x86/x87 space, thus the performance gap.