ARM: vmull.p8
Intel: PCLMULQDQ
Power9: I forget, but trust me, its in there somewhere.
IIRC, they're all 64-bit carryless (or polynomial) multiplications. The funniest part is that ARM and Power9 still claim to be "reduced instruction set computers", despite having highly specialized instructions like these. Power9 is the funniest: implementing hardware-level decimal-floats (for those bankers running COBOL), quad-floats (for the scientific community), Polynomial-Multiplication and more.
Here's a CRC32 impl using POWER8/9 polynomial instructions: https://github.com/antonblanchard/crc32-vpmsum — coincidentally, you don't have to do this on ARM, ARM just has CRC32 instructions.
> hardware-level decimal-floats (for those bankers running COBOL)
IBM being IBM :) I've heard somewhere that they are sharing some internals between POWER and z/Architecture (mainframe) chips. No idea how true it is and how much is shared if true, but a scenario like "decimal floats were implemented because mainframes are using the same backend with a different frontend, so we exposed them in the POWER frontend too" sounds quite plausible.
To be fair, even back in the day decimal arthimetic could at least have an argument made for it on RISC chips. https://yarchive.net/comp/bcd_instructions.html
Ain't that the truth. Reduced never meant anything; they never defined it. Patterson said they cooked up the name on the way to a grant presentation. I think they should call it Warmed Over Cray. What was new in RISC that wasn't in Cray? It was like the Geometry Engine; it was possible for an academic department to accomplish but that doesn't mean they accomplished anything that hadn't already been done.
preferably fixed-length
Or fixed-length/2.
https://web.archive.org/web/20181126134850/http://uranium.va...
- 68010s were essentially a 68000 with this fixed - unix could do a real paging system
- 68020s dumped their entire microcode state on the kernel stack - to do signals correctly (ie recusively) unix systems had to be able to spill this to user space and validate it on the way back in
- 68030s and 68040s were saner and restarted exceptions in useful ways
After all, you don't need to be able to restart on a page fault if you just swap in whole tasks out and back into memory just before switching to them (which you know when you're going to do), and you still get virtual memory's benefit of the swapped in tasks being able to occupy different physical, and discontiguous, memory regions without the task itself noticing.
So did they only design for that, or did they actually make a mistake?
Sun solved this problem in an expensive way. The MMU would make the CPU wait while a second 68000 was used to process the page fault.
Now that you mention it, I recall reading that two 68000s story before (possibly from another comment you made here?), and reading it this second time, it still amazes me. Is it possible to find the source code involved? Not sure what of Sun's kernels is available.
You do need to restart instructions if you want your stack to extend automatically, there's a hack you can do to make this happen on a 68000
You could do a paging system on a 68010/68451 - you had to treat the 68451 as a software refilled TLB
Sun built their own MMU using SRAM, at the time there was a guy going around the Valley designing "sun-style" MMUs, for IP reasons every one was different, some were broken (one I ported to couldn't protect the kernel from user mode).
Someone (maybe Sun?) managed to get 68000s to do paging by having two of them and freezing the primary on a page fault and having the second one load the page somehow.
68020s had external MMUs, Moto made the PMMU which most people used Phillips tried to make one (it never really worked), Apple had a weird pseudo MMU that fit in a PMMU socket on the mac 2 - it emulated the 24-bit addressing on the 68000 (Apple had stupidly used the upper bits in their pointers)
Finally: I think that the 68000 not being able to restart a page fault was probably a mistake, something they hadn't thought all the way through
Though qualified with "until the design was almost complete", making the answer to my question (as often) much more fuzzier than I thought.
And another commenter speculates it was Apollo with the two-68k implementation, but without sources and the commenter is unsure of it themselves.
In any case, very interesting stuff.
EDIT: Someone else claims it was Masscomp: http://www.dadhacker.com/blog/?p=1383#comment-3582
The '451 is more functional than his power of 2 design (because it offers multiple powers of 2 that you can cobble together into arbitrary regions) - the 'throwaway instructions' he mentions is indeed what we used for stack extensions - you do a 'tstb offset(%sp)' in every subroutine entry where 'offset' is the extent of the stack the subroutine will use, the results of the instruction aren't used and the fault can be detected - by recognizing the instruction at the fault pc offset and returning to the address following it
Again as someone who somehow never got to touch 68k during its day, the level of creativity required in 68000 MM, and the internal state spilling of the 68010 and later, really got me curious (x86 of course always had plenty of curious quirks, but I think I'm too familiar with them at this point %) ).
I'd also still love to read more about the details of the two-68000 design, but everything I could find so far was essentially folklore.
Funny story is that this is still kind of a problem today, where a single pointer-dereference may be an mmap'd file over a network over SSHD tunneled through (increasingly arbitrary indirections).
foo->next = bar->next->next can be surprisingly complicated. I don't think most programmers think to handle SIGBUS errors.
Whereas to me, it sounds like this polynomial evaluation instruction needed all the 21 pages within one (very "large") instruction, meaning you actually need to have those (again, pathological case) 21 different pages present at once?
EDIT: A sibling comment mentioned keeping partial state of (apparently) single instructions instead of fully restarting them. So that might be an option, too.