Is MIPS Dead? Lawsuit, Bankruptcy, Maintainers Leaving and More
cnx-software.com
cnx-software.com
I remember reading a PCWorld article comparing various processors. This was the time the Pentium came out. If my memory servers, the MIPS won the overall performance crown vs Pentium, PowerPC, and Alpha AXP.
Not MIPS. Windows NT was initially developed on a more obscure RISC architecture, Intel i860. It was later ported to MIPS though.
I got to play with an Intel iPSC supercomputer in our Physics lab (at university) that used the i860. We had a Sun box that hosted cross-compilers for it and allowed you to reserve a specific number of i860 processors for you. The number had to be a power of 2. Fun times!
Learning MIPS was what originally got me interested in ASM programming since we had a class that was focused on MIPS code and another class that had us build a digital MIPS processor from scratch. The combination of these two classes really sold me on the magic of super low-level programming.
Its version of 32-bit MIPS is so simple, its whole instruction set fit in a 2-side cheatsheet (the famous "green sheet"). The design of the CPU is quite easy too. Given an instruction and its binary representation, it is almost straightforward to see how each bit contributes to the computation (setting the correct ALU operation, retrieving a value from the correct register, etc.).
We can make a parallel between those low-level ISA and high-level languages: a language like Lisp is lean and simple so it is taught and presented as good design (and people who went through that education keep that in memory), but when it comes to produce real program almost everybody chooses a much less regular language, which is way more practical. (Same could be said for stack-based languages like Forth, which present an extremely simple model to apprehend, but that doesn't mean at all that it is simple to program in.)
Or postfix vs infix for mathematical expressions/calculations. Same principle: the one which is based on a very simple model is praised by aesthetes, but almost everybody prefers the other one, which is simpler to use because it is more natural, despite being based on a more complex model.
In fact, the simplicity of the model is not of much interest for the user, it just makes the life of the implementer easier. But for 1 implementer, there are thousands or millions of users, who want ease of use, not ease of implementation.
First, mod-r/m addressing on x86 is fairly traditional and can often save considerable calculation over a "simpler" addressing mode (given the opportunities for add-and-scale operations).
Second, treating x86 machines as load/store architectures passes up the opportunity to achieve improved code density and increased execution bandwidth from "microfusion" - this is when a operation (e.g. "add") is done with a memory operand. Microfusion, for those not familiar with it, allows two "micro-ops" (aka uops) that originate from the same instruction to be "fused" - that is, issued and retired together (even though they are executed separately).
This can occasionally - in code that has already been militantly tuned to an inch of its life - yield speedups, as Skylake and similar can only issue and retire 4 uops per cycle. However, there are 8 execution ports (of which only 4 do traditional 'computation'). Carefully designed code can take advantage of the fact that issue/retire are in the "fused domain" while execute is "unfused domain" - so you can sometimes get 4 computations and 1 load per cycle even on a 4-issue machine.
I was trained on MIPS and Alpha, so of course old habits die hard, and it's always tempting to go old school and design everything to act as if the underlying machine is a load-store architecture. However, this (a) isn't necessary on x86 and (b) often won't be faster.
The other blow against load-store is that a modern o-o-o architecture can hoist the load and separate it from the use anyway - and it doesn't have to consume a named register to do it (it will use a physical register, of course, but x86 has way more physical registers than it has names for registers). This of course is a bigger deal for the rather impoverished register count of x86 so it is, in the words of a former Intel colleague on a different topic, a "cure for a self-inflicted injury".
Even a global optimum for one is unlikely to be an efficient solution for all.
Or how about a macro assembler where you can do that with custom pseudo-instructions?
Branch delay slots were a somewhat clever solution to reduce the complexity of the original implementation, but they baked implementation details into the ISA and became problematic when the implementation details changed.
Same reason why stuff like VLIW has failed to catch on. These things are so dependent on specific hardware implementation details that one can hardly call them general-purpose ISA's anymore.
The Onion Omega2S is based on a MediaTek MT7688 (MIPS32LE) and since the demise of the CHIP-Pro, is really the only inexpensive surface mount Linux SoM left.
[1] https://wavecomp.ai/wp-content/uploads/2018/12/WP_CGRA.pdf
(Stanford)
MIPS Computer Systems
SGI
Spun off, IPO
Imagination
Tailwood
Wave
I might have missed one.The question it's not if someone will use some old MIPS ISA in a few years from now on, the question is if someone will improve the ISA from now on, in the same way x86 and ARM are consistently being improved.
Some companies are still fabbing old Z80 CPUs, but that's not to say Z80 has a bright future.
CIP-United still promises to provide enhanced versions of the both the architecture and the MIPS Warrior cores for the Chinese market, regardless of what happens to MIPS Technologies. This may seem utterly futile now, but it is also the very thing that the US Committee on Foreign Investment was trying to prevent when it required MIPS to be spun out of Imagination Technologies when that got sold to Chinese investors.
The old 32-bit Arm (now called Aarch32) was quite different and only somewhat RISC-like. Arm's Aarch64 however is mostly derived from MIPS64 with a lot of modernization plus some parts (exception levels) from 32-bit Arm.
MIPSr6 was an attempt of modernizing MIPSr5 by removing all the ugly bits (delay slots!) but the incompatible instruction encoding prevented it from being widely adopted. You cannot buy a single MIPSr6 machine that a mainline Linux runs on.
RISC-V's design looked at all RISC architectures (Berkely RISC, MIPS, SPARC, Power, Arm, ...) for inspiration and took the best parts of each. Leaving out all the historic baggage means it's simpler (the manual is a fraction of the size), but most of the important decisions are the same as in MIPSr6 and Armv8/Aarch64.
One notable difference is the handling of compressed (16-bit) instructions: ARMv8/Aarch64 doesn't have them at all (like RISC-I/RISC-II, ARMv3 and MIPS-V), MIPSr6/microMIPS needs to switch between formats (like ARMv4T through ARMv6) and in RISC-V they are optional but can be freely mixed (somewhat like ARMv7 and nanoMIPS).
For example, the notion that condition codes interfere with OoO execution has been repudiated; Power and x86 both now rename condition registers. Lack of popcount and rotate in the base instruction set are glaring omissions. (That x86 got popcount late, and that the bitmanip extension will have them if it ever gets ratified, are no excuse.) It was silly to make the compare instruction generate a 1 instead of the overwhelmingly more useful ~0.
We only get a new ISA once in a generation. It is tragic when it is wrong.
It is possible, in principle, that popcount and rotate could be added to the base 16-bit instructions, but I'm not holding my breath.
As an anecdote, I have personally switched a MIPS core to an ARM core in an SoC revision specifically because ARM gave us better licensing terms than MIPS. It was a pain because MIPS big endian and ARM big endian are not directly compatible with each other.
MIPS has been the walking dead for a couple of decades now. It's really time to let go. Especially since there are actually interesting things happening in the ISA space with RISC-V.
Just take a look at their membership page for people funding or helping develop the processors, https://riscv.org/members-at-a-glance/
[1] https://www.westerndigital.com/company/innovations/risc-v [2] https://www.extremetech.com/computing/281891-western-digital...
There's dead and there's dead. MIPS is only dead in the way that internal combustion engines for transportation are dead.
In the '90s and early 2000s there were lots of CPUs, computer systems, operating systems.
ARM = RISC?
Intel = CISC?
? = MIPS?
What else is out there in large amounts?
But there is one notable CISC still out there: IBM's mainframe z/Architecture. Might not sell a lot of units, but it's still pretty important commercially.
Worth noting that POWER similarly lives on in radiation-hardened form for NASA; the RAD750 (and its predecessor, the RAD3000) both have a long history of interplanetary use.
MIPS also sees some use here, too; for example, the MIPS-based Mongoose-V is what New Horizons uses, and the KOMDIV-32 ostensibly is designed for Russian spacecraft use (but I don't know of any specific examples).
As for what else is out there in large numbers, I think a lot of 8bit Z80 and AVR microcontrollers (like the ATmega328 used in Arduino Uno) are embedded in larger systems and go unnoticed. More recently I've seen the Xtensa (32bit RISC) catching on with the popular ESP32 and ESP8266 SOCs.
I struggle to think that ARM can really be called reduced-instruction after neon or so.
> The term "reduced" in that phrase was intended to describe the fact that the amount of work any single instruction accomplishes is reduced—at most a single data memory cycle—compared to the "complex instructions" of CISC CPUs that may require dozens of data memory cycles in order to execute a single instruction.[23] In particular, RISC processors typically have separate instructions for I/O and data processing.
So yeah, if SIMD instructions can execute in a single cycle (or maybe even a small number), then it still counts.
- Predication as a major architectural feature -- every instruction can be conditionally executed - Complex load-store instructions: ldm/stm can operate on a large set of registers in a single instruction, including performing a branch by loading into the instruction pointer - 16-bit Thumb instruction format (also optionally present in RISC-V and newer MIPS)
64-bit Arm mostly drops all of the above and is basically a traditional RISC implementation.
It's also a thing to fuse smaller operations into macro-ops in many microarchitectures.
All high performance chips avoid running microcode, it's reserved for eg emulating seldom used legacy instructions. Microcoded CPUs (where all instuctions are implemented with microcode) were a 80s/90s thing.