AI chip startup Wave to buy Silicon Valley old-timer MIPS
cnet.com
cnet.com
From https://www.eetimes.com/author.asp?section_id=36&doc_id=1332...
"... if a company seeks to manage MIPS correctly, what’s the strategy? “The value is in restoring the IP roadmap, the MIPS brand, and coming at the market in a way that doesn’t put MIPS head-to-head with ARM, but concentrates on the parts of the market where ARM is weak and there is good opportunity to license,” one executive said. Citing Mediatek’s recent decision to use MIPS instead of ARM in its modem, he stressed, “ARM is in the modem by default, not for any particular suitability or technological advantage. The apps processor is where ARM has the real advantage, as all of Google Android is built on it, so you avoid this area completely.”
From https://fuse.wikichip.org/news/1373/wave-to-acquire-mips/
"[Wave] ... raised $56.7M in funding led by Tallwood Venture Capital – the current owners of MIPS ... While Wave Computing holds over 60 patents, MIPS holds 100s ... in March, Wave announced that they will be integrating the 64-bit MIPS IP cores into their future DPUs. The integration is done in order to allow existing MIPS-based RTOS to handle the control and management functionalities of the chip. Conversely, Wave might be seeing an opportunity by injecting their IPs into the MIPS existing ecosystem. MIPS software and tools are quite mature and can certainly accelerate Wave’s development."
The no. 1 reason ARM went ahead
What made them special is how they combined their hardware with their own version version of UNIX.
Hence why even trying to write portable code across UNIXEs wasn't as easy as many believe.
Irix had IrisGL, Iris Inventor, XFS and a few other goodies, before they became the Open variants.
SGI was also a big sponsor of the C++ STL work.
Domain Specific Architecture talk by Hennessy & Patterson, 2018 talk for 2017 Turing Award: http://iscaconf.org/isca2018/docs/HennessyPattersonTuringLec...
However I can say yes, we are returning to those days.
Thanks to the thin razor margins you see OEMs now selling un-upgradable phones, laptops, tablets, 2-1, leaving the old desktop PC for a niche market that gets smaller every year, mostly targeted at gamers.
Also OEMs got envious of Apple as the surviving icon of those days, all of them want to be Apple of their market and sell experiences.
Not business model per-se, but not being crippled by pointy hair boss types from C-levels to lowest level managers.
Second to that what got them ahead was their lack of greed and ambition.
Few people who had firsthand experience with buying IC IP licenses I know tell of MIPS 10 years ago as "we will not simply sell you a core, but force feed you all peripherals, and make you pay for 20 more licenses for god know what along the way." Compare it to ARM where you wire money today and get a guy coming to your office to drop off harddrives with netlists a week after.
For support, documentation, software, and everything else ARM was light years ahead.
Is this a specific reference? What device would that be?
I'm still impressed with my dual-R12000 Octane today.
https://www.cavium.com/octeon-III-CN7XXX.html
Their stuff turns up on eBay sometimes. Last one I saw was an Octeon II PCI card for just over $200. I was going to use it for network/firewall offloading. Blocking things like Intel ME, too.
Can you elaborate on this? How were you using these to block Intel ME?
I was thinking last year about building one with new hardware given all the hardware/firmware vulnerabilities showing up. I have it on back burner fof now.
1. It needs a communication channel. My guard denies it that.
2. It will create or receive packets using standard format that existing channels support. Using obfuscated formats before traffic hits the guard means its packets will just be dropped.
3. It can still be programmed to just attack the system somehow. Now, we're in realm of targeted, sophisticated attacks more likely to be noticed. That's already an improvement to status quo.
I still think someone should try paying either company's semi-custom division to make one without a ME. Add an extra must-have feature that accelerates common workloads, esp web cache or database, to generate further sales. If they refuse to remove it, then that would be really admitying something if they accepted other modifications for money.
The real future is going to be in reprogrammable hardware, something akin to FPGAs but with cores as building block instead of gates. We need massively parallel chips with an order of magnitude more cores than we have today (at least 256 as a baseline), with better interconnect and smarter routing that can do something like content-addressable memory.
MIPS is an ideal candidate for this, as it's a textbook example of the minimum number of transistors needed to implement a pipelined CPU. They should be able to fit hundreds or even thousands of cores in the number of transistors wasted on cache today in mainstream processors from Intel and AMD. Assuming they don't just kill MIPS to do their own proprietary AI DSP..
There is a massive trade-off. We don't have the power budget to make 100s of cores, so you need to start ripping stuff out of cores to make the power costs tolerable. In turn, that means that your single-threaded performance is going to go down. If your code is massively parallelized, you can easily cover those costs. But for high scaling, you start having problems with the communication costs. The problem isn't fitting the ALUs into the chip, it's filling the ALUs with data.
The most likely future is that we move to a more heterogeneous world: you have a few big cores for handling unparallelizable code (and handling things like interrupts) intermeshed with various kinds of accelerators. One of them could be a very high-throughput (at cost of higher latency) LINPACK-style accelerator kind of like GPUs (but not living off of a slower PCI bus). You could attach an FPGA as well. But the very-high core count processors just haven't worked well (ask Intel how well Xeon Phi worked out).
CPU performance increases since 2000 have mainly come from longer pipelines, bigger branch prediction logic and larger/deeper caches. Those are all extremely important for single-threaded performance but are not much use for embarrassingly parallel (low branch) computation like in MATLAB/Octave, R, shaders, AI, physics and so on.
You're right about Xeon Phi and make an absolutely valid point about the trouble of filling ALUs with data. It was probably destined to fail because most mainstream languages just fall down terribly trying to do multiprocessing. Only a handful like Erlang and Go get it right, but at the time just weren't on the radar.
I think the world is treading water with the current CPU/GPU divide and heterogeneous processors like the Cell. Few developers fully utilized the Synergistic Processing Elements (SPEs). These paradigms provide powerful hardware but leave the intricacies to the developer, or sometimes the compiler if they're lucky.
That's the part that I'm getting tired of. Companies tout the power of their various technologies without really addressing the hard problems in computer science. I want a paradigm that presents itself as a single unified CPU and memory. I want the compiler and operating system to optimize my code and handle moving my data with copy-on-write. If I had something like this and a decent Actor Model language, then it could trivially emulate things like shaders and we could get rid of all of these domain specific languages.
What I'm getting at is that with a few hundred short-pipeline cores from the 90s (say MIPS or PowerPC 600 series) and a content-addressable memory, compute-bound problems become tractable. Why bother with rasterization when you can write a ray tracer in a page of code and get better results? Or imagine being able to run not just one neural net but many in parallel, or let genetic algorithms run for 2 or 3 orders of magnitude more generations. I want raw computing power, not a subset of it filtered through someone else's notion of what I might need.
And better misaligned memory handling (Haswell has no latency penalty reading a word that crosses cache lines), better speculative execution, wider SIMD units, more complex SIMD swizzling operations.
> That's the part that I'm getting tired of. Companies tout the power of their various technologies without really addressing the hard problems in computer science. I want a paradigm that presents itself as a single unified CPU and memory. I want the compiler and operating system to optimize my code and handle moving my data with copy-on-write. If I had something like this and a decent Actor Model language, then it could trivially emulate things like shaders and we could get rid of all of these domain specific languages.
Speaking as someone who writes an optimizing HPC compiler... the idea that you can write a simple program and have it magically parallelize to ultra-high-throughput code is fantasy. We had that idea back in the 1960s. We've still not solved it 50 years later, and we've learned that there are giant barriers towards making these fantasy compilers that are rather fundamental.
One major problem you have to consider: communication is expensive. Already, on a single chip, it's impossible for data to move from one end of the chip to the other within a clock cycle (the speed of light in a vacuum takes about 10 times the clock cycle time to do so, and the speed of light in a physical wire is slower than that in a vacuum). So you start paying latency costs to have to move bits of data around. But your interconnect mesh isn't all-to-all, not when you have hundreds of access points. So that means that you have to pay attention to who is communicating to whom, and how much they are doing so. In a pathological case, like reversing a list, you've got all of your data moving across a single link in a 1-D network or only sqrt(#processor size) links in a 2-D network.
You do realize that ray tracing is ultimately a problem whose performance is determined by the external memory system?
Unless you’re thinking about each mini CPU rendering a bunch of reflective balls where the full scene can be stored locally in each CPU.
Because once you don’t, you’ll have to find ways to cover the access latency to the shared memory pool, and before you know it your super simple CPU will looks suspiciously like the shader core of today’s GPUs.
Your other examples have similar limitations.
The truth is that there are not many problems that can efficiently be mapped to an architecture with tons of small CPUs, some local RAM, and nothing else.
that is a kind of a step in the similar direction:
https://www.electronicsweekly.com/news/risc-v-processors-mak...
Some other related companies had modified GCC to automate the process of co-processor offload.
These massive number of little cores were supposed to be the next evolutionary step going from customer transistors to standard cells to ...
When I asked about power consumption the answer was that this was one of the remaining problems, but it would eventually works itself out.
It’s still one of the biggest problems, but it has only gotten worse.
Most companies would love to see their competitor adopt an architecture like that. :-)
The only place where it might find a place is as an FPGA alternative, but since FPGAs are one or two orders of magnitude less efficient then ASICs, that’s really not a high bar to clear.
I keep expecting this to be a retro branding exercise on something fundamentally different, like a Fiat 500 or a modern VW Beetle.
MIPS has been relevant only in the embedded space for the last 10 years, if not more, and even saying that it is relevant in the embedded space is kindda pushing it at this point. The widespread availability of reference designs has ensured a pretty long lifespan in industries where that's especially relevant (see e.g. Atheros) but I think they're pretty much near the point of no return, where either they find a few niches that ARM can't (or won't) fit in, or they finally go the way of the Dodo.
I think there are places in the market where there's room for them, and MIPS has its strongpoints (e.g. multithreading) that make it useful in some fields, such as real-time systems.
To be honest, I have doubts that the people running the show will be able to pull it off but I really hope I'm wrong.
The last I heard of MIPS was some Chinese-fabbed chip that used the unencumbered set of MIPS opcodes--probably MIPS II--and a mostly-GNU OS to build the only computer on Earth (at the time) completely free of patent or licensing restrictions. It sounded like a stunt then, and I haven't heard anything about it since, so it probably was.
Wave looks like a pretty cool startup, but I'm not sure how helpful MIPS architecture would be to them. The biggest benefit is probably not getting sued by Intel or other large players.
https://www.nytimes.com/1992/03/13/business/silicon-graphics...