NUVIA: New Server CPU Startup Going After Intel and AMD
anandtech.com
anandtech.com
With the first problem, they can simplify things considerably by paying an ARM tax. The RISC-V tax is less ($0) but then it offers them less as well. If they design their own ISA, well, good luck with that. Also, clouds like tweaks, so one size won't fit all.
With the second problem, there's fab space to be had for sufficient coin. But there's more to manufacturing than filling out a webform and sending a tape with a check.
If they can get past the first two hurdles and actually deliver silicon which is significantly better than Intel's then marketing to the big clouds should be the least problematic.
Gonna take some money and time. Gulati left Google in March.
This investor+customer approach has a history. Yahoo did this with Google (search) and Apple did this with Adobe (laser printers).
If they could shed historical ballast that the established vendors cannot get rid of due to compatibility constraints, it would make sense. But they need to be compatible x86 or ARM (maybe RISC-V) anyway, or else you're talking about boiling the ocean.
NUVIA can't do everything and be everything because they're a startup. They have to make choices. I'm incapable of thinking they'd choose anything but ARMv8. Transmeta learned that x86 is really hard and then you get to compete with Intel. RISC-V is incomplete and I haven't heard much from Esperanto of late.
So I think they're going to build a cloud worthy ARMv8 chip. There's a lot of historical ballast to be shed in just getting rid of x86. Indeed, I think that is both their market opportunity and their market risk.
I don't think the world revolves around Spectre, Meltdown and MDS but designing a microarchitecture that makes them impossible would be gain a lot of market good will.
But at the end of the day, they have outperform Intel.
For one example, if there are many workloads reading same shared resource, the CCPU (Cloud CPU) can use that to reduce access overhead. Even the program might be a shared resource - you can easily run several dozens of copies simultaneously in vectored fashion.
This might be exploited even further - the queueing operations, for example, can be transformed into parallel scans, yielding less than 1 clock cycle for a synchronization on the work queue.
Etc.
Cray tried similar things to somewhat good results with their Athlon-pin-compatible accelerators. But I think we can do and get more.
Understatement of the year. The number of man hours Intel has sunk into optimizing each of these boggles the mind. And this is coming from someone who worked there for 2 years.
Unless there's a minimum friction to migrate, most companies won't make the effort even if they can save a few $100 per server. It takes me back to Intel's VLIW attempt with Itanium/EPIC. Even when they got compilers up to snuff, too many high end tasks (video encoding) either required special instructions or were written in assembly that couldn't easily be ported to EPIC instructions.
It made me wonder if it was a VIA spinoff; maybe licensing some of the interdependently developed x86_64 IP they bought with Cyrix from the entire Via C6 era.
I'd much rather translate from that than create the compiler backend and the entire toolchain.
[1] It's not perfect alas, as there are still many things the compiler knows that is lost in translation, such as the complete static alias sets of memory operations, true range limits on values, etc. For some of this we burn power today trying to (re-) discover at run-time (like the memory disambiguator).
I, for one, would appreciate an explanation of how WASM differs from the original JVM, if someone has a moment.
The JVM is also very far from language independent. Good luck making an efficient mapping for any language that doesn't look like Java.
WASM in contrast is designed explicitly for JITting, in fact, one-pass JITting is possible. The data structures maps well to the compiler backend. The representation is restricted in ways to avoid requiring expensive analysis, most notoriously, no arbitrary branching.
If you're looking for what WASM does differently resulting in the JVM not being a good fit it's mostly around browser integration. WASM runs as part of the browser, designed so that the same VM running the JavaScript portions of a page can run the WASM portions without enormous modification. It's also more tightly coupled with the data model of the browser but it doesn't have full direct access to everything, still better than ferrying everything to a second VM.
As for why people are so excited to use it outside of the browser it probably has to do with the level of support and investment around engines like V8 and that it's not designed for a particular language.
It’s a transitional step between the current mass market mode of chip productions selling one size fits all, and the way the market is heading towards task specialized hardware.
Data crunching Fungible is already on that
Distributed services a lot of fan in and fan out some kind of chips that can combine IO networking and moderate general computing instructions can be useful
Massive code data storage
Catching servers?
Overall, I see no reason to take AMD Intel heads on... It's not necessarily anyway, no one needs a third x86 player. We want to have true architecture disruptor...
Nybody want to
NUke them from high orbit?