Tachyum’s Prodigy CPU Specs
wccftech.com
wccftech.com
Interesting bit I found researching them: https://www.nextplatform.com/2020/04/02/tachyum-starts-from-...
The processor pipeline has its out of order execution handled by the compiler, not by hardware, so there is some debate about whether this is an in order or out of order processor. Danilak says that instruction parallelism in the Prodigy chip is extracted using poison bits, which was popular with the Itanium chip which this core resembles in some ways and which are also used in Nvidia GPUs. The Prodigy instruction set as 32 integer registers at 64-bits and 32 vector registers that can be 256 bits or 512 bits wide, plus seven vector mask registers. The explicit parallelism (again, echoes of Itanium) is extracted by the compiler and instructions are bundled up in sizes of 3, 8, 12, or 16 bytes.
exposed: The results of operations that take more than one clock cycle don't necessarily appear at their destination the cycle after the instruction is executed. Think branch delay slots but potentially for multiplies and loads too.
skewed: Loads, processing, and stores can happen on subsequent clock ticks so simple loops don't necessarily need prologues and epilogues.
A compiler can handle either of these for a particular CPU pretty easily but they tend to eliminate binary compatibility. Code morphing as in the Transmeta lineage like to use both of these. Other VLIWs with barrel multithreading, like the Hexagons DSPs in Snapdragon SOICs, don't need them.
EDIT: The Mill guys have a video on how this works on their system. They've got their own names for things for some reason but they do a good job of explaining how this all works: https://millcomputing.com/docs/execution/
In supercomputing, this sort of thing sometimes appears and works well (like the Pezy-SC2 chips that recently came out). However, they usually fail on general purpose computing tasks.
Ie. you transpile your x86 code to native code for your VLIW machine. Then you run that code for a few hundred clock cycles till bam "cache not ready on time exception" which fires when you try to execute an instruction that is expecting to read memory from the cache, but the cache isn't yet populated with that value. Then you re-run your transpiler which will produce new code which either does a better job of reading the data into the cache ahead of time, or issues different instructions which take more time and read data direct from RAM.
Remember a software transpiler sometimes has more information than a typical deep OoO CPU, because it can use a lot more memory for state (eg. remembering that a particular branch or memory access won't be cached), and it can even persist state across system reboots. It can also do far more expensive optimizations and save the results, something a deep OoO CPU can't do because all the optimizations need to be doable in hardware.
On Itanium you essentially had to double check your values for poison before operating on them in a lot of cases and that caused big performance penalties on conventional code. Other systems restrict the shenanigans you can get up to with your MMU and solve the problem that way but that means you can't run a traditional OS with mmaped files and memory paging and such. I'm not sure what the Tachyum people are doing or if they've got some clever idea to get around all of this.
"speculative load" can be seen as data prefetch to a register, but with a extra instruction to put before reading/using the register. This instruction check that the memory content has been received, otherwise the instruction stall the pipeline. It allows to reorder a load before a conditional branch.
To determine whether the load has been received, there is an additional structure in the microarchitecture that tracks speculative loads in-flight. When a context switch occurs, this data-structure is overwritten, so there are additional mechanisms to replay the load in such cases.
Did they make any performance claims about the cross-ISA support? It might just be QEMU port with qemu-user and TCG.
"Lots of money" and interesting challenge can be the due diligence, it's not like the product has to be realistic to get paid (it helps in the long run, but in the short run VC money pays the bills).
Hell, they might even believe in the product. Linus spent 6 years at Transmeta.
But if you actually look at the thing, it does not actually seem like an out-and-out scam, a 5.7GHz VLIW design is not even remotely in the realm of the impossible.
And rosy promises which end up crashing and burning are exactly what transmeta achieved.
They also talk about air-cooling four of the 600W model in a 2U chassis! I'm not convinced how possible it is to handle 600W for any reasonable socket size with air cooling at all - the watts per mm2 ratios are going to be very high - but the idea that you can air-cool 2.4kW from the CPUs alone (and at these speeds there are going to be a lot of other very warm components in there) in a 2U chassis is simply crazy.
Actually the article says it’s on N5P the same process as the Apple M2.
I think TSMC does give smaller companies access to cutting edge processes in small volumes but it won’t be cheap.
I’m still sceptical.
Besides, most of the work in reaching a certain clock speed or target can be owed to the foundry (in this case TSMC which is world-leading, certainly beating Intel on most metrics at the moment.)
Comparing to the over-tuned enthusiast SKU of 2018 is not a fair comparison for either. Also 600W is not impossible to cool, there are GPUs at that level of power for a while now.
> The processor pipeline has its out of order execution handled by the compiler, not by hardware
This is not a direct competitor to general purpose CPUs, which makes the frequency claims a lot more realistic but also completely useless as a basis for comparison.
I wonder how compilers improved at generating good ILP with VLIW designs. The Itanium suffered from poor performance because it was very difficult to generate optimally ordered code for it.
OTOH, with enough GHz, even a simple, inefficient, in-order architecture can be fast enough. That was what made RISC attractive in the 80's and 90's.
You’re comparing apples and oranges. The 9990-xe is on 14 nm process. This is on 5 nm. Beyond that, this appears to be an in order VLIW core. The 9990-xe is a deeply out of order core. In a modern processor, much of the power budget is taken up by the structures used for out of order execution, such as reorder buffers and schedulers. These structures often grow non-linearly with the amount of reordering capability of the CPU. Jettisoning them entirely saves huge amounts of power.
Doesn't anything confuse you?"
For example, it is possible, if we consider some fantastic scenario, in which some big rich country will totally prohibit all other architectures, and enforce all people to buy only this. But I only know big poor countries, who could consider such experiments.
Everything, even claims to be better than CPU,GPU and TPU at once.
But to make all claims in one business, in one solution, looks scam. Because every vector needs it's own very effective leader, and in real life they will compete for shared resources. It is extremely difficult to organize in such way, so will got overall system working smooth and be economically viable.
In best case, I think possible for this company, to deliver some sort of very good number-crunching accelerator, and only then, to consider something more.
But now I see too much self-confidence and too much hidden details.
For example, all commodity hardware suffers from slow commodity DRAM and buses.
It is possible to make complex solution, with custom DRAM (HBM), custom high speed bus, etc, but its net cost will be prohibitive for commodity market.
So need some killer app, which will tolerate so high cost for some more important reason. This could be something for big corp, or govt, or military, or anything, but I cannot see any fit, which is not already handled with already existed solutions.
For example, I could imagine, we got knowledge about 11-dimensional travel, but for this need too much computational power for current best hardware, so need something much more powerful, and it will pay off. But I don't see any signs of such knowledge.
Runs binaries for x86, Arm, and RISC-V in addition to native ISA
I am quite a bit skeptical. OTOH, if it were aarch64 compatible with 1024 bit SVE support, it could have a nice potential.
When we talk about speculation and out of order execution we imagine the CPU is a kind of giant dependency-graph engine but this isn't actually true. The genius of modern hardware design (although kicked off by Tomasulo in the 60s) is that you can "compress" this idea into a real circuit with a finite number of gates and SRAM etc.
The quality of the branch prediction lets you get away with a lot on a fairly dumb "throw stuff on a buffer, speculate across memory accesses, pop (commit) off the buffer when the pie's finished cooking" model inside the processor.
The draw of this and originally Transmeta is that you have an extremely wide dumb processor in front of a smart software frontend that can do much more complicated work and scheduling based on the all-important runtime information that static VLIW sorely lacked (and thus lead to Intel doing all kinds of stuff with Itanium, hence it ending up EPIC rather than a true VLIW spiritually).
Now, software is slow, the way Transmeta got around this is by having a physical cache (the "Tcache") for storing translated instructions in.
https://www.cs.cornell.edu/courses/cs6120/2019fa/blog/transm...
https://www.realworldtech.com/crusoe-intro/5/ Some reverse engineering from the time. Note the amount of nops in the firmware.
https://safari.ethz.ch/digitaltechnik/spring2019/lib/exe/fet...
It was very underwhelming.
NVIDIA. For example, the Carmel core, shipping on Tegra Xavier.
> The agreement grants to NVIDIA a non-exclusive and fully paid-up license to all of Transmeta's patents and patent applications,
https://linuxjedi.co.uk/2021/12/27/800mips-amiga-with-emu68-...
That aside, I don't really get this nostalgy for these systems. I don't care about Doom, or some port of Quake. While 68K assembly was much nicer for me than anything common today, what do I get from that without a usable Browser, Office, "daily driver" apps? Show me how to port Firefox, Chromium or something functionally equivalent to these, and how those perform! :-)
Or Blender.
(Or Android 68K! (Giggle))
Apart from the nostalgy factor, I suspect there would be no actual benefit from such a system. I doubt m68k would compare well to ARM or x64 in terms of compatibility or modern-app performance.
They are missing critical statements like "The emulation allows booting x64 Linux and gets XXXX score on [industry benchmark]".
They have no documentation on their native ISA, nor any indication of any software or even compilers ported to it.
I would guess the 4 year delay is them struggling to get the software side working with any reasonable performance. It isn't too hard to make a toy processor with a high clock rate in a simulator... But getting it to run linux and win benchmarks is much harder.
This triggers my bullshit detector however even delivering anything in this sector is deeply impressive so hat's off to them even if they're cheating as long as they ship.
They can maybe do that if the architecture is configurable and they are using different configurations for their CPU, GPU and TPU competitors, but then it's still highly implausible that they can deliver more than 2x improvement over competitors, unless perhaps if it also costs proportionally more (but even then, you get decreasing yields for larger chips and inefficiencies if you make an SMP system).
Unseating x86/AMD64 is going to be hard because of the inertia they have. Whatever you make, even if it's faster, at least has to be able to run x86/64 binaries at a reasonable speed or it won't get wide adoption, unless it's a very captive market (mobile/Apple).
I'm skeptical of their claims but I would just say, the M1 has proven that if you have an appetite for throwing away compatibility and a great engineering team you can accomplish quite a bit. It seems you definitely can throw away the baggage that comes with needing compatibility and make fast, successful processors.
Might have something where hand-tuned assembly code can beat top tier CPUs and GPUs for very specific workloads (Although I guess it would have to be something straight-line enough that it can be statically scheduled, while not falling into the bucket of things that GPUs blow through).
Think about how wide the execution unit of a modern Intel/AMD CPU is compared to it's average IPC, if you could magic up an instruction stream that filled every port, and you could remove all the Out-of-Order/reordering/hazard logic it would make the cores a fair bit smaller and power efficient while giving the "same" peak possibly performance. But more ALU ports don't seem to the the bottleneck in current consumer software - keeping them filled is. Hence all the OoO/multithreading/big caches.
And the "Sufficiently Advanced Compiler" making up for OoO scheduling on the hardware itself has been a pipe dream for many years. Maybe this will be the one to break through the apparent barrier? But it's not like there haven't been a LOT of money and very clever people working on this exact thing before - and they have all fallen by the wayside compared to current large OoO CPU cores.
So to me it's just another cookie cutter attempt at VLIW, nothing special.
I agree that the VLIW approach will likely fail to deliver in most other applications, but I do not see the product marketed to them either.
Modern high-end Out-of-Order processors feature over 160 physical integer registers. This number is high because OoO processors use a very large instruction window, and every instruction in this window have to be free from register name false-dependency (Write-after-Read and Write-after-Write dependency). That's the register-renaming job to introduce extra physical register to hide false-dependency. If they really emulate the behavior of an OoO processor we should expect more registers.
The large instruction window form OoO high-end processor provides latency tolerance, with VLIW processors, to increase latency tolerance, we usually increase the register count and employ more aggressive static scheduling techniques.
It is known that register-renaming technique allocates on average more registers than necessary, due to early allocation and late release, so with static scheduling it should be possible to make more efficient allocate/release. But here the register count difference is too large, low register allocation/release efficiency of register-renaming is not enough to explain the difference.
As a comparison, Transmeta VLIW core has 64 integer register and 48 integer shadow-register which is equivalent to the physical register count of OoO processors of its time.
So either they don't really emulate an out of order processor, or they do some magic, or they don't really perform that well.. (or documents that mention 32 integer registers are incorrect)
The most likely hypothesis to me is that their benchmarks numbers are based on a selection of benchmarks that can exploit their vector unit, and other uncompetitive bench results have been left out. Meaning that performance on typical x86 general purpose applications would be extremely low with this processor.
[1] https://www.tachyum.com/media/pdf/tachyum-tries-for-hypersca...
Hypothetically you can do all kinds of things in software wrt to speculation and changing ISA on the fly.
So even if think Russians are not smart enough for such things, but also unsuccessful, how small private company could success without huge support?
Fake news, yeah maybe. What makes me incredulous is how bad the Tachyum name is. Although people without multiple startups before them can't come up with good names, like https://fgemm.com (coming soon). Apart from that!
Wow, that's high.
How can something that should e.g. Visual Studio whose solution should easily fit into RAM take about 5 seconds to start, 5 seconds to search the MRU and 10+ seconds to load intellisense.
How can "search the combo box" on Octopus deploy take notceable seconds sometimes even though it is purely an in-browser operation?
That's just two examples, there are plenty of others.
Sorry hardware guys!
This was published nearly 33 years ago:
OUSTERHOUT, J. 1989. Why aren’t operating systems getting faster as fast as hardware? WRL Tech Note, (TN-11)
I guess the most obvious parallel is that current OSs were designed for hardware that worked the way it did 30 years ago and now that those bottlenecks have disappeared, it is not trivial to rework the design to take advantage of these things.
I suppose the OS suffers from the tragedy of the commons in that each application developer assumes access to enough RAM and CPU to just work without considering how their large resource usage might affect all other applications/services on the system.
You could start all three in three threads and have everything finish in 10+ seconds, provided you have cores and memory bandwidth available.
Maybe Microsoft is providing too good desktops for the Visual Studio team. ;-)
I frequently advocate developers should, as an exercise, work on machines with more cores and lower single-thread performance. Core counts are only going to increase and being able to put all cores to work at the same time is an important competitive advantage.