Esperanto Champions the Efficiency of Its 1,092-Core RISC-V Chip
hpcwire.com
hpcwire.com
At this point I think we need to go back to descriptive old-school 70s company names like, "West Coast Microprocessor Solutions", "Digital Logic, Inc.", "Mountain View Artificial Intelligence Laboratories", etc.
You know, something that would blend into this map: https://s3.amazonaws.com/rumsey5/silicon/11492000.jpg
Edit: Looking at that map, some of the company names are fantastically generic! "Electronics Corporation", "California Devices", "General Technology", "Test International".
Edit: OK, I can see this has "network on a chip" architecture but I think that only answers some of my questions.
~100 MB of SRAM on chip, 8 GDDR busses to DRAM off chip.
It's purpose designed for parallel sparse matrix ML problems. It's more efficient than both a CPU and GPU at these, as well as faster in absolute terms, taking their numbers at face value.
The rest of those lanes come from MIMD techniques.
-------
CPU cores are more MIMD today than SISD because of out of order and superscalar operations. So honestly, I think it's about time to retire Flynn's taxonomy. Everything is MIMD.
The rest of those lanes come from MIMD techniques.
Not sure what you mean here. Are there places where different groups of kernels can simultaneous execute different code?
Yes. Lets take AMD's GCN / Vega, since I'm most familiar with it.
The Vega 64 has 64 "compute units" (CU for short), where a CU is the closest thing to a "core" that the Vega GPU has. So lets really look at how one of these CUs functions.
1. A kernel in GCN executes 64-wide SIMD assembly language. This is the programmer model, but this is not what's going on under the hood. This 64-wide SIMD is executed every 4 clock ticks, 16-at-a-time.
2. The CU has 4x16 groups of vALUs (where a vALU is a set of 64-wide registers and 16-wide arithmetic units). The CU also has 1x group of sALU ("scalar ALU" has 32-bit registers). A kernel has access to upto 256 vGPRs (each executing in SIMD fashion: 16 wide over 4 clockticks for the 64 lanes) + 103 sGPRs, which are shared. (For example: a function call is an sGPR. Because function calls are "shared" between all 64-lanes, its more efficiently implemented in sGPRs than vGPRs).
------------
There can be up to 40-simultaneous instruction pointers being executed by the AMD GCN GPU. Aka: occupancy 40. The exact instructions will be switching between sALUs (ex: branching instructions, function calls, push / pop from the stack), and vALUs (ex: multiply-and-accumulate, which will be SIMD-executed 64x in parallel over 4 clock ticks). (Really: each vALU has up to 10 instruction pointers that its tracking)
As we can see, the modern GPU isn't SIMD at all. Its executing multiple instructions, from multiple different instruction pointers (possibly different kernels even) in parallel.
Even at a minimum number of threads... there's 4-wavefronts per CU getting executed (4-instruction pointers, one for each vALU). That's MIMD in my opinion, since there's no smaller unit than the CU in the whole of AMD Vega.
Just because its executing SIMD kernels / SIMD assembly code doesn't mean that the underlying machine is SIMD.
---------
NVidia is similar: except with 32-wide SIMD programming model and 32-occupany per SM (though these magic numbers seem to change every generation) NVidia is also beginning to implement superscalar SIMD (2-instructions per clock tick, if those instructions go to different execution units).
So I'd also classify modern NVidia as a MIMD machine.
--------
CPUs of course have superscalar units: something like 4-way decode and something like 6+ instructions per clock tick (if executing out of the uop cache. 4-instructions per clock tick otherwise).
If you have fewer than 4-wavefronts running at any given time, you've got an underutilized CU. That's 256-SIMD lanes you need at a minimum to fully utilize any CU from an AMD Vega.
They happen to scale up to 10x wavefronts per vALU (aka: 40-wavefronts maximum). At least, if your kernels use few enough registers / __shared__ memory. But even if we remove all the SMT-like features in an AMD Vega CU, we still have 4x instruction pointers being juggled by the underlying system.
Which should be MIMD by any sane definition. (4 instruction pointers is MI or "multiple instructions", each of which is a SIMD instruction, so MD is also happening).
---------
Note that SMT / Hyperthreading seems to be considered MISD in Flynn's original 1966 paper. So SMT + SIMD == MIMD, in my opinion. These modern GPUs didn't exist back in 1966, so we don't know for sure how Flynn would have categorized today's computers.
But honestly, I'd say its somewhat insane to be reading a paper/organization scheme from 1966 and trying to apply its labels on computers and architectures invented 50-years later.
So each CU runs natively 4-wavefronts (and scales to 40-wavefronts as resources allow). Each wavefront is a 64-wide SIMD, with 256-SIMD lanes total running per CU.
That's 16,384 "threads" of execution on a 4-year old GPU before all cores are "lit up" and utilized... with the option to have up to 163,840 "threads" at max 40-way occupancy (useful if you have low register usage per kernel and a dependency on something that's high-latency for some reason). These are mapped into 4096 physical SIMD-lanes / "threads" that physically execute per clock tick.
---------
At any given time, there are only 256 instruction pointers actually being executed. Which is where and how a GPU manages to be efficient (but also the weakness of a GPU: why it has issues with "branch divergence").
EDIT: The general assumption of the SIMD model of computers (like GPUs) is that line#250 is probably going to be executed by many, many "threads". So these "threads" instead become SIMD lanes, and are batched together to execute line #250 all together, saving power on decoding, instruction pointer tracking, the stack, etc. etc.
As long as enough threads are doing similar things all together, its more efficient to batch them up as a 64-wide AMD wavefront or 32-wide NVidia block. The programmer must be aware of this assumption. However, the underlying machine still has gross amounts of flexibility in terms of how to implement it. So it could be an MIMD machine under the hood, even if the programming model is SIMD.
-----
There's also the issue that when you KNOW 64-lanes are working together, you have assurances of where the data is. Things like bpermute and permute instructions can exist (aka: shuffle the bytes between the lanes) because you know all 64 lanes are on the same line of code. So in practice, cross-thread collaboration (such as prefix-sum, scan operations, compress, expand...) are more efficient on SIMD model than the equivalent mutex/atomic/compare-and-swap style programming of CPUs.
Something like a Raytracer (bounce these 64-simulated light paths around), and organizing which ones go where (compress into hit_array vs compress into miss_array) are fundamentally more efficient on the GPU/SIMD style programming, than the heavy CPU-based threaded model.
I should make it clearer - I'm looking to run that's data-intenive, structurally MIMD but with some SIMD aspects. I'm trying to figure out the most cost-effective chip with which to do this.
The other question is whether chips like Esperanto's product have effective primitive for reduce operations and how much general memory through-put they have compared to a GPU.
Don't miss the most important detail: low voltage. By running the Minions at a low voltage, you are reducing performance, but disproportionately increasing the efficiency (up to a point). It's an interesting trade-off of power, area, performance, and efficiency.
Add: it's clearly not indended to drive a GPU, but is presented as an accelerator. They have mentioned that it can also run standalone, thus, can be configured both a PCIe host and target.
I’d love to see this misused as a workstation. We’d need to do something with `htop` though, because showing all these cores would require an insanely tall terminal.
People definitely do that, the voltage you chose to power your chip is always a consideration, it's one of the variables used by laptop manufacturers to manage thermals, and it's also pretty common to lower the CPU voltage so that you can run a desktop system with passive cooling.
Anyway my point is that you can lower the voltage on any chip so the fact that this runs at a low voltage isn't noteworthy.
To the OP: yes this is trading silicon (aka area) for efficiency. That’s the whole point.
> The ET-Minion core, based on the open RISC-V ISA, adds proprietary extensions optimized for machine learning. This general-purpose 64-bit microprocessor executes instructions in order, for maximum efficiency, while extensions support vector and tensor operations on up to 256 bits of floating-point data (using 16-bit or 32-bit operands) or 512 bits of integer data (using 8-bit operands) per clock period.
So sounds like at least 8736 SP FP operations per cycle.
Brilliant move
> the entire chip would consume just 8.5 watts ... Ditzel said, one chip would take about 20 watts
This seems to contradict each other.
> But if, instead, they followed the peak of the energy efficiency graph, the entire chip would consume just 8.5 watts [...] and operating at about 0.4 volts, Ditzel said, one chip would take about 20 watts.
"the peak of the energy efficiency graph" isn't stated explicitly, but is at around 0.32 volts.
Basically they want to use up all 120W of a PCIe slot, at the highest efficiency possible. They could have gotten higher efficiency (at 8.5W per chip), but that would have resulted in not being able to use the full 120W and thus actually having worse performance overall, even though it's more efficient.
Right, at the end of the day they were constrained due to not being able to fit more than 6 chips on a single PCIe card. If they could fit 14, they would've been able to make each chip use as little as ~8.57W while still using up the 120W available to the card.
- CPU on chip
- System on chip
- Network on chip
I'm looking forward to the next X on chip which I'm not aware of.
Compilers are hard, device support is hard, the compiler community is small and closed source compilers quickly become weird tech islands.
The rationale for open-sourcing here, in addition to the general recruiting/hiring benefit, is that we want vendors to target a common interface so that it's easy to make direct comparisons amongst different hardware.
I'd say, though, that ML is moving somewhat away from the "graph compiler" approach. PyTorch (and users' experience with TPUs/XLA vs GPUs) has suggested that static graphs aren't desirable for usability or necessary for performance. These days, I'd say write a PyTorch device backend and a fast kernel library.
More extensions are in the process of being ratified before the end of the year. The big one is the Vector extension, which allows code using vectors/SIMD to execute completely unchanged on machines with vector registers anywhere from 128 bits to 64k bytes in size.
I know of one apparent implementation compatibility bug that's just been diagnosed. GeekBench, for whatever reason, is using the fairly new FENCE.TSO instruction in their code. A bit pointless as I don't think they are yet any RISC-V cores that implement TSO memory semantics, so a FENCE RW,RW would do just as well. The FENCE opcodes have a largish field with the lower bits giving pred and succ R, W, I, O bits. In the base FENCE instructions the upper bits are all zero. Future FENCE instructions are supposed to be designed so that if a CPU ignores the upper bits then they devolve to some slightly stronger standard fence. So FENCE.TSO for example is FENCE RW,RW with one of the upper bits set. It seems that the Alibaba C906 core (in the Allwinner D1 chip on the Nezha board) is not treating unknown upper bits in FENCE as if they are all zeros as the spec says, but is instead giving an illegal instruction trap.
Fortunately, this can be worked around by adding a FENCE.TSO emulation handler (or FENCE with upper bits set in general) to OpenSBI, alongside the emulation of things such as misaligned loads and stores. These can be safely present in the M mode software for every CPU type as they will never be triggered if the hardware handles those things directly.
Of course the trap and emulate performance penalty will be far far greater than any minor performance improvement from using FENCE.TSO instead of FENCE RW,RW on a (hypothetical at this point) machine that actually implements the TSO extension.
So you end up evaluating both sides of the decision branch and adding the results. But this is fine if you've got a dumb number of cores. And often traditional cpus wind up evaluating both branches anyway.
That's actually really overstated. Evaluating both sides isn't really something CPUs tend to do, but instead predict one path and roll back on mispredict. This is because the out of order hardware isn't a tree fundamentally, but generally better thought of a ring buffer where uncommitted state is what's between the head and tail. Storing diverging paths is incredibly expensive there. I'm not going to say something as strong as "it's never been done", but but I certainly don't know of an general purpose CPU arch that'll compute both sides of an architectural branch, instead relying an making good predictions down one instruction stream, then rolling back and restarting when it's clear you mispredicted.
This concept, of exploring like a tree vs a path was explored under the name Disjoint Eager Execution. You know what killed it? Branch predictors. In a world where branch predictors are maybe only 75% effective, DEE could make sense. We live in a world where branch predictors are far better than that. So it just isn't worth speculating off the predicted most likely path.
The more cores you have, the more memory is needed to keep the chips processing.
Datastructure and locality matters a lot for GPGPU programming. While PCI-Express is really fast, it's got a lot of latency and is limited.
I'm curious though, apart from huge NNs and raytracing scenes, what use cases call for so much RAM. I mean, apart from content lookup tables?
For NN, it’s quite easy to use a huge amount of memory by making back-propagation chains huge, which is natural for deep models or a recurrent architecture. The model doesn’t have to be huge in a conventional sense; it just has to maintain state e.g., the recurrence.
So, a big image model on video classification/segmentation (e.g., self-driving cars) is probably the ideal combination of memory consumption.
I personally think those chips would be an absolute monster for solving MILP problems as they tend to have enormous parallelism and a lot of linear algebra (in particular simplex iterations).
However there is no hype for funding in Mixed-Integer Linear Programming so machine learning it is.