Tenstorrent unveils Grayskull, its RISC-V answer to GPUs
techradar.com
techradar.com
- Grayskull e75 | drawing 75W | 96 Tensix cores | 1 GHz clock | 96 MB SRAM | 8GB LPDDR4 @ 102.4 GB/s | $599
- Grayskull e150 | drawing 200W | 120 Tensix cores | 1.2 GHz clock | 120 MB SRAM | 8GB LPDDR4 @ 118.4 GB/s | $799
It will be interesting to see their inference performance, compared to graphics cards. Will they be interesting for home labs?
I found one interview with unboxing a preview version (if I understood correctly), with some background info, but no performance numbers: https://morethanmoore.substack.com/p/unboxing-the-tenstorren...
I'm probably dreaming; new CPU core designs are one thing, a completely new memory design is likely high fantasy..
Are the optimizations for inference basically identical to graphics?
I imagine the workloads are different, and that inference has been the main market for years now...
It will have to compete on price/memory in the crowded inference space.
This is purely a developer kit, letting people get used to the hardware configuration and the software stacks before the big hardware comes later in future generations.
I'd like to know what different precisions/ quantizations if any are supported. For LLMs 8GB is fine for playing around if weights are quantized. And as the article mentions it's more than enough for lots of computer vision models.I don't know how and with what performance those are supported.
FP4 has 221 TFLOPS on the Grayskull e75, and 332 TFLOPS on the Grayskull e150. [0]
While wormhole has 82 INT8 TOPS per card. [2]
I haven't found any other numbers.
Edit: Some more info about data formats: https://docs.tenstorrent.com/tenstorrent/v/tt-buda/dataforma...
[0] https://docs.tenstorrent.com/tenstorrent/add-in-boards-and-c...
[1] https://tenstorrent-metal.github.io/tt-metal/latest/tt_metal...
Also the CPUs on Grayskull is 32bit. Memory is addressed through the bank address so it works for now. But they'll have to upgrade to 64bit soon.
I always assumed that was what killed Xeon Phi. Hard to justify a GEMM card in and of itself. However clever you get, NVIDIA will be back next year with twice as much bandwidth.
If you haven't yet, buy some bitcoin and stop letting the government steal your hard earned money through inflation.
1. The hardware isn't nVidia's primary moat. CUDA is. Currently all machine learning tooling assumes you have a CUDA device.
2. The hype cycle may be real, but nobody really knows how much better AI will become. The uncertainty is driving the hype (which is obviously on the optimistic side), but it could well be that models become so good that everyone "needs" an AI chip like how we "need" a smartphone today.
3. The gaming/graphics market is probably going to slowly shrink. There's only so many pixels you can push to the eye. There's very real diminishing returns there. At some point people literally won't be able to tell the difference between the newer/better/faster graphics and the older generation, at which point the GPU vendor's margins drop significantly.
People really don't appreciate this. So many times I see people comment that AMD should implement CUDA and not try to get software to move standard.
AMD have a massive task ahead because they NEED to get CUDA unseated just to play on a level playing ground, let alone win.
Melting connectors, obscene pricing, marginal performance increase, zero availability; then when hardly any of that was addressed, Ti versions release which are equally expensive, and equally lacklustre. Combine that with patchy support for frame generation and 30xx series cards still going really strong, and you’ve got not a lot of incentive to upgrade.
In units sold?
I’m sure they don’t mind having both revenue streams of course! But on the other hand, they have to serve two sets of interests.
I put those in quotes, because theoretically both can do training, it's just that grayskull isn't well suited for it, because of the internal fp19 format, and no good support for scaling to multiple cards. wormhole will have internal fp32, and is designed to scale out with across multiple cards, and servers.
At least that's how I understood it.
In LLM-land we've rapidly gone from training bespoke models to doing fine tuning to RLHF to zero-shot prompting. The better the underlying model the more you can do without additional training, so hardware that fails to scale up to the largest training runs will have limited practical utility despite technically supporting training.
And yes, I know, I've been working on LLM pretraining for about 4 years now, since 2020. The number formats themselves so far are mostly scale invariant or improve with larger scale - you can quantize a larger model and see less performance drop than a smaller model.
Indeed.
GPGPU means "using a GPU for things that are General Purpose (GPGPU)". The aforementioned IC is not a GPU that is used for general-purpose computations, but a specialized chip for AI computations.
The H100s feature a minuscule amount of ROPs and TMUs (mainly for graphics); and doesn’t come with NVENC/DEC, etc. DirectX, Vulkan is not supported.
You’ll just see more and more tensor cores, optimised CUDA cores for bfloat, etc.
https://www.servethehome.com/wp-content/uploads/2023/10/NVID...
Tbh it’s mildly surprising they even removed NVENC considering the overall size of the chip (in the absolute we are only talking about low-single-digit mm2 savings) and then H100 is still advertised and has features targeting VM graphics/visualization still… remember they also still put a full graphics pipeline with ROPs/TMUs on the chip, just no actual display hardware.
These days, simulation runs basically nonstop on all the FPGAs the company is able to throw at it until tape out. Then, in post-silicon validation, you have engineering teams working 24/7 to validate the device on the bench. Finally, test engineering tries to achieve yield in HVM.
Anything that you don't need on the spec sheet or the designer insists on is cut to save time to market.
Apple initially developed OpenCL with Khronos as a similar-ish analog for GPU acceleration a while ago. There were a few partners that invested in it, but it suffered the same lack of demand as CUDA and languished for a while. Now Apple doesn't support OpenCL, so the purpose of a multi-platform acceleration library has kinda been scuttled. Nvidia played their cards well, and their competitors are going to feel the pain for a while unless they work together again.
Not much. GPUs became array processing engines in the early 2000s and the amount of residual fixed function graphics hardware is tiny now.
The biggest thing is that a lot of network evaluation is being done at lower precision (BF16/ FP8 / ternary) now, but my understanding is that training the networks still requires FP32. As these companies are really targeting the training market the high precision modes needed for graphics can't be thrown away. If a company was to target the end user, then you could optimise for lower precision, but end users find dedicated neural net acceleration a luxury they can get for free by buying a GPU.
That being said, it’s a Grayskull Processing Unit. ;)
1. https://docs.tenstorrent.com/tenstorrent/v/tt-buda/hardware
Are you saying proximity here more than offsets this vs. e.g. each core having its own cache as I think they do in a "normal" CPU? And if so, is this more true of ML inference workloads than other workloads, for some reason?
https://www.realworldtech.com/includes/images/articles/snbep...
https://en.wikichip.org/wiki/intel/microarchitectures/sandy_...
e-cores do have a CCX/core cluster, but the clusters themselves go on the ringbus lol
Hopefully some of y'all tinkerers with time and dough can bear some of these ideas to fruition, keep Nvidia on their toes ;)
64gb is a good RAM amount IMO, cheap yet still vastly underutilized since we play to the LCD of users... guessing Linux will be able to make that pivot much faster/so little baggage.
Plus..."Grayskull"
What about them do you like as a design decision? (genuinely curious, as again, I don't understand it)
It doesn’t have 64gb of RAM, it has 8. The system requirements otherwise need 64gb of RAM for model compilation
Your core needs to be fully programmable so you can do things like kernel fusion. The simplest form is to load quantized weights and dequantize them to bfloat16 as you go. Llama.cpp and it's gguf files support various types of quantization and most of them require programmability to efficiently support them.
Upside is each core has its own code and is fully Turing complete and independent of eachother. You can handle conditionals much better. And you lose the latency of having network hops for workers.
Downside is you need to break down your process to map onto specific nodes and flows.
(Assuming it is in fact manycore - which is not the same as multicore)
Because you can have a core set up for each branch and just pass to that vs. “context-switching” your core to execute the branch that ends up being taken?
> Because you can have a core set up for each branch
and
> because each core has its own instruction counter
You can have each core be a tighter loop because you can limit the types of cases it handles.
But you also have your own instruction program counter you can take all sorts of branches in a way that you can't with SIMD, because in SIMD as the name implies you get only a Single-Instruction to deal with Multiple-Data. So if you need to change the behaviour depending on the exact value of the data, you are better off using separate cores than wider vectors.
CPU design moved away from this analogy a long while ago because the tasks being done with CPUs involved more dynamic control flow structures and arbitrary workloads. But workloads that are linear batches of brute force math don't need that kind of dynamism, so gridded designs become fashionable as a way of expressing a configurable pipeline with clear semantics - everything on the grid is at a linear, integer-scalable distance, buffers will be of the same size, etc.
The basic system consists of the cards consists of a bunch of Tensix cores and shared memory:
> Each Tensix core contains a high-density tensor math unit (FPU) which performs most of the heavy lifting, a SIMD engine (SFPU), five Risc-V CPU cores, and a large local memory storage. [0]
> The cores are connected with two torus-shaped, going in opposite directions. [0]
The RISC-V cores in the Tensix cores, are tiny rv32i cores, that can control the FPU, SFPU, and are also used to prepare/move data.
The FPU, does "dense tensor math", so I think it's probably a matmul engine, but I don't know any more specifics. [1]
The SFPU is a more general purpose SIMT engine, that can be driven from the RISC-V cores.
There is a SFPU simulator you can play around with on their github. [2] See the low level Kernels example for how the programming model works. [3]
The grayskull SFPU has 4 general purpose LRegs, which each hold 64 19-bit values. Wormhole has 8 general purpose LRegs, which each hold 32 32-bit values.
I've been told that wormhole SFPU has a ~3x IPC increase over grayskull, and a few new SFPU instructions.
You can probably find out more by browsing the docs and rummaging through the github repos. [4]
[0] https://docs.tenstorrent.com/tenstorrent/v/tt-buda/hardware
[1] https://docs.tenstorrent.com/tenstorrent/v/tt-buda/terminolo...
[2] https://github.com/tenstorrent-metal/sfpi/blob/master/tests/...
[3] https://tenstorrent-metal.github.io/tt-metal/latest/tt_metal...
I wonder why they are starting with these models. One could speculate that they are going for power efficiencies but that does not quite add up entirely.
The grayskull cards only have 8 GB of RAM and don't have a fast enough memory to NoC bandwidth to make working with multiple cards practical. The next generations, that is wormhole and newer don't have this limitations and are specifically designed to work in server racks, see the galaxy system. [0]
The Grayskull really only is a devboard. It has some other quirks that will be improved by wormhole, like native 19 bit floats in their SIMT engines, instead of 32 bit, in wormhole.
Disclosure: I work for a Tenstorrent customer.
They were building Greyskull in 2020/2021 initially. BERT_Large and similar were the SOTA at that time.
Those models do seem to be the best currently available (IMO), so yeah power efficiencies for guaranteed common use cases should be a 0-day integration.
The purpose of these boards seems to get people acquainted with their programming model. Not as in using board and model as a turnkey solution (if that happens, fine, but this is not the goal), but as in getting potential customers for future boards to learn how to make their own models or third party models run on the board. The more models they supported out of the box, the less well the goal of building buyer-side expertise in the programming model would be served.
( 1 Tensix core = 5 RISC-V core : https://docs.tenstorrent.com/tenstorrent/v/tt-buda/hardware )
If you had to communicate it orally, what words would you add?
ADDED. I think I figured out what it means: a Grayskull e150 contains 120 'Tensix cores', each of which contains 5 RISC-V cores.
"TT-Metalium is a platform for programming heterogeneous collection of CPUs (host processors) and Tenstorrent acceleration devices, composed of many RISC-V processors. Target users of TT-Metal are expert parallel programmers wanting to write parallel and efficient code with full access to the Tensix hardware via low-level kernel APIs."
https://tenstorrent-metal.github.io/tt-metal/latest/tt_metal...
[1]
"The software stacks come in two varieties – a high level and a low level. The high-level is called TT-Buda, using higher-level APIs to get things up and running, along with interfaces into modern machine learning frameworks. The lower level is TT-Metalium, which provides fine-grained control over the hardware for custom operators, custom control, and even non-machine learning code. Tenstorrent states that there are no black boxes, no encrypted APIs, and no hidden functions." https://morethanmoore.substack.com/p/unboxing-the-tenstorren...
> It has more to do with the memory resources required during model compilation. The requirement will vary from model to model, so 64GB is a safe limit for all.
I imagine it's just to leverage bandwidth correctly? Or is there a necessary feature in there?
What am I missing?
I assume they are sold to hardware manufacturers and maybe software developers so that they check out and test their systems against real processor. Processors are probably manufactured in some relatively old process technology.
The real thing will be manufactured using different 2nm process according to the article and performance will be different.
Obviously immature ecosystem, expect a couple of years at minimum before adoption if it's the real deal.
Can they really compete in FLOP/$ with the likes on Nvidia even if it is a more bespoke architecture than a modified GPU?
* Everything is relative. It's hard to make even a simple microcontroller - but also it isn't hard, you know?
It's hard when you can't buy TSMC capacity
All the details you need!
And it can play games too.