1. https://docs.tenstorrent.com/tenstorrent/v/tt-buda/hardware
1. https://docs.tenstorrent.com/tenstorrent/v/tt-buda/hardware
Upside is each core has its own code and is fully Turing complete and independent of eachother. You can handle conditionals much better. And you lose the latency of having network hops for workers.
Downside is you need to break down your process to map onto specific nodes and flows.
(Assuming it is in fact manycore - which is not the same as multicore)
Because you can have a core set up for each branch and just pass to that vs. “context-switching” your core to execute the branch that ends up being taken?
> Because you can have a core set up for each branch
and
> because each core has its own instruction counter
You can have each core be a tighter loop because you can limit the types of cases it handles.
But you also have your own instruction program counter you can take all sorts of branches in a way that you can't with SIMD, because in SIMD as the name implies you get only a Single-Instruction to deal with Multiple-Data. So if you need to change the behaviour depending on the exact value of the data, you are better off using separate cores than wider vectors.
Your core needs to be fully programmable so you can do things like kernel fusion. The simplest form is to load quantized weights and dequantize them to bfloat16 as you go. Llama.cpp and it's gguf files support various types of quantization and most of them require programmability to efficiently support them.
CPU design moved away from this analogy a long while ago because the tasks being done with CPUs involved more dynamic control flow structures and arbitrary workloads. But workloads that are linear batches of brute force math don't need that kind of dynamism, so gridded designs become fashionable as a way of expressing a configurable pipeline with clear semantics - everything on the grid is at a linear, integer-scalable distance, buffers will be of the same size, etc.
Hopefully some of y'all tinkerers with time and dough can bear some of these ideas to fruition, keep Nvidia on their toes ;)
64gb is a good RAM amount IMO, cheap yet still vastly underutilized since we play to the LCD of users... guessing Linux will be able to make that pivot much faster/so little baggage.
Plus..."Grayskull"
What about them do you like as a design decision? (genuinely curious, as again, I don't understand it)
It doesn’t have 64gb of RAM, it has 8. The system requirements otherwise need 64gb of RAM for model compilation
Are you saying proximity here more than offsets this vs. e.g. each core having its own cache as I think they do in a "normal" CPU? And if so, is this more true of ML inference workloads than other workloads, for some reason?
https://www.realworldtech.com/includes/images/articles/snbep...
https://en.wikichip.org/wiki/intel/microarchitectures/sandy_...
e-cores do have a CCX/core cluster, but the clusters themselves go on the ringbus lol