I wonder why they are starting with these models. One could speculate that they are going for power efficiencies but that does not quite add up entirely.
I wonder why they are starting with these models. One could speculate that they are going for power efficiencies but that does not quite add up entirely.
The grayskull cards only have 8 GB of RAM and don't have a fast enough memory to NoC bandwidth to make working with multiple cards practical. The next generations, that is wormhole and newer don't have this limitations and are specifically designed to work in server racks, see the galaxy system. [0]
The Grayskull really only is a devboard. It has some other quirks that will be improved by wormhole, like native 19 bit floats in their SIMT engines, instead of 32 bit, in wormhole.
The purpose of these boards seems to get people acquainted with their programming model. Not as in using board and model as a turnkey solution (if that happens, fine, but this is not the goal), but as in getting potential customers for future boards to learn how to make their own models or third party models run on the board. The more models they supported out of the box, the less well the goal of building buyer-side expertise in the programming model would be served.
They were building Greyskull in 2020/2021 initially. BERT_Large and similar were the SOTA at that time.
Disclosure: I work for a Tenstorrent customer.
Those models do seem to be the best currently available (IMO), so yeah power efficiencies for guaranteed common use cases should be a 0-day integration.