Grayskull is supposed to be A100 performance for $1000, with some cool features (horizontally scalable by plugging them to each other over ethernet, C++ programmable, sparse computation, etc.)
Over the last couple weeks, I've seriously considered purchasing AMD Instinct MI50 (32GB HBM2, 1 Gbit/s) card which goes for under a $1000. I know it lacks Tensor Cores, and can only offer 53 TOPS which sounds silly compared to Grayskull's 600 TOPS. However, isn't it the case that for the most part these cores are idle, & waiting for memory still? At any rate, you're not going be able to run, say Llama 30B which won't fit already into 16GB but must fit comfortably in a 32GB system— on a single card. But perhaps most importantly, amdgpu driver is actually open source and allows PCIe passthrough unlike other vendors so considering all of the above, it almost seems like a no-brainer for a trusted computing setup.
I wonder if my logic is correct.
MI50: https://www.amd.com/en/products/professional-graphics/instin...
Transformers naturally have a big part of the network that is unused on any particular token flowing through the model. You could see that by how little RAM ended up being used in llama.cpp when they moved to mmaping the model.
So my understanding is that the tenstorrent cards are drastically more efficient, even on "dense" models like transformers because of the sparsity of any specific forward pass.
Also: I wouldn't bet on AMD accelerators for ML. They've disappointed every time. I would trust Jim Keller, whose every project in the last decade ended up being impactful.
I think the internet access is just to download the driver. It's not some sort of DRM setup where it needs to be always-online.
Presumably the architecture keeps the model in CPU RAM and shuffles it dynamically to the PCIe card network? I'm guessing here.
Whatever quirks come out of their hardware, I want to keep an eye on it.
I also think that comparing the current generation of models, which are built and trained to maximize GPU or TPU bandwidth, could be improved if someone architected models to maximize their model to greyskull advantages. Given PyTorch runs on it, I don't think it'd be too hard to do.
Dense transformers (GPT-3 AFAIK is dense) don't.
I spoke with a PM at TT and he told me an important idea is that we're spending a lot of electricity multiplying things with zeros.
From what I've read that was just an error in reading memory consumption after switching to the mmap version and not more memory efficient at all in the end.
The author of the mmap patch chimes in here:
EDIT: Ahhh a bit below on the page it is written "13.5 million AI-optimized cores" and there it's plural. So it was probably just a mistake.
e.g.
"He commanded a ten thousand man army." (not men)
"Andromeda, a trillion star galaxy, is 2.5 million lightyears away." (not stars)
etc
> The Edge TPU ... supports only TensorFlow Lite models that are fully 8-bit quantized and then compiled specifically for the Edge TPU.
They even wrote "The products offered by Google are unrelated to the products offered under the CORAL trademarks" in the page footer.