MatX: Faster Chips for LLMs
matx.com
matx.com
It would be be better to write a post about how you're making faster chips for LLMs, that everybody can learn something from.
A lack of information about the technology is understandable as it presumably doesn't exist yet. But I see that it's Google TPU and software people striking out on their own with the backing of a bunch of VCs and prominent researchers, and it seems pretty clear where it's going. Likely an evolution of or progression from what they were working on at Google, positioned for easy acquisition by any number of competitors.
I think George Hotz is right that ML ASIC companies need to invest in software far more than they typically do. I guess the plan here is to sidestep the need for generic software by focusing solely on transformers for text, maybe even only a couple of specific architectures. I don't really think that's a good strategy but it seems to be what they're describing.
> Approach:
> * We target just LLMs, whereas GPUs target all ML models. LLMs are different. Our hardware and software can be much simpler.
> * We combine deep domain experience, a few key ideas, and a lot of careful engineering.
The OP is something between a landing page and a job ad, not an in-depth article. This is not a borderline call!
I wonder what's the end goal.
The ML/AI world seems to be changing fast. So one model approach might become "old tech" if something better comes.
Is our current state of the models is stable so the LLM glory would be the same in 1-2 years? Or suddenly there'll be new approach making this "tech" obsolete?
What looks state of the art will probably look to people in 20 years how the 1885 Mercedes Benz car looks like to us in terms of efficiency, right now it's a bit of a bruteforce approach in general
they were asic for BTC at a way earlier stage
+they can always pivot in a year if the market changes too much
I think you are downplaying a bit the timescale of going from ideation to tape-out in ASIC design.
https://geohot.github.io/blog/jekyll/update/2023/05/24/the-t...
He's still on the AMD drivers train, judging by his Twitter post from 4 days ago, so we'll see where things go.
https://twitter.com/realgeorgehotz/status/168616581138659737...
I feel like AMD's 7900 XTX strategy would be good too: a bunch of small, cheap (LPDDRX?) memory controller dies to form a massive bus for a central compute tile.
I don't see MatX ending up any different than the legion of startups that have come already - either they get acquired by a bigger player, or they fade into obscurity.
This is changing.
https://github.com/merrymercy/awesome-tensor-compilers
There are more and better projects that can compile an existing PyTorch codebase into a more optimized format for a range of devices. Triton (which is part of PyTorch) TVM and the MLIR based efforts (like torch-MLIR or IREE) are big ones, but there are smaller fish like GGML and Tinygrad, or more narrowly focused projects like Meta's AITemplate (which works on AMD datacenter GPUs).
Hardware is in a strange place now... It feels like everyone but Cerebras and AMD/Intel was squeezed out, but with all the money pouring in, I think this is temporary.