Apple already have the "neural" cores, is that more or less what they are?
Could there be a theoretical LLM chip for inference that is significantly cheaper to run?
Apple already have the "neural" cores, is that more or less what they are?
Could there be a theoretical LLM chip for inference that is significantly cheaper to run?
Problem is, inference costs do not dominate training costs. Models have a very limited lifespan, they are constantly retrained or obsoleted by new generations, so training is always going on.
Training is not just matrix multiplications, given hundreds of experiments in model architecture, its not even obvious what operations will dominate future training. So a more general purpose GPU is just a way safer bet.
Also, LLM talent is in extreme short supply, and you don't want to piss them off by telling them they have to spend their time debugging some crappy FPGA because you wanted to save some hardware bucks.
Just curious what the current bar is here and which of the LLM-related skills might be worth building.
Being very good at fine tuning for a particular goal. Its much easier to learn fine-tuning, so standards are higher to stand out.
Being able to come up with architectural improvements for LLMs, aka the researcher path.
Wages start at $250k for grads at the big AI companies.
1. For BERT scale model, all you need is a good codebase from GitHub (I had some luck with this one [0]) and a few weeks of trial and error. Want to try training T5 or LLaMA, but don't have the resources needed. Of course training models with more than 100B parameters is another level of labyrinth.
2. Finetuning is mostly related to how well you understand the task and the data you are dealing with. Since the BERT paper focuses on the GLUE benchmark, I've become very proficient in fine-tuning GLUE and eventually got sick of it.
3. Made some architectural improvements to BERT, got decent results so I wrote a paper, and got rejected because the reviewers want a head-on evaluation against some well funded papers from Google.
4. Not in my country. Damn, I am envious.
Capital outlays are tied to the derivative of compute capacity, so even if training just flatlines, hardware spend will drop significantly.
As far as I can tell that gamble isn't work out particularly well for any of the startups but that might be money drying up before they've hit commercial viability. I know the hardware is pretty good for Graphcore, Cerebras and the software proving difficult.
I would say they are more like GPUs or DSPs: programmable but optimised for a specific application domain, ML/AI workloads in this case. Sometimes people call this ASIPs: application specific instruction set processors. While maybe not a very commonly used term, it is technically more correct.
As a rule companies should only do their own chips if they are certain they can solve and overcome the cogs problems that low yield and low volume penalties entail. If not you are almost certainly better off just eating the vendor margin. It is very very unlikely that you will do better.
That’s why having 96gb (with another 480gb or whatever) available via high speed interconnect is a big deal. It means we can train bigger models faster.
When talking about the current ML industry it's more like nobody wants to invest significant amounts of money in hardware that will be obsolete before it's even taped out.
If you want more efficiency gain than general matrix multiplication hardware, you need to start getting specific about the NN architectures the hardware will support.
Effectively it'd require the entirely memory controller and the cache, and scheduling. At point point you got most of the GPU w/ a stuck, non-programmable interface of a designated compute. Likely you'd never have to compete for advanced nodes as well.
I wonder what they’re doing with that hardware now.
Ethereum is designed to bottleneck on memory bandwidth (while being uncacheable) so at the end of the day the name of the game is how many memory channels can you slap onto a minimum-cost board. You won't drastically win on perf/w - but as mentioned by a sibling, 30-100% over a fully general-purpose gaming GPU is likely possible, because you don't have to have a whole general-purpose GPU sitting there idling (and it's not a coincidence that gaming GPUs were undervolted/etc to try and bring that power down - but you can't turn everything off). "ASIC-resistance" just means an ASIC is only 1-10x more efficient than a general-purpose device, so general-purpose hardware can still stay in the game. It doesn't mean ASIC-proof, you can still make ASICs and they still have at least some perf/w advantage.
However, if your ASIC costs $100 to get the same performance as a 3060 Ti, that's a huge win even if you only beat the perf/w by 50%. Particularly since your ASIC is likely way easier and more stable to deploy at scale, and doesn't require a host rig with at least a couple hundred bucks of computer gear to even turn on.
Only plebs were buying up GPUs from retailers or sniping websites, buying from ebay was for the chumpiest of chumps. Gangsters were buying them from the board partners a truckload at a time, true elites just pay someone to engineer an ASIC and do a small run of them. Eight-figures (as mentioned by a sibling) is plenty, a $50-75m run of ASICs is quite a lot of silicon even on a fairly modern node (and some mining companies were publicly known to be using TSMC 7nm and other very modern nodes). And when you invest that kind of money, you don't flash it around and scare the marks.