How to make LLMs go fast
vgel.me
vgel.me
but also, on a broader scale, if a transformer model is presented with a long input that does not fit in its context (e.g.: you are building a chatbot, and you have a very long chat history), it must "compress" or "forget" some of that information (e.g.: repeatedly summarizing historical messages, dropping them and appending the summary at the beginning of the input).
Mamba/RWKV/other "recurrent" architectures, can theoretically operate on unbounded input lengths; they "forget" information from earlier tokens over time, but is that not comparable to what a transformer must do with input lengths greater than their context window?
There is definitely a use for this kind of model, but this also shows why Transformers are still the main architecture we use today.
Beware of feel-good traps like this.
If you were to map the human connectome to a computational neural network down to the ion channel, it'd be at least 500 quadrillion parameters*. That's at least 5-6 orders of magnitude beyond what is currently possible with SOTA ML which means that even if the human brain was 99% devoted to compressing those tokens, that 1% that could actually do work with them is still a thousand times bigger than GPT4. There be emergent dragons.
* This is a fascile argument to begin with since biological neuron signals aren't quantized and ion channels are far too complex to map to a single static parameter
Perhaps, but they inherently have imprecision due the nature of being biological. Expose the same neuron to the same input N times, and you'll get a range of output. The effect of noisy analog data is, broadly speaking, similar to the effect of low-resolution digital data.
Newer methods are going back to similar concepts but trying to get past previous bottlenecks given what we've learned since then about transformers.
Your assumption that the same effort could be given to MAMBA or RWKV isn’t necessarily true. I’m sure there’s research arms looking to see if they scale to the level OpenAI or Google have their transformers at now, but they’re still very much in their infancy.
This all ignores the risk of them not panning out and the lost time not focusing some energy on transformers.
New technologies aren't generally just created out of whole cloth by some genius, they're build up of layers of incremental improvements, each of which is modest in its own right but which when taken together are groundbreaking.
[I would be happy to say more but I just got a top-level comment flag killed, presumably because they though I was advertising, so I won't mention the company name.]
Probably because the pace of development is now no longer determined by ML researchers but by CS engineers. ML is "easy to pick up" until you have to understand the reasons certain architectures work or don't for different applications.
OpenAI's moat is the high cost of training as you mentioned but this might be obsolete in a few papers down the road.
They themselves realized this by trying to turn it into a platform but I don't think it's enough.
I wonder if some sort of diffusion hybrid is being worked on.
Something that approximates a complete text answer, but 'increases the resolution' of that text over time.
One key benefit to a 'shot-gun/top-down/all-at-once' approach as opposed to a 'guess the next token => add and repeat' is the ability for the tokens at the end of the text (a twist in a surprise mystery story) have a direct effect on the beginning of that story. which (as far as I know) is not possible in current LLM's due to their architecture.
how self attention and positional encoding would work in a diffusion model... that's the question I have if that could even work...
to be clear, I don't mean stable diffusion rendering text as an image. I mean diffusion run on raw (random) text: turning that into a legible response.
One of the issues is that you'd need to be really sold in this being a better paradigm before spending the money to pretrain a huge non-AR model.
What are some models that 'work okay'?
The implication there is that consumer hardware will have more and faster RAM than the current defaults. AI features are going to probably require more than the 8gb minimum that macbook airs have. Sure mistral 7b can fit, but I think generally useful models start at 32B, and mixtral 8x7B still is short of GPT-4. Mixtral requires about 28GB 4bit quantised.
It will be interesting to see if llama etc drive a step up in consumer hardware defaults, but always connected to cloud and paying a subscription for cloud services is the trend with commercial offerings.
And 2 days ago, I was able to use MS' copilot (not Github's) to generate Python scripts, and now it tells me that it is incapable. Just tried it right now and the functionality came back though..
They just keep removing functionalities. To me that is pretty much censoring or whatever you want to call it. The quality degrades even faster then Google search quality.
here is a very simple example: https://i.imgur.com/vydRUwn.png (I would never ask that question usually, but I can't remember which questions I previously asked because they deleted my history)
Either that or some massive efficiency breakthrough, probably (hopefully!) both. Otherwise AI will be limited to enormous computing banks at datacenters, and it will be severely limited in terms of privacy (because no one wants to upload their personal data to datacenters, especially not to power some black-box "AI")
Our hardware is an "LPU" (language processing unit). The silicon has a very regular/uniform/homogeneous design with compute located close to memory, which allows us to get a much lower latency for language models than a GPU can. A lot of the smarts is in the compiler toolchain, which can fit high-level model descriptions (say Torch, ONNX) to the hardware without the need for handwritten kernels. I work on the assembler part of the toolchain, which is written in Haskell.
EDIT: To the downvoters, I guarantee that once you've tried the demo you'll be so surprised at how fast it is that you'll change your vote.