I feel like this is the greatest demand for LLMs at the moment too.
It's hard to believe we're only 8 months into this industry, so I imagine we'll start seeing smaller footprints soon.
It's hard to believe we're only 8 months into this industry, so I imagine we'll start seeing smaller footprints soon.
Gpt3 is 36 months old. Dalle-e is 28 months old. Even StableDiffusion is like 11 months old.
TBH devices just need more ram for coherent output though. Llama 13b and 33b are so much "smarter" and more coherent than 7B with 3 bit quant.