It’s mind bogglingly crazy that language models rivaling ones that used to require huge GPUs with a ton of VRAM to run now run on my upper-mid-range laptop from 4 years ago. At usable speed. Crazy.
I didn’t expect capable language models to be practical/possible to run loyally, much less on hardware I already have.