HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by orangepanda | Hacker News Reader
Parent
Full thread
orangepanda
·
Or maybe even middle class plebeian 24gb rigs?
View on HN
griomnib
·
At that point just run 8b.
pulse7
·
Or wait for the IQ2_M quantization of 70b which you can run very fast on 24GB VRAM with context size of 4096...
griomnib
·
At some point there’s so much degradation with quantizing I think 8b is going to be better for many tasks.
Reply on news.ycombinator.com