250$ is a lot of money to me though. I spent €290 on my 16GB GPU for my AI server but that was once off and I really had to think about it.
I'd love to see a llama model that fits now economically inside 16GB. The 8b is a bit too small when quantised even to 8 bits. A 16-20b model would be perfect.
But I think for 400b models to be viable, the hardware pricing really needs to catch up.