Nvidia is halving the production of its higher end GPUs, doubling prices and heavily segmenting the market. They don't give a single shit about individuals. One wafer that makes 10 5090s that might sell at 3000 each, or one wafer that it already sold 6 months ago for 200k to one of the incestuous AI companies it works with?
Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
A lot of llama.cpp contributions come from the community and ecosystem, like Unsloth. If something goes awry, I fully expect lots of forks.
There already are a lot of forks for things they decline to implement. TurboQuant, ROCmFPX, and more. I need to set up an agent that will loop on merging them.
Except when they have less than 16 gb of ram?
It'd be really nice if I could use my egpu 4090 on my macbook pro..