Nvidia might do for desktop AI what it did for desktop gaming
theangle.com
theangle.com
Meanwhile all the usual "desktop" players are still trying to find a way to make good on their promises to develop their own competitive chips for AI inference and training workloads in the cloud.
I'm betting on Nvidia to continue to outperform them. The talent, culture and capabilities gap just feels insurmountable for the next decade at least barring major fumbles from Nvidia.
Right now using a two RTX 2080 setup for a pilot project, where it runs ollama and qwen2.5-coder (14B quantized version), a more serious step up from there on a budget would realistically be an RX 7900 XTX (24 GB, though I've had issues with setting up ROCm) for 1k EUR or RTX A5000 (24 GB) for 2.5k EUR in the current market.
Honestly, I don't even need the best performance for what I'm trying to do, even an Arc A770 (16 GB) for 350 EUR would be enough to iterate, except that it's not actually supported by ollama and lots of stuff out there: https://github.com/ollama/ollama/blob/main/docs/gpu.md (I know ollama isn't the only solution, but it sucks when the tools you like aren't available)
Underneath everything, ollama can do a lot of the heavy lifting, but it still needs to run somewhere. For decent chat models you probably want a server with either one beefy GPU or a few regular consumer ones (it actually seems to split the load just fine, at least when it comes to Nvidia hardware), whereas for smaller models like autocomplete (where 3B or 1.5B models are enough) you can choose whether to use the same server or run locally.
Though the local market here is a bit bad. There's a Mac Mini with an M2 Pro and 16 GB for 1.9k EUR, so more expensive than just two GPUs. I'm guessing that the local sellers are trying to profit quite a bit, because on Apple's site, the M4 16 GB version starts at 600 USD (and for 24 GB it's 800 USD and for 32 GB it's 1000 USD). That actually makes it a good option!
But they don't because that would cannibalize the extremely overpriced enterprise offerings. The #1 reason people are forced into the tens of thousands of dollar cards is memory needs.
So considering that, ask what niche this device really fills: Is it a new "supercomputer" for the home? Not really, given that it is silicon and memory bandwidth restricted so much that their $600 GPU can beat it soundly on every metric (not surprising when you look at the power and airflow/cooling needs of "real" GPUs) but in scenarios requiring large memory. But while this can hoist those larger models, it is going to be far removed from state of the art.
It's neat, but the market for this is being grossly overstated on a lot of these hype advertorials. The large models you'll run on it will be quantized the point of absurdity, not to mention that for 99.9%+ of users, anything short of state of the art is basically useless.
It's a neat eGPU of sorts for a Mac or something (they really hype the fact that you use CUDA for this). Still really trying to figure out what value it possibly brings outside of trying to lure a bunch of enthusiasts to blow money on this so they can fiddle with Llama for a week and then realize it's a waste of time.
Outputs at 8k resolution! (+)
...at 5 fps (+)
For comparison H100 can read its memory 40 times per second, so if you use it all you can get around 40t/s.
Of course in either case you don't have to fill it up, but instead use smaller model or more GPUs.
AMD Strix Halo has a 128GB config with 256GB/sec.
Apple has a MBP m4 max with 128GB ram for $4,700 with a 546GB/sec interface.
Hopefully Nvidia can do better.
Which is way slower than GDDR6x or GDDR7 let alone HBM. I don't expect these machines to be anywhere near as fast as the hype.
256-bit LPDDR5X is impressive, don't get me wrong. But it's impressive for a CPU platform. It's actually pretty bad for a GPU.
> Huang also revealed ‘Project Digits,’ a new product based on its Grace Blackwell AI-specific architecture that aims to offer at-home AI processing capable of running 200 billion-parameter models locally for a projected retail cost of around $3,000.
> There are many exciting things about Project Digits, including the fact that two can be paired to offer 405 billion-parameter model support for ‘just’ $6,000
My experience with running local LLMs is quite limited, but most tools can split the workload between GPUs (or more commonly GPU+CPU) with minimal fuss. It parallelizes fairly well. There may not be any actual secret sauce beyond just having the necessary gobs and gobs of fast memory to load the model into.
AI training feels like transport. You rent the capacity/vehicle you need on demand, benefit from yearly upgrades. Very few people are doing so much training that they need a local powerhouse, upgraded every year or so.
Even sharing the hardware in a pool seems more rational. Pay 200/month for access to a semi private cluster rather than having it sit on your desk.
I see similar coming for AI. Tagging local photos, reading/summarizing private documents, helping you code (without uploaded it to 3rd parties), using uncensored LLMs, maybe even playing NPCs once games support API for LLMs.
It's going to take awhile, but I wouldn't be surprised if NASs start supporting AI accelerators to provide local endpoints for AI.
The comfy ecosystem is rife with people that want local tools.
If you replace your GPU every 2 years, its 250 usd per month
If the price halves, still its 125 usd.
Even if price halves and use for 5 years, its 50 usd per month.
> I spent $745 in Kling credits to bring the Princess Mononoke trailer to life.
https://www.reddit.com/r/aivideo/comments/1fvchbf/i_spent_74...
Most of that is in failed generations.
Many startups have been trying for several years now, and eventually someone will succeed, but it's not easy. Even AMD hasn't been able to pull it off.
sending that over the network feels very idk icky because it's not just photos or emails
[1]: https://developer.apple.com/machine-learning/core-ml/
[2]: https://machinelearning.apple.com/research/neural-engine-tra...
[3]: https://research.google/blog/improved-on-device-ml-on-pixel-...
Seems like this is MUCH more likely to have at least decent support, I think the current DGOS is based on ubuntu 22.04 LTS.
> Nvidia will be introducing two new chips, the N1X at the end of this year and the N1 in 2026. Nvidia is expected to ship 3 million N1X chips in Q4 this year and 13 million vanilla N1 units next year. Nvidia will be partnering with MediaTek to build these chips, with MediaTek receiving $2 billion in revenue.. Nvidia will show off its upcoming ARM-based SoCs in Computex in May.
we're already at that stage now with AI / LLMs. this type of physical product will remain niche.
obviously companies will try to migrate everything to the SaaS model because it is better for revenue and bug fixes but not necessarily the best user experience
don't want games to stop if I dont't have wifi and i don't want ML stuff to require internet or feel comfortable sending my entire digital data to a third party or have models be updated because of some compliance policy
but it's still in the infancy companies will do everything to make it over the network one or two nice tech people will try to make it local/hybrid hopefully they succeed
It will take more than jeans and leather jackets to sell those
They are still selling in 2025 an Android 11 device whose last update is from 2022.
Sounds about right: Nvidia gives a shit about maintaining the devices they sell.
Further, there was a beta test update just last month.