With Nvidia's GB10 Superchip, I'm Running Serious AI Models in My Living Room
pcmag.com
pcmag.com
Memory bandwidth is 273 Gb/s. Which is nowhere near a GPU’s. It’s a 4K machine. Personally, I’d rather have two GPUs and run a quantize model. I have two 32GB AMD r9700 cards, cost $2600. Quantized models get me 120K ish of context window and TPS is about 60% of what I see with the same model on my 4090 (which only has enough vram to load weights and about 6K context).
Sure I can’t run a 100B+ model but neither can a single GB10 unless no context window is what you are going for. So you buy a second 4K machine?
At least this thing is actually useful, and there are $3k variants available.
I'm building and using this machine daily for building and using applications with LLMs, TTS, STT, ASR, and image generation.
Before anyone gives me grief my company has a strategic partnership with Nvidia, I do AMD under the cover of darkness. So I live in both worlds. I’m a bleeding heart for the under dog…if being a 360B market cap company makes you the “under dog”.
I don't know what you want me to tell you, you're welcome to believe whatever you want but that doesn't change the reality I experience actually using the thing.
Benchmark numbers and first hand reviews are readily available if you bothered to look.
For example: For LLMs, it's easy to do the math, and see how long you will be waiting for 50k input tokens.
https://github.com/kyuz0/amd-strix-halo-toolboxes
It takes all the work out of it, you just start llama-server in the container context and you're off doing inference without having to figure out dependencies.