And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
The electricity prices are relevant because if you paid $0 for your H100 and didn't use it a single time, you could buy millions of tokens in inference just on the electricity it draws while idling. If you can't keep that thing saturated through the night, you're probably underwater overnight. Likewise, it's too small to run even the frontier open source models so you need to be able to live with worse models.
Max power matters because you aren't going to run many of those H100s before you blow breakers in most houses. Newer houses in the US are 15A service to non-kitchen breakers, so 1650W (that might be peak rather than continuous, not sure). If you're plugging that into an existing run, you could maybe run 2 before you start blowing breakers? You can't just plug 4 H100s into the wall in a normal house.
Maybe I'm wrong, though. I'd be curious, it'd be neat to run my own inference for something more than what'll run on a 3080.
A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
My home still has 100 amp service (as does almost the whole city, most of it is a historic district), but it was built in 1910 so that's not shocking.
> A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
Oh yeah, I was mostly talking about having to upgrade the service to the house from the pole. A new circuit with a single outlet isn't horribly expensive, but needing to upgrade service is pretty rough.
But I never assumed it was to be cost competitive at current token rates. I assumed it was for the same reasons people might use open source hardware. Freedom to tinker, etc.
Models are too big now