What does the depreciation curve look like for nvidia cards purchased today? How long will it take to recoup the investment on this buildout? Will those datacenters pay for themselves before they're scrapped?
What does the depreciation curve look like for nvidia cards purchased today? How long will it take to recoup the investment on this buildout? Will those datacenters pay for themselves before they're scrapped?
Let's say we serve a Fable class model on 8x B300.
From Kimi K3 metrics, with 8x concurrent streams, we would achieve 55-60 tok/s per stream, matching Fable 5.1 throughput.
432 tok/s x 3600 => 1.555M output tokens/h x 50$/M API price = $77.76 revenue per hour.
Assuming total API billing at 2.06x output token bill = $160.2 / hour or ~$20 per B300.
A server with 8x B300 could be $461.5k.
At an obviously unrealistic 100% utilization we would look at 4 months of revenue to match the cost of the server.
About how model serving works at scale and actual utilization I know little.
And for all we know Anthropic could serve their model with 64 streams on the same hardware instead of the 8 we assumed here.