Where could I get a mapping of token / time vs hardware?
The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU cores (e.g. Tesla P40s also have 24GB of VRAM, for half the price or less, but are much slower than 3090s), or run on the CPU entirely, which I don't think will be viable for this model even with gobs of RAM and CPU cores, just due to its sheer size.
I would be curious to see relative failure rates over time of consumer vs Quadro cards as well.
I don't have any actual figures to back this up, but my gut tells me that the fact that enterprise GPUs are an order of magnitude (at least) more expensive than, say a, 3090, means that the payback period of them has got to be pretty long. I also wonder whether setting the max power on a 3090 to a lower than default value (as I suggest in my other post) has a significant effect on the average W/token.
Not necessarily saying that Quadros are cheaper, just that there's more to the calculation when trying to run 405B size models at home
I think it should work as-is with the components listed, but if you disagree please let me know!
I'm not necessarily saying that it's obviously better in terms of total cost, just that there are more factors to consider in a system of this size.
If inference is the only thing that is important to someone building this system, then used 3090s in x8 or even x4 bifurcation is probably the way to go. Things become more complicated if you want to add the ability to train/do other ML stuff, as you will really want to try to hit PCIE 4.0 x16 on every single card.
Will need more space, true.
Here are some TGI 405B benchmarks that I did with the different quantized models:
https://x.com/danieldekok/status/1815814357298577718
The 405B model is very useful outside direct use in inference though. E.g. for generating synthetic data for training smaller model: