To be comfortable you need 300 to 400 terabytes of storage. You need the GPUs and the cooling for them. They don't need a lot more. You put models in memory and you do math really fast and make tokens. You need some CPU and RAM infrastructure for like prompt caches and things like that, which does add up, but the dominant cost is by far the GPUs and VRAM. Everything else is a pretty easy step down to afford. You probably need to have multiple different models, right? Like if you want an Astra plus a Sol plus, you know, like a Luna-class model, say, you probably need like 10 terabytes for Astra, probably 3 to 4 terabytes for Sol, and 1 to 2 for Luna, maybe less, maybe one or half a billion even for Luna. That's RAM. You need So all total you need 15 to 16 terabytes of active VRAM to keep those models resident. That's not counting what you want to keep in cash and whatever hosting constraints these models have that aren't public. That much VRAM is five to eight million dollars right now.
Plus what models are you running? The biggest models you can run are like in the terabyte range or Qwen or GLM. Kimi K3 is there, but GLM 5.3 is very competitive. Using voice dictation, hands hurt, sorry for any errors. You still want many terabytes of RAM to run say multiple GLM+DeepSeek V4.1 for some actual concurrency.