That another 400 to 700k.
It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.
That another 400 to 700k.
It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.
40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.
A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.
Who is going to fix it when the api does something weird ?
Who is going to proactively make sure it’s not overheating?
Chat GPT has enterprise contracts for a reason.
> A company of the size that this is worthwhile for, probably has dedicated devops on staff already
I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.
The companies which have the will and the budget to host their own LLMs typically require a ton of other, much smaller, models as well. They have internal security requirements, guardrails, audit, critical workflows start depending on your onprem setup, downtime is now something that’s not even allowed. There’s going to be a zoo of tooling, lots of bespoke work with internal clients who have no clue about docker, but now want their vibe-coded app to access a model they downloaded yesterday, and this model better be served and monitored 24/7 because now C-levels use it.
No one is going to budget several millions to then look at an email admin who hears ‘cuda’ for the first time in their life and ask them to just support the entire thing somehow.
And god forbid it’s an AMD setup.
This obviously isn’t relevant for a 10-person startup and their second-hand xeon with a single H100.
From my experience only the companies which are REALLY interested in privacy and data security bother with hosting their own models
I agree that what you've described is probably worth a dedicated person. The earlier comment describing one rack running one model, is not.
(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)
Why not? Do you trust AWS with things that need to be private?
Not a devops but I'd say one full time is already too many.
Add to this number another $1.5m/yr in opex, so not sure I’d call such an enterprise wealthy enough to spend those kinds of sums on LLMs an “SMB”.
Having these models in the open caps the inference margin.
Adding another rack beside the VMWare cluster, managing any storage/networking issues, etc. will be incremental costs; they already have a pager (probably not a pager not anymore just an app on their phone) like rotation schedule etc.
I'd update this to
'I don’t LLMs for anything that needs to be private'
What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.
A NN per se is a file... The executable that runs it can "act"...
Or that perhaps you have a perfect alignment algorithm which you are unwilling to share with the broader research community (evil)?
If you're doing the training yourself, you at least have a verifiable supply chain and an audit trail. Today, we have no idea if a black-box model handed to us is coded to recognize specific domains or patterns and back door an application in a sneakily targeted way.
Black box is a black box. Many open-weight models clearly haven't been trained or created in the manner their creators claim---which raises the obvious question: if they lied about the recipe, what else did they lie about? Putting those models in a production capacity scares the living crap out of me.
That said, building from scratch isn't about achieving mathematical perfection or manually auditing 15 trillion tokens---that's impossible. It's about eliminating third-party supply chain risk and having actual governance over the pipeline.
Of course, that doesn't mean we can magically guarantee gradient descent won't produce weird emergent behaviors, or that we can blindly trust OpenAI not to backdoor things. But at least with the latter, you're making a calculated operational decision rather than blindly trusting an opaque black box of entirely unknown provenance.