Can you host the model for a lower cost per token than you'd pay Anthropic or OpenAI for a similar level of intelligence? I doubt you're beating their efficiencies of scale.
Ok you can host this model once. What if I want a dozen subagents? Ok you can host it 12 times at once. What if we go a whole week only using max 4 at a time? Etc etc. The limits imposed by self-hosting might be bearable for a variety of reasons, but it's going to be more expensive and less convenient/useful.
The marginal cost goes down significantly if you have datacenters. Baring in mind that US API pricing is kind of absurd. Even if you, say, only utilize your DC 1/10 of the time... you might still be ahead of API pricing by a wide margin.