While my colleagues are running 6 parallel agents at 50-100t/s each, with an actual SOTA model? Don’t you think I‘d get fired after a few weeks of that?
An other reason would be because your company does not allow any source code leaks and thus every developer either has local models or none.
6 T/second local model while your colleagues have 200 EUR/month claude does not make much sense. At least I can't see a use case.
Incase it's not clear, you will be generating 10,000,000 a second. Good luck verifying it. Token generation is not the bottleneck for creative work. If you are doing a predictable work and have a good workflow and massive dataset to process, then speed of token matters. If you are performing creative work like coding, it doesn't.