But what is it doing for you that you couldn’t do yourself at that speed? I‘m really curious and on the fence of partly going local.
An other reason would be because your company does not allow any source code leaks and thus every developer either has local models or none.
6 T/second local model while your colleagues have 200 EUR/month claude does not make much sense. At least I can't see a use case.
Incase it's not clear, you will be generating 10,000,000 a second. Good luck verifying it. Token generation is not the bottleneck for creative work. If you are doing a predictable work and have a good workflow and massive dataset to process, then speed of token matters. If you are performing creative work like coding, it doesn't.