If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.