I don't see more drafts as a real differentiator.
Well the person I replied to did..
Google has invested in custom AI hardware for some time now and does not run their workloads on nvidia cards
To do so, you need to split the matrix multiplies across the new machines. You also need more inter-machine network bandwidth, but with GPT-3 that works out to 48 kilobytes per token predicted collected from every processing node and given to every processing node. Even if Bard is 100x as big, that is still very doable within datacenter scale networking.
However, OpenAI doesn't seem to have done this - I suspect an individual request is simply routed to one of n machine clusters. As they scale up, they are just increasing n, which doesn't give any latency benefit for individual requests.
They are claiming to be the first to achieve >50% saturation during training. Pretty sure I recall Midjourney is using TPUv4 pods too
https://cloud.google.com/blog/products/ai-machine-learning/g...
https://cloud.google.com/tpu/docs/system-architecture-tpu-vm