I suspect that some/a large portion of the build out is due to OpenAI/Anthropic/etc. just throwing hardware at the problem -- why bother trying to make training and inference efficient when you can just throw hardware/money at the problem. I remember NVIDIA talking about their supercomputer clusters with a large number of interconnected GPUs.
I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:
1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.
2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.
[1] Though the original idea for mixture of experts comes from a 1991 research paper (https://huggingface.co/blog/moe), so maybe a third area is research from Universities, etc.