I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:
1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.
2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.
[1] Though the original idea for mixture of experts comes from a 1991 research paper (https://huggingface.co/blog/moe), so maybe a third area is research from Universities, etc.