1. Im not sure who you believe will produce these chips, or who will use them.
You are correct that specialized inference chips will get you a 10x gain.
So what?
If you want several million of them in a datacenter, that's a tall order.
That's on top of
A. the hundreds of millions it will take to get a design to production.
B. The complete and total lack of allocation at anybody who could make you chips, except at very very very high cost, if at all - have you forgotten that automakers still can't get cheap chips made on older processes?
Most allocation of newer processes is bought out for years.
While there is some ability to get things made at newer process, building a chip for 7nm is 10-100x as expensive as say 45nm.
C. The fact that someone has to be willing to build, plan, and execute putting them in datacenters.
This will all happen, but like, everyone just assumes the hard part is the chip inefficiency.
We already can make designs that are ~10x more efficient at inference (though it depends on if you mean power or speed or what). The fact that there are not millions in datacenters should tell you something about the difficulty and economics of accomplishing this.
People aren't sitting around twiddling their thumbs. If Microsoft or Google or anyone could make themselves a "100x better cloud for AI", they would do it.
2. Dead. Dennard scaling went out the window years ago. Other scaling has mostly followed. The move to specialization you see and push to higher frequencies is because its all dead.
3. Will take a while.
The economics of this kind of thing sucks.
It will not change instantly, there isn't the capability to make it happen.