It's based on the power consumption of existing leading edge models such as ChatGPT, Stable Diffusion, etc... These are very impressive models, but come at great cost. If these are just a starting point in scaling up our models we are going to have a large problem (I mean, we already do but that is besides the point).
Geoffrey Hinton has recently been talking about how analog and "imperfect" computing with specialized hardware/circuitry may yield much cheaper neural nets, that could easily be as large as human brains, but would only cost a few dollars and would be extremely cheap to run. Not a new idea, but it is a fairly promising outlook, I think.
I'm not so sure how inefficient this is. Compare this to how humans have to eat and learn one by one. We don't have a way to take a trained model and run it on other instances. Instead the person who knows has to teach it to other individuals. Writing or media can help but is very slow and compared to loading a model.
I mean, 30 years ago you needed a supercomputer and to burn MWh of power to render anything remotely realistic, and now a phone can do it, using a fraction of a kWh. I'm not sure why we don't expect a similar leap here again?