IMO we are currently in the ENIAC era of LLMs. Perhaps there will be a brief moment where things get worse, but long term the cost of these things will go way down.
Cumulative AI capex will hit $2T this year. Cumulative opex is on the same order. Unless the models get real good (as in: can fully replace many engineers) right quick, nobody is even going to see interest getting paid on those investments. The only alternative is model access costing 5 figures per (replaced) seat.
But yes, once GPU racks can be had at auction for pennies on the dollar, inference of open source models might be an... OK low margin commodity business.
A huge difference is early computers were not subsidized. It took decades until most people could afford to own a computer at home.