This is true, but there's also the cases where slight differences in prompt yield wildly different results. In any programming language, if I add a clause to a conditional like "if car is red or car is blue", that behaves predictably--and if it doesn't we can dig into the debugger, assembly, etc. If I do that with an LLM, that can change everything, and there's no way to "debug" it.
This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.
So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.