AI will be used to select the net to automatically load. Nets will be cached, branch predicted etc.
The future of AI software and hardware doesn't yet support the scale we need for this type of generalized AI processor (think CPU but call it an AIPU.)
And no, GPUs aren't an AIPU, we can't even fit whole some of the largest models on these things without running them in pieces. They don't have a higher level language yet, like C, which would compile down to more specific actions after optimizations are borne (not PTX/LLVM/Cuda/OpenCL.)
In 5 years, the specialized model would still beat the generalized ones. They would just be different from the ones today.
Also, it's not only the prompt engineering. Training ChatCPT was a decades long process, with billions of dollars invested, and it takes a small cluster to run the thing. How would the other models compare if given similar resources?
Besides that, on the context of this thread, the bitter lesson has absolutely no relation to the specialized vs. general purpose model dichotomy. It's about the underlining algorithms, that for this article are all very similar. (And it's also not a known truth, and looking less and less as an absolute truth as deep learning advances, so take it with a huge grain of salt.)