I'm not convinced, because I think there will be a drive to distill down models and constrain them, and try to train models with access to "premade blocks" of functionality we know should help.
E.g. we know human voices can be produced well with formant synthesis because we know how the human vocal tract is shaped. So you can "give" a model a formant synth, and try to train smaller models outputting to it.
I think there's going to be a whole lot of research possibilities in placing constraints and training smaller models, and even training ensembles of models constrained in how they're interacting and their relative sizes to try to "force" extraction of functionality.
E.g. we have reasonable estimates at the lowest bitrate raw audio that produces passable voice. Now consider training two models A, B, where A => B => audio, and the "channel" between A and B is constrained to a small fraction of the bitrate that'd let A do all the work, and where the size of B is set at a level you've first struggled to get passable TTS output from.
Try to squeeze the bitrate and/or the size of B down and see if you can get something to emerge where analysing what happens in B is doable.