I would suspect what you have instead is a single model attached to data sources, where the model doesn’t have to have so much compressed fact, and instead can rely on higher level summary.
I ended up with the premise that each of them had their relative strengths and weaknesses, and it would actually be best to use all of them, but only in their own areas of strength. Then have them use something akin to a shared blackboard where they could all read the results from the other systems and write their results as well. That the sum of all the available algorithms working together, each in the areas in which they were best, would result in better outcomes than any one algorithm could achieve on its own.
My professor was not impressed. I only got a C.
Now, the story of how my dad had to work his ass off to finally get the College of Engineering to force the professor to actually give me a grade that he owed me, when all the professor really wanted to do was focus on his new job at one of the big airlines -- well, that's a story for another time.
And part of the reason why single-hidden-layer networks aren't enough even in continuous memoryless Euclidean cases is, again, because of how loss functions work; you're unlikely to converge on a good approximation with very few hidden layers.
TLDR, the essay explores how LLM could evolve into front-end routers that connect users with specialized tools, leading to a future where federated models determine the best-suited system to answer specific queries. Not too different from today's federated search approaches.