TL;DR: It's a very interesting line of thought that as late as Q2 2024, there were a couple thought leaders who pushed the idea we'd have, like 16 specialized local models.
I could see that in the very long term, but as it stands, it works the way you intuited: 2 turkeys don't make an eagle, i.e. there's some critical size where its speaking coherently, and its at least an OOM bigger than it needs to be in order to be interesting for products
fwiw RAG for me in this case is:
- user asks q.
- llm generates search queries.
- search api returns urls.
- web view downloads urls.
- app turns html to text.
- local embedding model turns text into chunks.
- app decides, based on "character" limit configured by user, how many chunks to send.
- LLM gets all the chunks, instructions + original question, and answers.
It's incredibly interesting how many models fail this simple test, there's been multiple Google releases in the last year that just couldn't handle it.
- Some of it is basic too small to be coherent, bigcos don't make that mistake though.
- There's another critical threshold where the model doesn't wander off doing the traditional LLM task of completing rather than answering. What I mean is, throwing in 6 pages worth of retrieved webpages will cause some models to just start rambling like its writing more web pages, i.e. they're not able to "identify the context" of the web page snippets, and they ignore the instructions.