I believe embedding-based RAG, everybody is using, will end. As chips advance, you would use a big llm instead of word embedding for retrieval. It's much more accurate and extensive covering every topic.
Still need ~2 years to be replaced.
Still need ~2 years to be replaced.
With agents, the prompting could be dynamic for maximum accuracy for every retrieval.
This absolutely would beat the best of the best embedding-based RAG models.
Nobody uses this now mainly due to speed. An llm retrieval would be 10x or more slower than embedding.
You can try that now
Take some failing cases or bad retrieval from your current system Prompt an llm wisely like a perfect prompt to get what you want and provide it the context to it. And see the results.
For context, you are limited now by models contexts (1m), so mostly you would need to split what you have and prompt twice....or more...and so on