Fine-tunes lead to catastrophic forgetting.
RAG is only irrelevant if you’re completely disinterested in cost and latency.
We also don’t have enough data to gauge performance of models >200k context window size when reasoning over inputs of that size, much of which will be irrelevant to any particular user. Multiple random needles in haystack tests work flawlessly, but rarely applies to real world activity.
Just last Friday I took the contents of the 2024 folder of one of the teams at the company I work for, for which we use RAG at the moment. I dumped the text index, concatenated it and used Google’s API to return the token count, to see if it would fit in Gemini’s 1M context window; turned out it was 5.7M tokens. And that’s less than 3 months worth of documents for that team.
So yeah RAG is not dead yet, although I do question its usefulness, but that’s a separate topic.
The claim that RAG is dead is obviously wrong.
The internet is full of blog posts about this. That doesn't mean they're actually good - I'd love to be pointed at one that has proven itself useful for someone (and definitely isn't just LLM blog-spam).
I don't care if it's trivial to fine-tune and get crap results - I care about fine-tuning where the result was worth the effort.
For the record, my favourite guide to fine-tuning is the section of this Jeremy Howard video that shows how to train a text-to-SQL model: https://www.youtube.com/watch?v=jkrNMKz9pWU&t=4850s
The person asked for citations, leave it be, stop dudexplaining how Internet works for you please.
I call bs on your 2 comments.
Even when people sincerely ask for a citation on a debatable topic, on an internet forum, it's effectively saying "I won't be hear any opinion that doesn't match my own unless it's as water tight as a law of physics". Another form of this is "show me the data".
Disclaimer: I work for Google.