I feel like tool calling killed RAG, however you have less control over how the retrieved data is injected in the context.
Heck, the RAG Agent could run cosign diff on your vector db in addition to grep, FTS queries, KB api calls, whatever, to do wide recall (candidate generation) then rerank (relevance prioritization) all the results.
You are probably correct that for most use cases search tool calling makes more practical sense than embeddings similarity search to power RAG.
or maybe even "cosine similarity"
RAG looks linear (constant per lookup) while tools look polynomial. And tools will possibly fill up the limited LLM context too.