I've had more complicated set-ups than that, but they still just looks like extra moving parts.
Instead, I have a (pi.dev) tool called `do` which forks the agent and spawns a task on top of it - then returns the last message to the parent agent. (or `delegate` to start with a fresh context).
That's it.
claude/codex use it for a lot of things (impl/review/search/misc), and thus including situations RAG would otherwise be used. By skipping RAG, these search sub-agents instead use tools they've been trained on to use (find, grep, etc), and can decide to look deeper if required.
I'm not saying RAG couldn't be made to work, but with my set-up i've automatically gotten improvement every time a new range of models was released.