I think it’s sort of self-defined. If a model is able to use bash well enough to not need specific tools.
The research seems to agree with you, though. The paper calls out that for “bash capable” models, adding tools to do things bash can already do doesn’t improve performance.
Vaguely the same result as RAG. Unless you’re in specific domains, you won’t beat handing the agent a shell and grep.