Case in point; I'm mostly a Claude user, which has decent background process / BashOutput support to get a long-running process's stdout.
I was using codex just now, and its processes support is ass.
So I asked it, give me 5 options using cli tools to implement process support. After 3 min back and forth, I got this: https://github.com/offline-ant/shellagent-tools/blob/main/ba...
Add single line in AGENTS.md.
> the `background` tool allows running programs in the background. Calling `background` outputs the help.
Now I can go "background ./server; try thing. investigate" and it has access to the stdout.
Stop pre-trashing your context with MCPs people.
If you do like rolling your own MCP servers, I've had great success recently refactoring a bunch of my own to consume fewer tokens by, instead of creating many different tools, consolidating the tools and passing through different arguments.
e.g. as I am dropped into a new codebase:
1. Ask Claude to find the section of code that controls X
2. Take a look manually
3. Ask it to explain the chain of events
4. Ask it to implement change Y, in order to modify X to do behavior we want
5. Ask it about any implementation details you don't understand, or want clarification on -- it usually self-edits well.
6. You can ask it to add comments, tests, etc., at this point, and it should run tests to confirm everything works as expected.
7. Manually step through tests, then code, to sanity check (it can easily have errors in both).
8. Review its diff to satisfaction.
9. Ask it to review its own diff as if it was a senior engineer.
This is the method I've been using, as I onboard onto week 1 in a new codebase. If the codebase is massive, and READMEs are weak, AI copilot tools can cut down overall PR time by 2-3x.
I imagine overall performance dips after developer familiarity increases. From my observation, it's especially great for automating code-finding and logic tracing, which often involves a bunch of context-switching and open windows--human developers often struggle with this more than LLMs. Also great for creating scaffolding/project structure. Overall weak at debugging complex issues, less-documented public API logic, often has junior level failures.
I use Sonnet 4.5 exclusively for this right now. Codex is great if you have some kind of high-context tricky logic to think through. If Sonnet 4.5 gets stuck I like to have it write a prompt for Codex. But Codex is not a good daily driver.
It's ridiculous and ties into the overall state of the world tbh. Pretty much given up hoping that we'll become an enlightened species.
So let's enjoy our stupid MCP and stupid disposable plastic because I don't see any way that we aren't gonna cook ourselves to extinction on this planet. :)