One thing you can do in pi that you can't in Cursor: Have a 5-prompt conversation, jump back to prompt 3 and have a new conversation [call this convo2], then jump back to the original point 5, then jump back to convo2.
So fork is really only needed for when you need to interact with the conversation tree in two separate processes.
1) system prompt in pi is quite small (way smaller than the one from OpenCode)
2) when your agents.md file changes pi does not re-spam it (preserves cache, good trade-off!)
3) only 4 tools, every tool comes with a description for how to use it and causes reasoning overhead (fewer tools is good)
all of these things add up
here are pi, opencode and smol working on the same tasks in 9 fresh runs
https://smolenv.com/t/nested-template-includes-60636/
you can step through the traces and see how the system prompt + tools steer the agent in a certain way
with GPT 5.6 Sol you can even get away without a system prompt (see smol) and only 1 tool (sh)
Basically, it doesn't handle the context "better", it barely does anything special to it, which can actually be better for cost efficiency.
So unless it’s been fixed or someone knows a work around, Pi is DOA - I’ve found that on a MULTI tool call (ie one prompt firing off multiple tool calls until it prompts you again) that’s close to hitting the auto-compaction limit (default compactor or extension) it will either keep going until your context spills over and you OOM, or it interrupts itself to compact but then loses the context.
From reading issue after issue on GitHub, I think it’s because Pi doesn’t let extension writers (nor the built-in compactor) hook in between each tool call and so the only place to check if it can compact is when it finishes a request and is about to wait for the next prompt - too late by then
it also have soft/soft compaction limit, it tries to compact on turn boundary when possible. with combining with above this can get you about 35% more context (at least it looks like this with the sol)
codex when shell command is executed, will pull output with hard cap at max 30s, so for running compilation it will burn tokens without any benefit.
I have some tasks where agent will have to run some suite that can take over an hour, and codex burns about $20/h just waiting and reasoning every 30s "yep, that's still running". And what is going to happen after compaction, when whole context was just waiting? it will loose the plot and when I'm back it just does completely different thing that I asked it to do.
codex also have a bug, that opening refuses to resolve that adds your last steer after compaction, so imagine that you asked it to cleanup some tmp files or refactor/simplify something. it will do that again and again after each compaction, best case it just burns tokens and figures out, this is already done, or worse do it again and mess up everything and forget about it's task