Eg when I have the AI do a self-code-review before bothering a human, I also want to the AI to give the draft PR to a sub-agent that only has the publicly available context that a viewer of the final PR would have; and not all the accumulated reasoning that lead to the writing of the code in the PR.
Opus directs and starts haiku/sonnet subagents. Much more efficient than opus reading it all.
Without this explicit direction, Opus burned through usage to create complicated regex parsers & 8 excel sheets.
And watch 10 hours of football on Sunday for our DraftKings bets.
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.