8k diffs PR is incredible, is it even sane for a human to comprehend that much?
There has to be a name for such scenario were the handoff from Agent to Developer is so difficult, its not even worth pursuing but to instead wait until the Agent comes back online.
No matter how good of a developer you are, we are still constrained to our minds working memory. Lots of people I have spoken to are aware of this. Agents produce lots of diffs in a fast pace, our minds don't get the time to grasp the changed logic. And then, we an Agent is unavailable all of the sudden, say in the case of it being an outage or usage limit reached, the human now have to go through massive amounts of changes. They cannot continue as usual, there is a spike, they have to learn the new changes and grasp everything before they can continue towards the goal.
Opus isn’t even allowed to decrypt thinking blocks from Fable if you switch models mid-conversation.
https://platform.claude.com/docs/en/build-with-claude/preser...
Other than trivial PRs; everything I do with Opus/Fable gets reviewed by Astra; and everything I do with Astra gets reviewed by Opus/Fable.
Using only models from a single vendor, is like testing your website/webapp only on Chrome.
Could you share a bit more about how you do it?
After claude or codex finishes a commit, I switch console tabs and ask the other to review it; with the commit ID. I often find it helpful to inject a bit of human knowledge, and callout any areas of attention I see from a quick skim. (But that could just be me wanting to not abstract myself away from software engineering that much :)
Sometimes, I do ask Codex to read my ~/.claude/; and vice-versa. But, generally, I try to keep as much knowledge (e.g. investigations, reports, deep dives) inside the git tree as possible; so that is not necessary.
I don't use skills, but I do have ~/.codex/AGENTS.md and ~/.claude/CLAUDE.md. These are high-level instructions for what (1) I consider readable, maintainable code and patterns, and (2) workarounds for empirically observed model behavior IOdon't like, such as Astra being a bit of a "over-correct over-validation nit-picker". I keep these human-authored, and update regularly based on what I find annoying.
I do all of this before I submit a PR; but of course, for trivial stuff (e.g. CSS changes, copy/string changes, etc), I don't bother.
My practices and workflows do change over time. Back in the ~Opus 4.5 days I'd often define a rubric/criteria in a markdown file, iterate with AI to improve it, and that's the "spec". I've stopped doing that since GPT 5.6; partly because models have gotten a lot better at understanding high level intent from the context; and partly because nearly all models these days feel 'gradermaxxed' when working like that.
Finally, consumer $100/$200mo subs get you _so_ far, I get a lot of value from both. I used to have multiple Claude subs for a while, but trying to 'get full value' made me work on projects just for the sake of it; so 2x$200/mo is my cap :)
It's just feelings.
PRs shouldn't really be reaching 8k diffs no matter who authors them.
There's an ocean between what "should" happen and what actually happens in tech today.
(but yes if AI is going to propagate, they need to fix their uptimes)
Not sure how long the latter will remain true – centralization is well underway there; try not to think of Cloudflare too much...
We already got used to "internal going down globally" when AWS was the main cloud provider. Now its 3, AWS, GCP, Cloudflare, but "portion of internet that you care about is down for the entire world" is a casual thing to happen now.
As somebody in need of a doctor, on a plane etc. I'd definitely consider it much worse if all hospitals, air traffic controllers etc. stopped working at the same time instead of just a local subset.
Isn't that a good news!
Likewise while at Google, they didn't let us put Google Voice numbers down as our oncall phone number. Some folks even had multiple phones on different networks as backups.
I'm sure they have either other SOTA models or open source models around to help if things to completely pear shaped.