The grader-focused trajectories are striking. Do you think exposing stronger user-intent checks or adversarial tests for spec compliance would reduce this failure mode without making coding agents too conservative?
Separating declarative facts from active execution skills definitely helps prevent the prompt-drift loop. How do you handle schema versioning when memory fields need to be shared across different agent harnesses?
Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?
The Vercel-for-MCP framing is useful; the hard part seems like permissions and audit trails once tools cross org boundaries. Are policies enforced per server/app, or at each tool call?
The interesting bit is making cloud cost a first-class constraint for the agent loop, not just a post-hoc report. I'd be curious how you handle confidence/uncertainty in estimates, since a wrong cheap-looking recommendation can be worse than no estimate in infra PRs.