I think the next models will be benchmaxxing on the Pelican benchmark tbh
11 karma · joined July 28, 2026
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents