> [..] , and on a site like this one that matters more than architectural purity. [..] Nothing in it depends on a build step [..] That distinction is worth keeping in mind [..] What matters is [..]
1,098 karma · joined May 1, 2014
GitHub: https://github.com/StanAngeloff | Mastodon: https://mastodon.social/@stanangeloff | LinkedIn: https://www.linkedin.com/in/stanangeloff/ | X: https://x.com/StanAngeloff | Blog: https://blog.angeloff.name
> [..] , and on a site like this one that matters more than architectural purity. [..] Nothing in it depends on a build step [..] That distinction is worth keeping in mind [..] What matters is [..]
This is the full URL that does a composite average across DeepSWE, Terminal-Bench and SWE-Atlas-QnA. Models are measured in their respective harnesses.
What is surprising to me is that Claude Code + Fable 5 (max) is on par with Codex + GPT-5.5 (xhigh)... yet Fable burnt through 1M extra tokens.
> ### Rewriting Bun with dynamic workflows
> An example of what dynamic workflows can unlock at scale is the recent rewrite of Bun. Jarred Sumner used dynamic workflows to port Bun from Zig to Rust [..]
That's very interesting to hear!
> I'm working with Claude Code on session aaaaaaaa-bbbb-1223-3445-abcdefabcdef which I'd like to hand-off to you, do you know how to read the session, my input and Claude's output so we can resume where I left off?
gpt-5.5, medium effort. "Resumed" session fully in under 2 minutes. Outages like today's are so common that I've now got the time to re-evaluate Codex every other day.
Discussion: https://news.ycombinator.com/item?id=47753643
This was a first for me with Sonnet. It completely veered off the prompt it was given (review a design document) and instead come out with a verbose suggestion to do a mechanical search and replace to use this newly fabricated function name - that it event spelled incorrectly. I had to Google numey to make sure Sonnet wasn't outsmarting me.
I wonder if this is even more exaggerated now through Easter, as everyone’s got a bit extra time to sit down and <play> with Claude. That might be pushing capacity over the limit - I just don’t know enough about how Antropic provision and manage capacity to know if that could be a factor. However quality has gotten really bad over the holiday.
Another thing that worked like magic prior to Feb/Mar was how likely Claude was to load a skill whenever it deduced that a skill might be useful. I personally use [superpowers][1] a lot, and I've noticed that I have to be very explicit when I want a specific skill to be used - to the point that I have to reference the skill by name.
I was wondering if anyone else is also experiencing this? I have personally found that I have to add more and more CLAUDE.md guide rails, and my CLAUDE.md files have been exploding since around mid-March, to the point where I actually started looking for information online and for other people collaborating my personal observations.
This GH issue report sounds very plausible, but as with anything AI-generated (the issue itself appears to be largely AI assisted) it’s kind of hard to know for sure if it is accurate or completely made up. _Correlation does not imply causation_ and all that. Speaking personally, findings match my own circumstances where I’ve seen noticeable degradation in Opus outputs and thinking.
EDIT: The Claude Code Opus 4.6 Performance Tracker[1] is reporting Nominal.
EDIT: and inevitably 10 minutes into the Opus outage the dreaded 529 Overloaded on Sonnet appears.
Feels like half the internet is down.
https://www.bbc.com/news/live/cpvrrre4zlkt?post=asset%3A0240...