Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
I think the models are being optimized for wealth extraction from users and companies, instead of solving problems.
I don't know why Opus would try to create an entire library when I told it specifically to do something simple that would take 2-3 lines of Python.
Because it reasons in one direction. First it encounters some kind of issue with 2-3 lines of Python that might make it not work, and then it goes onto plan B, which is making a library, but it doesn't circle back and compare the effort of making the library to working around whatever might make the 2-3 lines not work. Except sometimes it does, because it's inscrutable.
Yeah, that’s my thoughts as well. I feel it’s great for benchmarks and some tasks while in other it tries to spend as much tokens as possible, tries to overcomplicate task and needs seconds or third round of steering that costs. With the scale Anthropic operates I bet it’s huge amount of extra money just to make sure their model works.
It seems to me ANTHROP\C have harnessed hard for not spending tokens to bring content into context. I wish they'd left us a "LEROY_JENKINS" flag: read and think before you code. In Claude Code anyway, it appears to default to:
export CLAUDE_CODE_LEROY_JENKINS=trueYES! They introduced the new tokenizer to increase token generation by upto 33%.
On top of this, Anthropic are generating almost twice as much revenue per paid user than openai - whilst their subscriptions have lower usage limits than openai's:
I don't think so. Expect that in a market with high vendor lock-in but that's not the case here. The market is extremely competitive and switching cost are near zero. Anthropic can't afford to pull shit like this and sacrifice quality.
Just checked the dashboard, and we seem to have the exact same $200 credit as others, enterprise or not. Token inflation affects us just like everyone else.
It feels a bit like buying the same box of chocolates every day, but the size / weight of the box is shrinking... the price remains unchanged!
Plus there's subjective stuff even for coding, people learning how to deal with it. Even on HN you can already see cloude/codex camps each strongly convinced that one is better than the other.
And yet, the Java language exists.
For market share, ANTHROP\C needs to optimize for the vast mediocrity that are mid-bell-curve users and enterprises.
This adaptation tends to come with significant drag for the right ends of the bell curve firms or teams.
Unless ANTHROP\C have a separate objective function by cohort and ensure that doesn't regress, improving results for the emerging middle will nerf tools from point of view of those with high in-domain expertise.
And this type of observation appears highly related to how people hold it.
So this API can be used in a UserPromptSubmit hook [2] in the harness, get the token count for any model, calculate the cost and compare.
[1] https://platform.claude.com/docs/en/build-with-claude/token-...
You have to test each task obviously but it is not a bad model on its face.
https://www.reddit.com/r/ClaudeAI/comments/1ukgqwr/looks_lik...
The explanation Anthropic gave for the update doesn't address how the x-axis needed to range up to $50 previously and only $10 now. In any case the pass rates are also lower.
Probably the difference between whatever it is people notice when they say models become "nerfed".
> Anthropic did post an official explanation, stating the original chart used a "simpler methodology" that "underestimated Sonnet 5's performance." The new chart supposedly uses their "standard methodology."
Oops!