Now disabling adaptive thinking plus increasing effort seem to be what has gotten me back to baseline performance but “our internal evals look good“ is not good enough right now for what many others have corroborated seeing
Now disabling adaptive thinking plus increasing effort seem to be what has gotten me back to baseline performance but “our internal evals look good“ is not good enough right now for what many others have corroborated seeing
> Claude Opus 4.7 (claude-opus-4-7), adaptive thinking is the only supported thinking mode. Thinking is off unless you explicitly set thinking: {type: "adaptive"} in your request; manual thinking: {type: "enabled"} is rejected with a 400 error.
https://platform.claude.com/docs/en/build-with-claude/adapti...
For my claude code I went with following config:
* /effort xhigh (in the terminal cli) - To avoid lazying
* "env": {"CLAUDE_CODE_DISABLE_1M_CONTEXT": "1"} (settings.json) - It seems like opus is just worse with larger context
* "display": "summarized" (settings.json) - To bring back summaries.
* "showThinkingSummaries": true (settings.json) - Should show extended thinking summaries in interactive sessions
Freaking wizardry.
Particularly when compared to Opus 4.6, which seems to veer into the dumb zone heavily around the 200k mark.
It could have just been a one-off, but I was overall pleased with the result.
I think i’m doing it wrong
I straight up skip all the memory thing provided by harnesses or plugins. Most of my thread is just plan, execute, close - Each naturally produce a file - either a plan to execute, a execution log, a post-work walkthrough, and is also useful as memory and future reference.
Is your CLAUDE.md barren?
Try moving memory files into the project:
(In your project's .claude/settings.local.json)
{ ...
"plansDirectory": "./plans/wip",
"autoMemoryDirectory": "/Users/foo/project/.claude/memory"
}
(Memory path has to be absolute)I did this because memory (and plans) should show up in git status so that they are more visible, but then I noticed the agent started reading/setting them more.
Is it... not aware of its current directory? Is its current directory not the root of your repo? Have you maybe disabled all tool use? I don't even know how I could get it to do what you're describing.
Maybe spend more time in /plan mode, so it uses tools and the Explore sub-agent to see what the current state of things is?
- Use the Plan mode, create a thorough plan, then hand it off to the next agent for execution.
- Start encapsulating these common actions into Skills (they can live globally, or in the project, per skill, as needed). Skills are basically like scripts for LLMs - package repeatable behavior into single commands.
`claude --thinking-display summarized`
The thinking is then visible with ctrl+o in the claude cli (shortcut available at least on mac).
Well you can't really trust the documentation I guess. I can't edit my original comment anymore.
But they made their own bed with that one.
As for distillation... sampling from the temp 1 distribution makes it easier.
For example, chat, cowork and code have no overlap - projects created in one of the modes are not available in another and can't be shared.
As another example, using Claude with one of their hosted environments has a nice integration with GitHub on the desktop, but some of it also requires 'gh' to be installed and authenticated, and you don't have that available without configuring a workaround and sharing a PAT. It doesn't use the GH connector for everything. Switch to remote-control (ideal on Windows/WSL) or local and that deep integration is gone and you're back to prompting the model to commit and push and the UI isn't integrated the same.
Cowork will absolutely blow through your quota for one task but chat and code will give you much more breathing room.
Projects in Code are based on repos whereas in Chat and Cowork they are stateful entities. You can't attach a repo to a cowork project or attach external knowledge to a code project (and maybe you want that because creating a design doc or doing research isn't a programming task or whatever)
Use Claude Code on the CLI and you can't provide inline comments on a plan. There is a technical limitation there I suppose.
The desktop app is very nice and evolving but it's not a single coherent offering even within the same mode of operation. And I think that's something that is easy to do if you're getting AI to build shit in a silo.
That's clearly a trade-off that Anthropic have accepted but it makes for a disappointing UX. Which is a shame because Claude Desktop could easily become a hands-off IDE if it nailed things down better.
Need to fall back to codex to keep things in sync, but that's a great opportunity to also make sure I can compare how things run - and it catches a lot of issues with Claude Code and is great at fixing small/medium issues.
Whatever their internal evals say about adaptive thinking, they're measuring the wrong thing.
It didn’t give me a line number or file. I had to go investigate. Finally found what it was talking about.
It was wrong. It took me about 20 minutes start to finish.
Turned it off and will not be turning it back on.
But I don’t know, man in my opinion you don’t fucking snicker about a malloc without a null check and only a conditional free that isn’t there.
Go to hell “Sprocket”.
I did try out codex before claude went to shit and it was good, even uniquely good in some ways, but wasnt good enough to choose it over claude. Absolutely when claude was bad again it would have been better, but thats hindsight that I should have moved over temporarily.
Currently we are all subsidied by investors money.
How long you can have a business that is only losing money. At some point prices will level up and this will be the end of this escapade.
edit: example: GLM 5.1, a 751B model, is offered for 0.6$/m in, 4.43$/m out. Scuttlebutt (ie. I asked Google's AI) seems to think that Opus 4 is a 1T/5T MoE model, so you can treat it (with some effort) as a 1T model for pricing purposes. Its API pricing is $1.55 in, $25 out, ie. 2x to 5x more than GLM. Idk what to say other than this sounds about right, probably with healthy margin.
It was terrible. You could upload 30 pages of financial documents and it would decide "yeah this doesn't require reasoning." They improved it a lot but it still makes mistakes constantly.
I assume something similar is happening in this case.
I faced the same issue using Open Router's intelligent routing mechanism. It was terrible, but it had a tendency to prefer the most expensive model. So 98% of all queries ended up being the most expensive model, even for simple queries.
With a small bounded compute budget, you're going to sometimes make mistakes with your router/thinking switch. Same with speculative decoding, branch predictors etc.
as long as you introduce plans you introduce a push to optimize for cost vs quality. that is what burnt cursor before CC and Codex. They now will be too. Then one day everything will be remote in OAI and Anthropic server. and there won't be a way to tell what is happening behind. Claude Code is already at this level. Showing stuff like "Improvising..." while hiding COT and adding a bunch of features as quick as they can.
If you vibecode CRUD APIs and react/shadcn UIs then I understand it might look amazing.
With the fully-loaded cost of even an entry-level 1st year developer over $100k, coding agents are still a good value if they increase that entry-level dev's net usable output by 10%. Even at >$500/mo it's still cheaper than the health care contribution for that employee. And, as of today, even coding-AI-skeptics agree SoTA coding agents can deliver at least 10% greater productivity on average for an entry-level developer (after some adaptation). If we're talking about Jeff Dean/Sanjay Ghemawat-level coders, then opinions vary wildly.
Even if coding agents didn't burn astronomical amounts of scarce compute, it was always clear the leading companies would stop incinerating capital buying market share and start pushing costs up to capture the majority of the value being delivered. As a recently retired guy, vibe-coding was a fun casual hobby for a few months but now that the VC-funded party is winding down, I'll just move on to the next hobby on the stack. As the costs-to-actual-value double and then double again, it'll be interesting to see how many of the $25/mo and free-tier usage converts to >$2500/yr long-term customers. I suspect some CFO's spreadsheets are over-optimistic regarding conversion/retention ARPU as price-to-value escalates.
A company providing a black box offering is telling you very clearly not to place too much trust in them because it's harder to nail them down when they shift the implementation from under one's feet. It's one of my biggest gripes about frontier models: you have no verifiable way to know how the models you're using change from day to day because they very intentionally do not want you to know that. The black box is a feature for them.
there's no contract. you send a bunch of text in (context etc) and it gives you some freeform text out.
And Claude have no idea why it did that.
I misread that as Atrophic. I hope that doesn't catch on...