This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?
1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?
How is 2 better in any way for anyone involved? Why would you, as a paying customer, holding it wrong, want other people to keep that information from you?
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
https://community.openai.com/t/why-does-codex-report-a-258-4...
I barely compact at work in a very complex monorepo (neither with Fable 5.1 nor Opus 5.5), and yet in my personal greenfield project Astra keeps compacting all the time, to the point of it being unusable.
yes, exactly
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
~/.codex/config.toml
model = "gpt-6.1-sol"
model_context_window = 700000
model_auto_compact_token_limit = 630000I’ve also found compaction not to be a problem when it does happen.
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I've always seen them be prompted, but this is something high priority on my list of unknowns to understand better...
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.
Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
However it's less willing to obey your instruction so it's less usable for general runtine flows.
Looks like 5.5 is the new 4.6
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
Going on a slight tangent, I find that I get the best results when I force Codex models (Astra/Sol) and Claude models (Opus/Fable) to consult each other (just have them build a simple skill). There are tasks that neither one can fully solve on their own, but their differences are large enough to make a difference when they collaborate.
My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now, but either way having 2 "smart" models available from 2 different companies is a boon.
Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.