A week of using Codex more than Claude
allaboutcoding.ghinda.com
allaboutcoding.ghinda.com
I mostly do very obsessive, tightly scoped, carefully thought out small changes on a fairly boring stack, one interaction at a time, verifying functionality and code. I know what I am doing, but I also know what I don’t like doing (the same exact set of things I’ve already done a dozen times in my career)
I see what you did there
codex is good, both cli and desktop app, you get lots of usage on any plan. sol is good! and gets the job done, write or dictate a very long and thoughtful prompt, and leave sol xhigh or max fast working on it for an hour or so
omp is an amazing harness, any feature claude code or codex is adding has likely already been here for a couple months. good harness which im suggesting to all my developer friends, but for everyone else codex is the better option due to its simplicity and being the plug and play option
claude is decent, but not great. all models are somehow getting restrictive. you get basically unlimited opus on max plans, fable is good but slow and the random guardrails suck soo much which is why i havent used it once in weeks now.
gemini 3.7 is great for speed. everyone is sleeping on it, including even me
kimi k3 - great for frontend, one of the few models thats willing to commit crimes for you AND has the intelligence to have a chance at actually succeeding;
ds pro and flash are fast but not something id actually use for important things, unlike sol, fable and maybe 3.7 here and there
glm 5.3 i haven't tested yet
honorable mention to local models which are actually getting good now! 5090s will continue to get more and more expensive in the coming months. sadly.
theres way way more than claude in this world and its taking people surprisingly long to figure that out. maybe its for the best!
And it can communicate, unlike the gobbledygook that comes out of Claude.
Not yet. Don't give the guy ideas.
Their cache read costs are $0.50 per million, or 25% of the cost of uncached reads.
The industry standard is a 90% discount, so cache costs you 10% of uncached. So that means 5.6 Sol actually costs less per million cache reads - $0.40/million.
If you are doing a lot of agentic work where the vast bulk of your token consumption will be cached input reads, you won't get the expected cost savings from Grok.
I imagine this is the result of some problem in their serving infrastructure that I hope they will fix, because then the pricing will become actually strong. (The other possibility is that they bet on distracting people with good headline prices assuming they'd miss the bad cache pricing, but I'll give them the benefit of the doubt on that.)
Here's a quick review I just posted if anyone's interested:
Competition is good. Excluding a leading player in the market because you don’t like Elon Musk is…something.
I’m absolutely “hot money” when it comes to coding models. These things are commodities.
Only one of those capabilities can actually deliver kinetic solutions. Meanwhile big tech revenue is delivering ad solutions.
I assume they're referring to the recent discovery that Grok Build was uploading entire repositories to their servers in the background, include .env secrets that had been excluded
https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75f...
That incident has put Grok on the no-fly list for a lot of people and companies
Frankly, the people who keep bringing this up are mostly engaged in motivated reasoning. I don't trust any company, and any product where I have to send my code to a third party to make it work is a devil's bargain. I don't trust any of the major labs, but it is what it is.
The only way forward is local models, but we're not there yet.
His grandfather wrote his tracts to raise an alarm about what he called “mind control,” on the radio and television, where “an unconditional propaganda warfare is carried on against the White man.”
~ https://www.newyorker.com/news/daily-comment/the-world-accor...Man, politics are a hell of a drug. Guilt-by-association tu quoque logic is just fine when it's someone you don't like.
Literally every commodity product is in a crowded space, and easily substituted. If it's a good product (Grok Build objectively is one of the very best in the space), it's a good product, and it's self-defeating to avoid it because you hate a guy for political reasons.
Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.
I also think it's callous to brush away "political reasons" as though it's some trivial abstract thing. Or perhaps it comes from a place of nihilism?
I choose not to give my money to people I think are enormously evil. That's really all there is to it. I don't see why this is "nonsensical".
Oh stop. Other people believe different things than you. If you cannot see how using a coding agent is not "devoid of personal or social responsibility", then you really need to step away from the keyboard.
What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.
Based on your reply, I'm not actually sure if you actually believe this, or if it's only a form of motivated reasoning because you have some positive feelings about Elon or whatever, and that we wouldn't be having this argument if the original commenter was boycotting some other product for some reason you agreed with.
That isn't what I wrote. There are tons of valid reasons to avoid a product, other than the "direct value you're paying for". I don't pay for lots of products because I don't like the past corporate behavior, for example. I'm disinclined to use a particular AI lab's products because they seem to be on a mission to scare the crap out of everyone, and usher in an AI regulatory state. I don't support that, so I don't use the product.
What I said was that it's spitting in the wind to do what you're doing, because it's based on personal dislike of a single man. You don't like Musk, for political reasons, and because of that you've ruled out a product line.
To date, SpaceX has done nothing that bothers me, other than have a bug that they fixed immediately. So I use the product. Musk's political associations are irrelevant to me.
Anyway, you do you. Hopefully you now understand my "bizarre take".
In a capitalist society voting with your wallet is one of the few powers consumers have to change corporate behavior.
Why would I give that up?
Do you have ANY idea about the SpaceX corporate structure? Elon is basically SpaceX's Sun God and the other shareholders don't matter.
Plus SpaceX is incorporated in Texas where I'm fairly sure the legal system is arranged in such a way that it's supremely hard to contest anything in terms of corporate decisions.
As far as the average person cares, every SpaceX shareholder and employee is basically an Elon sharecropper and they matter less than Musk's toenails in terms of corporate decision making.
I've had great luck with the ds flash v4, paired with prime-agent for the harness--I like the results a lot. And you get to see thinking tokens.
I haven't liked the model as much in opencode.
Sol & luna have been great everywhere. sol plans, luna builds.
I also like prime-agent's way of handling sessions better than any other harness i've used. You can run multiple agents from one instance, although the scoping could be better.
But they can interact with past sessions, so preserving context isn't as important all the time. I just tell them to search for [thing] in another session.
It seems to have no problem with all the skills and things the other harnesses are using. I use superpowers and ponytail a lot.
It's my daily driver now. I like it better than opencode. But it doesn't ask permission. So I put it in a VM.
Codex has been an excellent workhorse - doesn't feel like I have to dance around the guardrails, doesn't lose _everything_ when it compacts, and doesn't litter the workspace with a million and one planning to plan files.
I used to rely on Fable for research when it was first out, today it doesn’t seem to be much better than Opus, and it uses up the quota exceptionally fast - 1h Fable in a single short session, and there’s little left for Opus to hit the 5h limit in a second session. With Opus I get about 3-5h of relaxed use with a couple subagents to save the context, but there’s usually quite some disagreement between the subagents and orchestrator - Claude does some model routing with default agents and picks Haiku and Sonnet for subtasks - only later to disagree with them and redo the work - and burn extra tokens. With Claude, it’s really either Opus or Fable if you want some quality.
That said, their marketing is exceptionally effective. Virtually all nontech folks consider only Claude.
As someone who's used Gemini 3.7 Flash (Google sub mostly for the storage) and DS4 Flash a lot (~6B tokens), I'd actually place DS4 Flash (even pre-0713) above Gemini 3.7 Flash. Gemini has a tendency to leave some things unimplemented; perhaps it's agy which frankly leaves a bit to be desired as a harness.
Although I will praise DS4 Flash any day, it no longer makes sense for me after the price increase (GPT 5.6 Luna is a much better price point) and I have completely migrated my high volume workflows to Muse Spark 1.2 Contributor (which I find to perform better than DS4 Flash 0713, happily).
Is this Gemini 3.7 Flash by any chance? Then - No. Not sleeping on it. It’s just not good.
I had a Python package build fail this week due to an unpinned dependency. Gave it to Gemini spent 5-7mins before I noticed it going off in some tangent. Reran with Claude Opus 4.8 - fixed in under a minute.
I know anecdata of one. But something like this has happened every time I test a new model from Google.
Im not convinced to pay $200 for Claude’s models.
With Claude, I have to intervene every 15-20 minutes, it’s non-autonomous and it’s incredibly unreliable at self-correction. GPT is strong at self-correction but it tends to drift away from the plan to self-correct in a loop very often - a lot of tokens and time burnt on aimless churn. Opus tends to push its uninformed opinions and fake retrieval, drifting every turn increasingly farther from the intended and approved design. Opus skims over specs and makes too many mistakes.
As for closed frontier models, I prefer the GPT models over Claude’s.
I’ve started relying more on Grok, GLM, Kimi and DeepSeek models for subagents - I’ve ended up with a factory and am seeking to reduce my reliance on the closed frontier models - they’re just not SoTA on their own for development anymore.
I run a lot of SlopCodeBench - https://github.com/michaelasper/benchmarks
Fable/Sol/GLM 5.3/Kimi are its league (in that order) Deepseek/Opus is solid Qwen 27B is the floor - there's no reason to use Sonnet/Terra/Haiku
For everyday activity - I don't think you need to be using Sol (xhigh) for everything - unless you're made of money - I've found using Luna from OpenAI to be more than enough - it'll outreach to Opus/Sol when it needs to
Haven't had access to Gemini 3.7 but we're getting it at work soon, will give it a go!
Codex CLI is pretty bare bones in a bad way (at least Pi is extensible). Claude code is vibeslopped to the extreme
To be precise, you need a while-loop, user input and bash.
It's about 50 lines of Python: https://minimal-agent.com/
I built my own agent based on this and use it every day.
It was "amazing" back when I first tried 4.6, but that's just my rose-coloured glasses speaking, I guess. I think I was one of the first few to call out Opus 5 for being hot garbage.
I hit my weekly limit on $200/mo Codex plan in about ~2 days. :/ I'm not doing anything custom/crazy/special. A lot of 5.6 Sol Ultra though, I'll give you that.
If you're not doing anything special there is no reason to use the Ultra mode.
Ultra mode is for applying the maximum amount of tokens to a problem without regard to conserving any quota.
I've found it consistently amazing at game dev. It sounds like something that would be difficult for an LLM to verify and iterate on properly, but it almost feels like having a mini-Carmack inside your computer once you try it out. You can throw it at broad, sweeping optimization passes, writing 5 different styles of eyesight sensor frameworks to see what works best in the game as it is, visual scripting integration problems/extensions, etc. with fantastic results.
Also found Ultra great for "get this local LLM working as fast as possible on this odd server setup with old GPUs and AMX support, writing custom kernels/modifications to llama.cpp/sglang/etc as you go while taking notes from relevant research papers and online posts"
I have the appropriate privacy settings set up but wondering how much I can trust each company with them.
But one thing I’ve noticed which I find a bit of a red flag: by default you only archive chats. If you go online, it says there’s a location in settings where can delete your archive. But it’s not that obvious where to find, and when I finally find some link, it was literally broken. It said it can’t find any archived chats, even though I archive them all the time.
Bit of a red flag for me. Both Anthropic and OpenAI claim that when you delete a chat, it’s gone after some retention period. They’re just words but if they’re secretly training on your traces and you delete your chats, then they would need to break two terms/conditions: ignoring your “train on my data” preference and ignoring your orders to delete chats. So it is an extra barrier.
But OpenAI, as far as I can tell, doesn’t let you delete your chats. Convenient then if they change their mind about training sometime in the future.
But these are not the same level of technical assurance you get from say, a zero data retention provider on OpenRouter.
Right now I am finding I have to tolerate substantial friction to use Hermes for personal stuff with a ZDR provider and ChatGPT and Codex for less personal stuff because the products and models are simply so much better.
Does this mean it can break your code faster now, or have they actually worked on making it good? Every single time I've given Gemini a chance (in older point versions) it would almost immediately break something and throw itself into a loop. I have not experienced it being useful for programming and almost never heard an account of somebody else doing so.
Remember those stories of LLMs catastrophically deleting entire repositories or databases? It was always Gemini.
I'm amazed that you'd trust Gemini over DeepSeek, which I've had very good experiences with after some tuning, though still on a relatively short leash.
Something with this combo works really well for Rust dev. The model doesn't really annoy me at all and I have not switched to Opus or SOL. And the monthly token bill is much lower...
jcode is a very very peculiar harness, but has some out-of-the-box thinking built in (by the devs, thinking ...)
Crush is also very well put together, and, IIRC, can do "mid turn" interruption, so can be driven from the outside.-
Great analogy for some reason. At fist I felt Codex Sol was a bit more cold. But now that I've worked with it for several weeks it has grown on me, even shown some personality. I appreciate that it is a bit more business-like, Fable is a bit too friendly sometimes when it ought to be focused on work. Codex can be a bit more nit-picky.
I agree with most of his other observations. I've already started to bin tasks based on which model I feel is best suited. In general, for well scoped and straight ahead tasks where banging out code is what I want I reach for Codex. For less specced tasks where I need a broader view and want the model to fill in more details I reach for Fable.
Both are great and they make a good team together.
This post needs an edit. Author is not comparing "Codex" and "Claude". They are comparing Codex TUI/CLI with (presumably) gpt-5.6-sol, against Claude Code TUI/CLI with (presumably) Claude-Opus-5.
Ctrl + f > [5.6, sol, sonnet, opus or fable] yields no results.
"Claude" is a product family, which includes Models, and Harnesses (and probably more). "Claude code" covers both the Claude Code TUI, and CC in the Claude desktop app.
"Codex" is the same, and could refer to the Codex TUI, or Codex in the ChatGPT (formerly codex) desktop app. (And well, historically, gpt-5.*-codex.)
Hearing "Yea Claude is great for coding" takes an hour off my life.
Something something "Honey why don't you finish up with your Nintendo and come to dinner?"
really feels like discussion spawns only off post title and as a second or third order effect, post content
Sol is for routine work, Opus for frontend/design, and Fable for more complex / ambiguous / architecture work. Fable works extremely well to drive Sol as a subagent.
Fable is the only one you can actually trust to not look at the code, but Sol is somehow still more pleasant to work with, especially in fast mode. Opus is the enemy, and it will make you insane if you talk to it for too long.
Curious what method you like for doing this? I've tried a few options and I haven't found one I'm happy with yet.
https://github.com/steipete/agent-scripts/blob/main/skills/c...
The important bit I found is to explicitly remind Claude that Sol 5.6 is a very smart and good model; otherwise, Claude performs its normal condescension towards any non-Claude model behavior and insists on reading all the diffs in full and testing all of Sol's work, negating any token savings.
Wow, I made exactly the opposite experience. Codex loves to make things as complicated as possible, even ignoring instructions and predefined skills. Claude behaves way more pragmatic. Maybe depends on the type of work one does, or even which programming languages/frameworks are used?
With Opus 5.0 being kinda crappy vs 4.8, I think Anthropic is in trouble.
It's expensive but it's doing in hours what no one's done in 2 decades.
I plan on releasing all of this at one point. It's crazy it hasn't been done in 20 years!
Sol medium is a great balance between speed and being thorough, but it’s quite expensive. Luna xhigh seems to compensate for slightly lower intelligence by thinking and reasoning for longer, so tasks can take more time to complete. But it’s crazy cheap.
I also have some custom evals using promptfoo to make sure I’m not introducing regressions when switching models. So far, Luna xhigh has been really, really good for the price.
Don’t sleep on it. Give Luna a try.
Makes me wonder if the current Luna prices are sustainable.
It likely is. Going by the performance of very competitive small models, Luna is likely pretty small (do they publish sizes?) to the point it might be runnable locally like Qwen 3.8 27B.
The specialized hardware cloud runners have can likely run a small model very cheaply.
A large part of what pushes developers toward these products appears to be the billing model. Pre-paying for tokens is some kind of ideological red line for a lot of developers. I think this is a strategic error. The subscription models have so many more perverse incentives baked in. Those paying $100/m+ for subscription access are almost certainly getting taken for a ride based upon my experience with prepaid tokens.
Yes. But, as your "so many more" implies, there are also perverse incentives in pay-per-token. And now, for the first time, the companies with the perverse incentives also happen to own the intelligence needed to, ad-nauseam, evade market and customer oversight. Potentially, this is a war where one side can inflict a thousand paper cuts in one second and the humans are on the other side. I think this is going to be an interesting test, a taste if you will, of what AGI means.
I don't understand what this means. Are they overpaying and getting less? Typically "taken for a ride" means, exactly "The seller got more out of the deal than usual sellers would".
Buying a burger for $3000 == "taken for a ride".
Paying $30 for all you can eat != "taken for a ride".
If I can get a few subscriptions for 100 - 200 EUR month and NOT have to pay 3000 - 9000 EUR (based on ccusage and some other stats) in tokens then it’s a no brainer for me to do that.
I don’t get why paying per token would be better if it’s economically disadvantageous.
The agents then go for several rounds criticising each other plans and implementations, catching big and small issues on each other’s work. The end result is not perfect, but it is a lot better than what I can get from relying on only one model.
> We expect that agents coordinating in the wild will act in higher variance ways than we see here, because they’ll have different backgrounds and therefore different contexts. They also, presumably, won’t all be Claudes.
Claude chat itself called it 'Claude-on-Claude action', which I found cute.
What a brave new world we're in, where this is necessary. Regardless, it's appreciated. Although, I have the feeling that those using an LLM to do most of their writing will be less likely to include such a disclaimer.
Claude's models in my experience do a better job of inferring my intent, or to say it does a better job of giving me the result I imagined in my mind. A recent example was a UI prototype I was building for a desktop application. I had asked GPT's 5.6 Sol to update the open document in the prototype to better reflect the context of the feature I was designing, and 5.6 Sol took it very literally and had just added some text to the currently open document, not what I had in mind. I tried again with Claude Opus 5 and it added a completely new tab with a complete new document that, although imperfect, much better matched my expectations.
You could say this was a prompting skill issue, but seeing how many people are prompting their AI I believe the labs are incentivized to continue to improve their ability to infer intent.
When it comes to the desktop applications though, I find Claude Desktop's output to be incredibly verbose and full of jargon. I feel like it hits me with an entire essay and the UI doesn't have enough typographic hierarchy to make it easy to scan. ChatGPT Desktop is much better in this regard, I feel the output is concise, clear, and gives me just enough info to feel in the loop without being overwhelmed. Even though I have the setting on for technical language, it feels more understandable than Claude. I also feel that ChatGPT's desktop app has a better design and much more polish.
I do not really like how bloated both applications have become though. This weird segmentation of Chat, Work, and Code all just seems like it's pushing a technical limitation onto the user. The other day I opened a document in ChatGPT and asked it to do something, then it told me it could only do it in work "mode", so it then created an entirely new conversation with a reference to the previous conversation. It wasn't a completely new area of the UI either, it just added a "Work" badge to the new conversation in the list. Feels a bit unnecessary, like couldn't you just keep it all within the same conversation?
Claude generates more code, tries to build things that are not planned or needed, repeats same mistakes over and over.
Chatgpt on the other hand just does enough, within a project would not repeat the same mistakes, tries to guess your workflow, so you don’t have to ask it to run the same.
If you need more control over your code/projects, if you know what you want to get done, use Codex.
If you have no idea what you’re doing and are happy to let LLM drive the thing, use Claude.
Codex/Chagpt is for the competent.
Chatgpt is a better planner.
Have a discussion in Chatgpt, have Claude to plan work chunks and have GLM to deliver, Codex to review and fix, delivers an overall a better version.
Problem - It’s just too much of a context switching.
Solution - I am thinking about a new product, a collaborative workspace where I can run this workflow.
A lot of people in the comments do have a software engineering background. People at different skill levels in different backgrounds are going to be using these tools in different ways, and that's going to heavily impact their experiences with these models.
Sure, there are differences between Fable and Sol. But I've even seen people on here saying that they're getting better mileage out of Qwen models they're self hosting.
I think the driver is just as important than the car, when it comes to this sort of stuff.
For iOS I used Gemini 3.7 flash to build almost everything with good success, but had to reach for Opus to fix a rather tricky audio engine problem (missing function annotation moved audio to the main actor).
Reached for codex after google banned accessing Gemini from OVH servers for some reason.
I’m sure there are edge cases but for your daily vibe coding or “make my printer work with the Tailscale instance running here” business I think they’re all fine.
Why is fewer comments a good thing?
It writes out stories describing what isn't there or what used to be there. It's usually not helpful, just noise. It also likes to write it in very verbose AI-styled prose.
Problem is today's LLMs don't have the long term memory that humans have, and so remembering the reason behind a given change/decision has to be preserved in some way if it's non-obvious. Hence why there is {AGENTS|CLAUDE}.md, the auto-memory system, and 1001 variants of memory implementations in the wild. All are trying to ensure that LLMs can have the context they need at the location and time they need it. And you want to block Claude from using a technique that it natively finds helpful.
Dude, just talk about the current state of the code!
That's your perspective. For Claude that's an extension of its thinking, which makes it work better. Just like the person who takes notes so they have references for later. Take it away and you're negatively impacting outcomes.
Claude very often litters code with comments about decisions that were made within a single session/pull request, its just noise.
You'll ask it to do something and it'll comment the code with an answer to what you asked it, rather than just explanatory comments to whoever comes after.
There's also a second issue that if the code is actually incorrect, the comment can nevertheless bolster the case for it.
Not to Claude – its own, old comments have helped me/it solve new issues on more than one occasion.
Useful for the LLM to know the "why", but not something a human would do, unless it's a very critical and confusing part of the code.
It felt like it was commenting on the diff sometimes instead of what the code was doing.
Fewer AI-generated comments is generally a good thing.
//add returns the sum of x and y
//per section 2.1 of addition-implementation-plan.md sum is designed as the seam for user addition interfaces.
//previously sum added numbers, now it adds numbers
def add(x, y):
return x + yIt's really time to move to OpenAI...
Digital ocean particularly looks promising as well.
I can really recommend the book Clean Code, here is a summary: https://gist.github.com/wojteklu/73c6914cc446146b8b533c0988c...
I think more than anything else, I don't get a headache conversing with Sol. That alone is enough reason for me to stick to Codex.
Experimenting with adding open source models to the mix to get more execution done while using Sol as the brain.
Claude Code seems more generous with its quota, which is why I use it as my main driver.
That said, Codex does seem more capable, terse, and faster. There are some tasks that Claude can't handle but Codex can. One example was a WinForms binding/project deserialization bug. Sorry, the code is a mess, so even I couldn't quite figure out which part was causing which problem.
I initially thought the bug would be difficult to reproduce in a unit-test setting. Claude could only narrow down the problem and tell me where to put a breakpoint. Codex, on the other hand, actually managed to create a reproducible unit test first, and then used that to fix the bug. That impressed me.
The only problem is the quota. Codex burns through it very, very quickly, even when I'm just using Terra 5.6 Medium. That's basically why Claude Code remains my main driver despite Codex seeming more capable.
I do agree claude looks for more things to do in your repo, whereas codex is more likely to do what its old and stop. Which is better is personal preference as far as I can tell.
I paid for Claude for only 2 months, each several months apart. Including a week of trying both Codex & Claude side by side on the same tasks with the same prompts.
I had been subscribed to ChatGPT/Codex at $20 for over a year and now I'm on my second $100 month. The whole experience is just so much better.
For just about every other harness it’s either alert fatigue answering permission asks all the time, or spending too long time hoping that you know the tools well enough to scope out a permissions file that actually works. Then there’s the “yolo in a VM” approach which also is a time eater and overkill.
Until someone solves “auto mode” with the other harnesses, I’m with Claude.
Its not an out of the box feature unlike Claude Code.
.codex/config.toml
---
approval_policy = "on-request"
approvals_reviewer = "auto_review"
[auto_review]
policy = """
Your prompt to the permission classifier here, e.g.,
Allow requests by default except requests involving...
Obtain user approval for denied request.
"""
I hated this so much. Both acli, which my agents have to rediscover how to use from --help every time, and the MCP, which I have to reauthenticate against frequently. I have replaced both with a "skill" that just describes where to find an Atlassian API key and which version of the API to use. Works perfectly every time.
Damn, my experience is the complete opposite of this. I have posted about it a few times, e.g. https://news.ycombinator.com/item?id=49348265
tl;dr I gave GPT 5.6 a small-medium sized ticket, which should have been several hundred lines plus tests. It ended up creating a 25,000+ line diff. Another GPT 5.6 Sol with fresh context looked at the worktree and said 98% of it should be thrown away. Claude thought the same, and suggested that several dozen compactions the model went through over several hours must have caused it to go adrift. I guess that's one consequence of having a relatively small context window.
I still use Sol quite a bit. I find that it's consistently the opposite of what the author describes: it's too relentless. It doesn't know when to stop. Opus is the opposite: it'll give up a bit too easily. If everything goes well that's not an issue, but often times it'll say things like "task is done, btw I couldn't do X Y Z" and X Y Z will be some important verification step that failed because another agent was using that resource or something.
At this point I trust GPT 5.6 mostly with surgical changes, or general codebase exploration tasks. It is a faster model, so it's easier to get small things done with it. For everything else I prefer Claude, despite its annoying tendencies.
The harness provides the model with its tools, context, and environment for execution.
However, I prefer not to have the model within that harness also bear the responsibility for remembering the process steps—like plan, implement, review, fix, and verify—deciding when to move from one to the next, and keeping track of the loop's state.
For tasks that need to be repeated, I’ve been moving that part into a reliable, deterministic runtime. Inside it, Claude, Codex etc simply take on interchangeable roles.
This setup makes changing models much simpler: the overall process remains consistent, and each role can be optimized independently.
I’ve been developing this approach as ctx.traits, if you're interested, here are the docs: https://ctx.company/traits/docs/quickstart/example/
Which Claude is this about? Sonnet, Opus, Fable? All totally different beasts.
One thing I don’t love about codex/sol is I find it tends to overengineer and be overly cautious.
I was using it to do create some scraping + data processing.
It went kind of crazy on the provenance, need at least 3 sources of consensus before promoting facts type bullshit.
defined a bunch of enums and gates.
I just wanted scrape some site data and put it into a SQLite dB. Like chill codex.
I feel like Claude is better at that.
I feel like codex/sol is better at well scoped hard technical problem.
Where it can sort of run this brute force analytical loop.
Like doing performance optimization or other search type problems. I think the math proofs are good examples of this.
It also doesn't have a clear idea of what the actual threat model is, and builds all kinds of extremely defensive systems to account for imagined hostile actors. I'm like "Dude, it's only our systems that are creating these SVGs, they're never going to be user supplied, so you don't need to write an entire validation and sanitation framework here."
It also seems to treat the desired initial state of something as a permanent invariant and designs elaborate tests to ensure that it remains that way. Then when you make one little change it has to go and update a ton of tests it created.
I've had to rip out a bunch of overengineered jank from several feature implementations, and in doing so I ended up having to create retrospective documents that warn against this kind of behavior that I'll have the model review whenever a plan begins to go sideways.
I wonder if it’s an artifact of OpenAI’s values or rl training approach.
Also, it prob does make it perform better just not more efficient.
Great for the OpenAI employee working on security scanning who doesn’t have to pay for their tokens.
Not so much for the dev building their web app who is trying maximize their subscription.
Like hiring an aerospace engineer to build you a shed.
I mean personally I'd just like to use OpenCode with all the providers, if their desktop app was a bit more polished and Anthropic wasn't so restrictive. There's also Paseo, but it has some issues with OpenCode sub-agent liveness checks (I've seen them hang, though the same happened with Kepler, might be a GLM 5.3 issue idk).
> It felt to me that Codex created a much simpler solution in terms of code architecture than Claude.
This feels odd, cause I've seen people say the exact opposite thing, that the GPT 5.x models seem to love overengineering etc.
> The output of the Codex agent harness is much more “technical” than the one from Claude. Claude feels more like your colleague in a Tuple session writing to you while Codex feels more like a version of Data from Star Trek.
This is very much preferable to me omg, maybe I should give OpenAI a look again.
The speed is the first big contrast; I have a routine multi-step skill that I run several of per week. Opus 5 was routinely taking 2 hours to do it, while older Claude models took around 20 mins; Codex restored that speed.
Second is legibility. Somebody wrote in one of the related discussions yesterday that Claude's current linguistic contortions could legitimately be considered damaging to mental health, which doesn't seem (too) hyperbolic to me. Codex (Sol) isn't perfect but it's much more direct. And so far I haven't seen it display much of an attitude, vs Opus's infuriating passive aggressive sulky know it all personality.
I slightly prefer Anthropic to OpenAI as a company, but I will vote with my wallet and discontinue my max subscription unless Anthropic does some serious damage control within the next week or two.
Also whole heartedly agree with the Opus 5 output being incomprehensible. It’s written in a way that I’d call spaghetti-tech-English. You can unravel it but it’s painful. Fables explanations is effortless and smooth. Opus 5 is user hostile.
Yes I have tried different settings already.