Sonnet 5.5
anthropic.com
anthropic.com
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
You’re right about people not watching plumbing videos and doing it themselves. But the equivalent example would be open-source software in the tech example. But instead of reading open-source code to see how different features were implemented, AI can go and dig into the code and figure it out.
The fact that you think this is the way software is priced is telling.
The model being destroyed here is that every piece of software is something that needs to generate recurring revenue.
I was more pointing out that software gets enshittified. A plumber necessarily doesn’t and if the plumber does get worse, you can call a different one next time. Software, especially B2B, has switching costs and lock in. So you just have to put up with it.
Another point is that most software started with a few features and to get more market share and support more use cases, it became worse for the users using the early features. That’s why they try to build their own so it’s not bloated with features you will never use.
You couldn't move the goal post further from "Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons".
So this one falls under denial for me.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
One of the quotes which might help as well (I think I have this even in my HN profile): The only thing we know about the future is that it will surprise us.
Not even experts are much more likely to predict for what its worth than a coin toss in many cases (especially if they believe that only one theory/idea will mostly predict the future)
It's a blend of things and ideas and the sheer interconnectedness of them where a small pocket can grow large and then also shrink and taking into account all variables and factors is just simply impossible for a mind. I think that although we feel we are being more informed about the world, that in it of itself doesn't prevent things in the future from happening. It just makes us alert and sad and anxious about it.
Yet this life is one which shouldn't be lived with sorrow and anxiety. It is one of beauty and greatness. In many ways, we humanity have come so far from the past (Our medicine is something that not even the mightiest of kings could get) and yes, there are many problems in the world and some things feel as if they are staying just the same or getting worse real-time.
But even then, worrying about it could lead to nowhere other than a path of misery. Also these problems are complicated enough that its extremely hard for a single person to bring change (not that I wish to demotivate that person but rather seeing the system as a complex nature)
So to me, its also a form of inward action. I can work on myself to be better prepared for the world that comes next. In the same time, I think that the present for me as well is good. I have great family and friends and have many qualities that I am proud of and I wish to share that gratitude to the people who have helped me along the way (my family/friends/ Hackernews!.)
Within the hustle culture, there is no time to relax but it is within the time of relax that I believe some of the most fruitful actions can come. I believe it just makes my mind more productive being in a calmer state.
here's a quote from how to measure your life that I hope can help some people:
I genuinely believe that relationships with family and friends are one of the greatest sources of happiness in life. It sounds simple but like any important investment, it needs constant attention and care(...)
You'll be tempted to invest your resources elsewhere but if you don't nurture these relationships, they won't be there to support you in hardships or as one of the most important sources of happiness in your life.
So thank you hackernews and have a nice day and please, please try to say gratitude towards someone close to you (within these tough times) and try to keep a balance towards inward focus, sharing time with friends/family and also writing on hackernews (as is my past time nowadays), balance is necessary :-D
So once again, I hope that its a call to action to say gratitude towards anyone. Just send them a big message thanking them and make their day as well as yours memorable, have a nice day!
Also you are allowed to make mistakes (everyone makes them!) and even though I am saying (preaching?) these things, I have found myself sometimes failing to act on these things as well but I just think that these help in being more mindful about them hopefully and can help provide a perspective. I wish to adopt more of these things in my life myself as well hopefully :-D
[Pardon me for the long post]
but the AI thing is on one side using lots of energy to keep up with it, and on the other side gives you an edge, because most people are not aware what even is already possible. so suddenly you are the "AI-Expert" just because you try to keep up to date. so if it all goes to shit, at least we have a chance to sniff it in the wind a couple moments beforehand. Or make memes from it.. that helped me cope with it: https://t.me/RobotComrades
I hope you find a peer group to talk with and exchange and build community. it is so rewarding to talk to likeminded people that have a similar knowledge base and soothe some fears that someone might have, and have them help you with the ones i have... (i recently did a deep dive in custom DNA synthesis, and how connected those services are already to API https://www.twistbioscience.com/tapi )
as always accepting what is seems to be a healthy strategy
Not sure how atheists are coping with the existential risks we are facing.
You either keep in mind all the horrifying little possibilities the future could hold for us and try to prepare, accept and cope, or you push it out of your mind using whatever techniques available to you to not drive yourself into nervous spirals.
Ultimately it comes down to what your brain chemistry allows in combination with ways you practiced dealing with stress, existential dread, cognitive dissonance, etc.
I guess submission to a higher power is one way to deal with it? That way it's no longer your problem (alone).
Even some very basic questions, like “are humans inherently valuable?” have been thought through and discussed thoroughly in many faith traditions. For many people not part of such a community it’s a question that’s suddenly very important and they lack the tools to address it.
The proof of the pudding.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding.
Yes. Our career is over, as is our economy. Soooo... FYI :(> the economy is over
Hackernews' neuroticism remains undefeated
It depresses them. Significantly. White collar jobs constitute the bulk of global purchasing power. What happens to the economy when aggregate purchasing power drops? The naive response is "prices fall until equilibrium is reached again"
But what if the needle continues moving so quickly that equilibrium is never reached?
This is the K-shaped-economy concern. The ultra wealthy and those who own the "AI means of production" will become unfathomably wealthy at the expense of everyone else.
Why should this not be a concern? Historically, this trend has always precipitated bloody conflict.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A lot of old-school software engineering is about how to deal with this reality.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
> literally hand-write code they don’t understand even as they write it
I find this literally impossible. How can you even start typing anything without knowing what to type?
Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?
People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.
It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.
Etc.
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
So accelerate the vibe coding of shit nobody wants or asked for, just to see some metric go up somewhere.
Are we still getting bonuses for the number of tokens we can burn?
Frontier models today don't really write incorrect code at the micro level. They do miss edge cases at the high level though, and that's what we want to test, is the scenarios.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
Just because robots can do stuff doesn't mean the human power structures or service preferences evaporate.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
We also have open weight models too, and ways to host those at home.
Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.
Where's the super awesome 100x turbocharged software that's a result of everyone here having been being a 100x turbocharged programmer for the last 6 months and a 10x supercharged programmer the past 2 years?
I still use the same software I used 2 years ago, but a bit less reliable.
I'm not angry, but yes I'm being deeply sarcastic to illustrate the extreme case you seem to be advocating, where we relinquish our understanding to the machines.
In a competitive business like software you need an unfair advantage and for very few people that's having near unlimited tokens. Even then I doubt that's going to produce good software.
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
I do agree that they're not great at program design by default and that's where we as engineers should spend our time. Data structures and data flow are king. But once you suss that out, they're pretty good at writing the resulting code.
This is also where I disagree with dhh about just using lower level languages. Good abstractions make for excellent program understanding and we should continue to build extremely good building blocks that make program design naturally solid.
For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.
Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.
It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.
Continue this improvement process long enough and you may find yourself with an AI Software Factory.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
You sure that wasn't just working at Microsoft?
And then on top of that you need another agent to manage merging all of the subagent code together?
Not complex at all, only one extra session other than the ones doing work and it's on a dumb model and can be thrown away & restarted because it only dispatches work, not doing anything.
I do everything in there, collecting requirements, kick off research, branching, merging, not one other agent on top. I considered making that orchestration command llm-powered but it's not justified at my current use.
It's not more expensive, in fact I could have just chugged along with the slow and manual session by session work but I have a claude subscription and another GLM one (the most low cost basic tier, not even much), that just sit there collecting dust if I don't put them to use in a more efficient way.
And doing session by session would face your problem when context switching too much become unscalable.
For a concrete example, check out this random plan [0]. A detailed spec followed by the exact implementation tasks that will be executed by the subagents.
[0]: https://github.com/bensyverson/woodcase/blob/main/project/20...
It does involve letting go and not micromanaging every code convention and implementation detail, but that is the same skill you need when leading engineering teams.
Last year I held off on implementing a few features knowing that a model like Opus 5.5 was around the corner. I'm now implementing them in a much more efficient and quality manner than I could have fall of 2025.
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
(With exceptions for what I can only call the "manic vibecoders" with like 10 simultaneous weird slopprojects they're spewing out at once. Generally with each project itself being something related to vibecoding. Steve Yegge being an example of a "manic vibecoder-actual programmer" hybrid.)
Also, I’d imagine the token-maxed user is a programmer that lives in chat. I’ll admit to having asked the LLM to move a method up/down in a file, and watched it burn tokens for a minute thinking and executing a menial task.
That's a big if though and the blank page syndrome was already getting worse long before AI.
With age, it becomes easier and easier to get angry at someone or something until they work as expected than it is to actually do it.
I think this is why we've been seeing the genius coders from two generations ago embracing vibe coding even before it was cool or any good.
Plus you get a bonus random line "methods are all on the top" in the commit message that makes no sense to anybody.
But, in my experience, the projects where I have a constant pulse on the core design and abstractions in the code end up moving much faster than the ones where I don't. And I've been working on one of each at work recently, so I have a decent point of comparison.
On browser based front ends it seems to be the case for me even though I still impose certain guidelines. On my C++ backends, no fucking way. Even the best models produce working but absolutely disastrous non scalable (performance wise and design wise) code unless watched over like a hen. Having said that - the value I get in either case is enormous.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
2. Make sure tests pass
3. <Every now and then> Review code for quality and fix - make sure tests pass.
4. Go to #1
Overly simplistic? Yes. But I would wager that this can go a long way, even for vibe coders.
"Review quality and fix" doesn't mean a lot without context.
Does it mean to remove unused features and simplify the underlaying code? Does it mean changing the data structures to better support future development? Does it mean improving performance because of bottlenecks?
You are supposed to tell an LLM what your codebase needs, but if you just vibe code without knowing the code, "review quality and fix" will have unexpected results
AI is amazing, but people need to realise meaning and intention can't exist in a vacuum.
Even after documenting decision and specs you need somehow to replay these in the correct sequence after you've validated these specs are still updated. Imagine you spec a feature, it works well, but in an edge case while doing a separate work you see something wrong, will you stop, fix and update the related spec? You will trust the AI will do this, and you guessed right, the counter-probability of success also has effect here, so eventually you will do undocumented changed in the codebase that won't reflect in the spec.
Now imagine all this but in the hands of someone that is an expert in their field but has zero notion what we are talking about here. Just look at the state of packages in R, the programming language, you'll see that technical competency and intelligence don't translate immediately to efficiency in a programming role.
I haven't had time to dive into it yet, but I think it might structure things in the way you want.
This is something I didn't think about. Have a tech illiterate friendly harness that will keep asking technical questions the mainstream user won't be aware they needed be addressed, until there is enough evidence to either start an implementation or outright reject the project with suggestions where the user might look into to better prepare for another session.
I still code "by hand" sometimes (mostly Ruby/Rails, C#, and random languages for code golf) but just for fun at this point. Serious projects started being 95-100% AI over a year ago.
It's not perfect but any means but it helps manage ones sanity.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and essentially the same sentiment, three decades later:
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.
If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.
Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.
Honestly, if I simply fed it a sense of presence (I would repeatedly tell it what's going on right now and ask it to react if it thinks it should), it would feel eerily like AGI.
It is useful. It may be dangerous. It has an impact. I care about that.
One of the problems is that by default, they'll avoid changing data structures or architecture that is already written down.
Like a junior dev, they're correctly cautious about breaking things, so they prefer to write more code instead.
Indeed I realized recently that, when we complain about LLMs producing slop, that's in part because we dont ask them to refactor.
Coding agents won't, on their own, make a big change the user did not ask for. And this is fine.
Long story: we have a big legacy desktop app. It uses a big legacy UI component (a grid control), which we had a license for in an old version. Fast forward 20 years, and to be able to move to a new runtime for our app, we need to update the component. Someone had bought the company making the component and now charges north of $1k per developer per year. So instead of doing this, we had just lived with the very old version.
We had long thought of writing our own control to replace the proprietary, but it was always going to be a man-year of work we thought. But I thought I'd give it a try with AI now. I told Opus: look at our uses of that control (tens of thousands of lines of code, it has over 100 instances across our User Interface). Write a new control that would compile with the exact same app syntax. First just make a dummy implementation that throws on every call. Then start implementing. Make a test suite that can run both with our new control and the proprietary control, and test everything, every function that can be called in its public interface and every state that can be inspected from the public API. Verify that everything behaves exactly the same, and lock it in with thousands of tests. Finally, check that the control _looks_ exactly the same as the proprietary one. Render to bitmaps, figure out the rendering logic from observation, such as arithmetic for padding, font sizes, and so on. Compare pixels until it's exactly the same.
Basically: it was a mammoth coding task, but it was so extremely well specified that an LLM could easily just do it. It's a clean-room implementation of something with no tests, but we had a test double that could provide 100% of the expected behavior. The description was extremely short. "Make a new thing that works like the old thing, and prove that it does". Opus 5.5 finished this in a number of hours. 500 source files, several thousand unit tests, and html reports with image diffs from the reimplementation and the original control. It did not use any disassembly or such "cheating". Only observation of the public API and the behavior. Do we need to deeply understand the implementation? Does the architecture matter? Not much in this case I'd argue. It was a black box to begin with and it remains a black box. If we notice a bug, we can always point it to the original proprietary control and say "there's a behavioral difference when doing X" and it will fix it, and lock it down with tests.
As a programmer it's kind of chilling. I had recreated for a few tens of dollars something that would cost $1000 per year to buy. Obviously it's not a complete implementation only the parts of the API we use. It likely still has some bugs. We don't get support, we get to maintain it ourselves. But the rate of reverse engineering this thing "black box" was frightening. It hasn't created anything novel. But we must realize that as programmers some times we have man-years of work that just isn't novel. And in the past, we didn't do this work at all.
I wonder if those who write and sell libraries like this will start having explicit no-reverse-engineering EULAs soon? Perhaps even explicitly mentioning AI/LLM use in analysis and reimplementation?_ Obviously the library we reimplemented was from 2005 so didn't mention AI... (It doesn't mention reverse-engineering either, luckily).
The law that covers this (in the EU) is EU Directive 2009/24/EC, where Article 5 is the reverse-engineering-without-decompilation.
> The person having a right to use a copy of a computer program shall be entitled, without the authorisation of the rightholder, to observe, study or test the functioning of the program in order to determine the ideas and principles which underlie any element of the program if he does so while performing any of the acts of loading, displaying, running, transmitting or storing the program which he is entitled to do.
This is pretty difficult to parse, but luckily there is a ruling from the European Court of Justice on this: SAS Institute Inc. v World Programming Ltd (Case C-406/10), delivered on May 2, 2012.
SAS Institute claimed that World Programming Ltd (WPL) infringed its copyright by studying the behavior of the SAS software system and writing a competing program (the World Programming System) that emulated its exact functionality and used the same data file formats. WPL did not have access to SAS's source code and did not copy any of its literal text or internal structural design.
CJEU:
> "It must therefore be held that the copyright in a computer program cannot be infringed where, as in the present case, the lawful acquirer of the license did not have access to the source code of the computer program to which that license relates, but merely studied, observed and tested that program in order to reproduce its functionality in a second program".
Which is a good find. But this is where I wonder if LLM-based reverse engineering is going to creep into either law (via lobbying) and/or EULA's, because this "observe every single state of the program for every single mutation" was simply not a viable mode of reverse engineering in the past. Or, it was at least always cheaper than just buying the software! Not so any more.
Or alternatively, that programs stop having so many observable states, making more things public. But for libraries as in this case, the whole product IS the public API. Without a rich public API, the library can't be sold. And with it, I can observe it and copy it - because it's internal workings are "too simple" not to be deduced from the public API. In short: a UI control is a ton of hard-to-write but easy to copy boilerplate code. And selling this has been an industry, but I wonder if it will be for very long.
Yes, but my point is that.... go on github, you'll find tons of decomps. And many more done just privately too. One of the No Man's Sky devtalks start with "yeah we decompiled the terrain generation from this other game, implemented it in our prototype, it didn't work okay, here's how we've learnt from it to make something better". This was in 2016. More recently, this has been going on way more openly, even full AI-assisted decomps thrown up onto GitHub casually. It might be the letter of law or included in Terms of Service but no one cares really.
I've seen bunch of persons like this and that's kinda stupid because they're just blindly following AI's "suggestions" while they actually don't know what they're doing, then results on terrible code and architecture with "if it works, it works" mentality.
I mostly use Fable though, Opus only via sub-agents.
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.
Not every user of Claude is a programmer. Or even exclusively a worker. Claude has uses beyond work. Something that many in HN struggle to understand.
I am using Sonnet 5.0 in browser (btw Claude in Chrome extension works in Edge) to download Datadog logs with multiple filters. It's running for about an hour, doesn't run out of tokens and does a splendid job.
I’m just hoping they didn’t “improve” sonnet too much or it will become annoying to wrestle into doing what I ask it to do.
Seeing 2000 years of history being replayed by the AI startups is pretty weird right?
also i find it interesting how the capabilities are growing. first speech, then code, then simple tools, then 3d objects, then desktop use
so whenever you are dealing with volume rather than independence. opus5.5 can define the goals of a sonnet well enough, that i would trust it with a group of 100s of agents
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
this is literally faster to do it yourself
> "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS"
these would not.
It's literally not unless you already know exactly where it is in the code
Also even if it is I find that the extra mental switching is not worth it, that's why I even have it do basic things like updating the text in buttons these days. There is no point in using my mouse and keyboard to track down a file and then make the change when I can just use my voice to tell it what to change and then wait a few seconds.
If your working with an active context, and changes you do then need to be conversed back to the agent, and even then, it might still find it jarring and wrong.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Fable 5.1 was pretty good. Even animating it:
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
Anthropic made it that way, and I'd say the lower score is accurate.
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.
I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.
The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
I use CC and Codex about equally. CC by far has the best harness but I prefer OAI's models.
- CC is pretty effortless in sleeping, monitoring, or waiting. If I ask CC to get CI green, it will correctly wait for the GitHub status check to resolve, check BuildKite logs, and iterate. Codex will usually just do one cycle usually because it gets confused on how to wait. I think Claude Code has a built in 'monitor' primitive which might help.
- With CC I can easily say here's a huge piece of work that will likely be 100k-200k LOC. Divide up the work and have cheaper models do the impl. CC will _usually_ do this right. With Codex it almost never works well for me. I have been able to get GPT 6 Astra to do this more successfully.
It seems weird to me that just using a ton of output tokens manages to produce a decent result in the end.
It seems to work well though. Sometimes I fear that it might be more likely to eg. run an incorrect, destructive command, but maybe that concern is not justified.
Most enterprise customers are paying per token at this point afaik, whether that’s to gh copilot, Anthropic, or running models on Vertex/Azure/Whatever
https://artificialanalysis.ai/models/comparisons/gpt-6-luna-...
GPT 6 Luna closed the gap significantly for sure (it seems to be about twice as expensive as DS v4.1 Flash), but Deepseek v4.1 Flash is still the best value model and capable enough for almost everything I need to do. Sometimes if it's babbling or can't nail down a solution I switch to Sol for one prompt, get the solution, then switch back to DS. I use 4.1 Flash almost exclusively though for both planning and implementation these days.
I was a heavy Kimi k2.5/2.6 user but since 2.7 Kimi has gone way downhill -- even the previous models. I think they got under heavy load and had to quantise their models to avoid going broke.
Is a zero-shot, zero-context prompt a useful benchmark? Yes, in the absolute sense. Does it reflect how teams would use it in the real world? I think in real-world use cases (IME), DeepSeek gets the job done.
The DS one used Claude Code via OpenRouter, the Luna one used Codex. I'd say a big difference in cost is coming from the harness, and quite possibly the different meanings of "one shot" in each of those harnesses. The Luna one probably spent less money, but might have also done far less real browser testing and testing/verification is the expensive part.
The better graphics is probably just Codex system prompt.
In terms of real world coding usage, I am finding that DS v4.1 Flash is about half the price of Luna for comparable workloads. G6L is super cheap for sure, and super capable. It's by far the best coding model from a frontier lab for everyday coding work, but DS v4.1 is even cheaper and no less capable in my experience.
Luna is obviously very competitively priced, and I'm expecting Haiku 5.5 will be strong based on this release and get back to more competitive pricing since the model family shrank this generation (though Anthropic has proven my expectations wrong on the latter before)
Where GLM, Kimi and co shine for me is when you need to offer near-frontier capabilities in your product and straight up can't afford frontier models: if you're offering Opus in a product with API pricing, $20 a month Claude Pro is offering about $500 of comparable usage in a harness that flexes to a lot of tasks.
Offering GLM/Kimi increases the max complexity of problems you can solve successfully compared to stepping down to Luna/Haiku, while letting you offer a reasonable amount of usage. Once $20 a month Pro is comparable to "just" $100 of usage in your product, it's much easier to close the gap with UX, a better constrained harness, etc.
There's something to say for flat pricing rather than per token. Even if its not a better deal.
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
> The Cyber Verification Program (CVP) is a free, application-based program that is designed to enable professionals to continue working on legitimate dual use tasks safely while minimizing interruption. If your use case has a legitimate defensive purpose and is being affected by these safeguards, we encourage you to apply for the CVP. See our Help Center article[1] on the CVP for more information.
[1] https://support.claude.com/en/articles/14604842-real-time-cy...
I use GLM directly from z.ai, they do not retain or train on your data accordingly to their TOS.
1. Opus 5.5 noted some concerns.
2. I asked it to write tests to safely test the concerns and write up a remediation plan
3. Opus 5.5 flagged and reset the conversation to Opus 4.8, murdering my usage quota and (possibly?) doing a sub-optimal set of tests.
NGL, it would have been at least polite for it to either:
1. Sanity check if I was the only/main committer on the repo (I'm the only committer, it's my side project) and then decide whether it was 'responsible' to help me fix it. (after all, I'm wanting to correct the problem, not exploit it!)
2. Warn me before just YOLOing the context to another model leading me to have to do a grace reset.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
The pricing model confuses me though (I presume by design, Hanlon be damned).
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
AI does not mean coders coding with agents all day long. AI integration is mostly for data processing, which is the promise and use of the APIs.
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
cfengine is 33 years old
Even 5.2 is doing really well in comparison here: https://labs.scale.com/leaderboard/sweatlas-refactoring
What's left are incompetent developers paired with people who get government contracts and leech off them. Ask me how I know.
That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit. This is why they can charge normal prices. Not $50 for 1M tokens.
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.
You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.
From my today's session with Sol:
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/amend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
- hallucinated several facts despite me asking beforehand to check online.
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to what Sol 5.6 was a month ago.
I can't deal with this sh*t anymore, I have no trust in the tools I use and both OpenAI and Anthropic do the same thing.
That seems hard to believe even with deepseek's own benchmarks. Not to mention for every person who says chinese ai is ahead of american labs, there's like 10 saying that they're benchmaxxed or that they're merely "decent value for money".
I haven't had these issues in multiple model generations of models.
Failing to understand English grammar? Give me a break.
> it: That still attributes two claims to SI that SI does not make. `KB` is not reserved for 1,024 bytes, and `KiB` is the standardized binary symbol.
> me: where does it say that SI reserves KB?
> it: Nowhere. You said *KiB* was SI-standardized; you did not say SI reserves *KB*. I misread your sentence and argued against a claim you did not make.
It was correct to point me out on my mistake in essence, but still misunderstood my bracketed "(alongside the SI-standardized KiB)" sentence.
Sure, it wasn't a grammar mistake as such, more like a logical one, but it still shouldn't make it. I had more than one such issues already with it, this one was most pronounced.
What it interesting, though, is the number of corporate apologists my comment brought in. It's like it doesn't matter how many times OpenAI and Anthropic have botched some of the models while keeping the branding, some people would still die on that apologist hill.
I use Deepseek 4.1 almost every day as well, it's nowhere close
It's not like rooting for a favorite football team. double woah!
Model and observed window | Messages | Retrieved actual $/message
DeepSeek v4-flash — all observed snapshots, 20 Jul–15 Sep | 4,610 | $0.0241
DeepSeek v4.1-flash — 11–27 Sep, before the 28 Sep billing change | 1,079 | $0.0576
DeepSeek v4.1-flash — 28 Sep, partial new billing window | 69 | $0.0365
GPT-6 Luna — 23–28 Sep, partial final day | 88 | $0.0329
I've subbed to Codex because I suspect at my usage rates the Codex Plus plan gives me more Luna messages than I'm using, and I've not really observed and better or worse intelligence performance. Interested to see how my $/message comes out after a month of usage on the Codex plan.
Something nice I've realised about my harness is that I can run different agents on different models so I can collect pricing data for a bunch in parallel.
[0]: pi-msg, run pi agents over xmpp https://github.com/zachpmanson/pi-msg
Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model. Their raw API prices have fat margins.
And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
i think they see what openai charges for luna and just don't want to try and compete
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
Their new Ember-1 model is pretty good, fine-tune of Kimi3 with way less thinking
Does this make it an American model or is it still Chinese? Does it matter?
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
That term is about hiding a system's design in order to secure something, rather than having secure design.
A secure system is impenetrable unless you have the key.
The world uses many security products that have false positives and false negatives (firewalls, intrusion detections, wafs, fraud detection, spam...). Those aren't generally considered security through obscurity.
They openly talk about the system and its drawbacks here if you're interested: https://www.anthropic.com/news/fable-safeguards-jailbreak-fr...
It's paired with rate limits, monitoring, account control, multiple classifiers, a deliberate safety margin, model design, and other things too.
(I think this was also why they faced the controversy over not having zdr in fable: they wanted to use logs to detect repeated attempts, etc. Possibly a bigger change with Opus/Sonnet is that it's zdr with the classifiers?)
The reason you see "dumb" refusals that "should clearly be allowed" is that they're using more traditional deterministic methods to deny prompts rather than just relying on the random LLM which you could bypass by luck.
> it’s the first Sonnet model to launch with cyber safeguards
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
low
27 input, 1,623 output, thinking_tokens: 0
1.6284
Duration: 10138ms (10s)
medium
27 input, 1,796 output, thinking_tokens: 0
1.7914 cents
Duration: 11266ms (11s)
high
27 input, 2,334 output, thinking_tokens: 745
2.3394 cents
Duration: 17376ms (17s)
xhigh
27 input, 5,730 output, thinking_tokens: 2535
5.7354 cents
Duration: 41882ms (41s)
max (failed to return response)
27 input, 128,000 output, thinking_tokens: 128000
$1.28
Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Score Tokens Reason Cost
Kimi K3 Max 44 48k 32k $2.00
Half reason 44 32k ? 16k ? ?
Opus Med 51 26k 12k $1.34
Opus High 54 36k 18k $1.82
Opus Max 58 119k 84k $5.98
Sonnet Med 41 ? ? $0.59
Sonnet High 47 ? ? $1.08
Sonnet Max 56 193k 142k $7.60
Medium is Anthropic's default.Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
PS: the next human that brings up pelicans on bicycles should try to draw them.
Any model release it’s the top comment, I do not understand why.
In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591
I find it useful (as well as a fun art project).
You can also check for any kind of degradation of them - you have the prompt, it doesn't use much $.
I’m not sure whether that’s a feature or a bug at this point though.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Now imagine you have a set of twenty of those tasks. You launch an Opus agent, give it the task list, and tell it not to do the work itself, but orchestrate agents to perform all of the tasks and do small spot checks to verify the work.
The overall task is completed much faster at a similar or cheaper cost with an extra verification layer inserted that wouldn't have been there if you just used Opus.
But yeah, it's completely true that you sometimes have the ironic situation where you actually pay more with Sonnet because it's worse at reasoning itself down a rabbit hole. Sonnet really should be capped at medium or high reasoning.
Ironically, one of the worst things you can do is to use Sonnet to organize sub-agents. It seems to be completely bonkers with what it asks agents to do. I tried to make a colleague of mine test the feature, and he, by accident, started a large bug hunt with Sonnet. It spawned 250 sub-agents and spent 5 hours looking through everything. It actually did find a couple of useful bugs, but not the one we were looking for, which is a stupid race condition probably.
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.
At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.
It's been out for an hour and you've already concluded this?
If you look at Anthropic's own benchmarks any thinking levels above medium quickly approach the cost of Opus 5.5 and even exceed them.
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
https://bench.killswitch-lang.org/
Claude Sonnet 5 17.8%
Claude Sonnet 5.5 7.4%On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
In xhigh effort it is a lot cheaper and possibly lot less impressive?
It's luna level, but 14x more expensive:
https://aibenchy.com/compare/anthropic-claude-sonnet-5-5-hig...
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
Is there any tool which allows for one model from Anthropic to call subagents or dynamic workflows using other providers?
How does one create a swarm of agents from different providers and get them to talk to each other, or otherwise hand over pieces of work to one another while being able to check that offloaded work status?
* 5.5x faster
* 4x cheaper due to using fewer tokens
* 1/3 the turns: batched reads, one-script edits & tests in the same call
Gemini 3.1 pro as a nutty professor, researcher, and design verifier.
My custom orchestrator for all local LLM work: https://vektormemory.com/vektor
Getting hard to keep track of opus vs sonnet vs....
Assume 6 will be announced on around an IPO?
if you're on free tier, using Medium settings is far more intelligent than Max, and tokens don't run out so fast. Max is cranky and verbose, Medium is patient and somewhat goofy, but does the thing as expected. At least within Sept 2026 this has been my experience.
Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose
lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
The communication and writing style also feels closer to Opus 5.5.
Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.
On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...
Also, these are benchmarks...
I'm very curious how do they know what requests could assist competing AI models.
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
OpenAI and Anthropic's lead is vanishingly small at this point.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
This is bollocks. Their safeguards are shit.
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
This is the way.
https://support.claude.com/en/articles/14604842-real-time-cy...
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md
It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.
---
For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.
I use AI the way I want, you don't force me not to use SOTA for this