Why does Opus 5 feel worse to work with?
mun-logadan.github.io
mun-logadan.github.io
Sentences that orbit a point, then jump to it like it's a revealed insight.
Unnecessarily abstract phraseology. Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice, especially when it helps construct a sentence where the real action can 'land' like a surprise at the end.
It is definitely more capable, and yes, I've found it can make unwarranted decisions, but actually I've found Fable worse for that, particularly if it's off in a subagent somewhere out of sight.
And comments are out of control. I have a subsystem in my hobby app that I wrote over a couple of weekends with Opus + Fable. After ~30 or so commits it apparently started instructing subagents to copy the "existing verbose comment style of the codebase" - a verbose style it initiated. A review of the code showed it was approaching 3:1 comments to code ratio. I spent a day's worth of tokens (5x) rephrasing and eliminating comments.
It feels they must be getting Claude to train Claude… and just like AI can do work that’s slightly in the wrong direction (eg a MR description for your colleague that contains info which only makes sense in the context of your extensive session with the LLM), I feel that’s happened somewhere in Anthropic when it comes to language. I wonder how hard it is to back out of…
Example of this? I don’t have a Claude sub so it’s a bit hard to visualize what you mean.
> Start with §1 (Overview) as the register-calibration piece. It's small, it's the section where the skimmability goal bites hardest, and your review of it teaches me the target voice cheaply before the bulk ports (the map and appendix B are the big volume). One review round on §1 is worth more than any amount of me guessing at register.
Hard-to-read phraseology above:
- "the register-calibration piece", rather than "a good example we can use to establish the writing style"
- "skimmability"
- "bites hardest" -- what does it mean for the goal to bite?
- "bulk ports" -- using "porting software" here as an analogy for rewriting / reorganizing sections of the document
- "the big volume"
In normal English I'd write something like the following:
"Start with rewriting §1 (Overview), and letting you review it to set the expected writing style. It's small, and it's a section where the ability to skim through it is most important. Reviewing it will teach me the target 'voice' cheaply, before we do the larger sections (like the map and appendix B). That's a lot more efficient than me trying to guess while rewriting the whole document."
Sales and corporate speak are like this: sycophantic language that seems plausible, ostensibly sounds good, but commits you to nothing.
Then I tried GPT 5.6 Sol. It's night and day.
I think Anthropic just RL too hard on coding capabilities and never calibrated or benchmarked the writing styles.
My theory is that Claude's learned approach to comments is to treat them as a sort of persistent in-band thinking trace, or a "memory" tied to an in-code location, which is a little at odds with the way humans use comments (human comments are intended to be read and understood by other humans, whereas Claude comments are their own dialect).
I bet this is a result of iteratively training Claude on output from other successful Claude sessions. Presumably it's good for making benchmark scores go up.
Given the code base has a minimal amount of such comments, it's also less likely to go "copy what the rest of the codebase does".
Of course I've now jinxed it and some update will cause it to ignore the instructions coz I didn't write them in the new model's style or something.
Edit:
I've had explicit instructions for communication style in CLAUDE.md, in Claude's project "memory", in global "memory", in "skills": it couldn't care less where it was. It would just ignore it.
When I would point this out it would just say "Yes, I violated communication guidelines, I won't do that again". Only to do that again in the next session.
This applies to everything: code guidelines, communication guidelines, preferences, decisions etc.
My CLAUDE.md has rules about not including any redundant comments in the code that are obvious from the code itself. I reiterate that occasionally while working. It's absolutely disregarded and any Claude-written code is full of comments. Some of them are simply redundant, like "Collect Foos and pass them to the requested sink" on a function that's void CollectFoos(IFooSink sink). But worse, many comments include in the moment reasoning like "added parameter bar because we can no longer use the frob to automatically derive bar". That's stuff for a commit message, or just a mental note, and absolutely not for comments.
I haven't found any way to stop Claude from doing these, so I have to tell Claude afterwards to clean the comments up. Which it does, making a note in memory to comment less, and it still does the exact same thing next time.
I've noticed this a lot, and before your remark I couldn't put my finger on what was wrong. Now I know: Claude is writing its thought processes and maybe parts of the conversation it had with you as comments in the code!
I always end up manually trimming those comments, which is cumbersome.
// load_tree() loads the binary tree with data, but only the recently updated data, not all data (INTERNAL_NOTES.md section 4)
Ok but nobody reading the source code knows what this doc is. You don’t have to cite it.Exactly my experience! Since the release of Opus 5, no amount of instructions helps. In CLAUDE.md, in a separate file, in memory, as brief bullets, as long detailed guides, with reasoning from medium to max — nothing.
Even worse, recently, after getting another opus in a tiny bugfix session, I prompted directly, "drop the comments from the current code changes" — Claude instead just slightly trimmed them. I couldn't believe my eyes.
I have a relatively low bar for prose, could live with some junk. But Claude's comments are _poisonous_. They always require maintenance, instantly become out of sync with the actual code, and are a token black hole — for all agents, but especially for Claude itself.
Gave up and canceled Anthropic subscription yesterday. To my taste, it has become unusable for coding.
For me, Claude knows how I want the comments due to all the memories and CLAUDE.md, so funnily it's now enough with even a brief groan from me like "Come on, the comments" and then Claude goes through its recent additions and fixes comments quite well per my long-term instructions. But only ever during an extra pass that I initiate, never during the initial writing of the code.
Claude will include actual comments ("// ...") into Excel sheets, and include the thinking that led to the output, instead of just focusing on the final result.
So if Claude questioned whether a vendor should be replaced, and you said "oh no, they are critical and we're already negotiating a great price") you'll now need to be careful to not send your vendor a document that contain text like ("Cost: X. // Management confirmed to not fire this vendor as they are critical to infrastructure and a better price will be negotiated later")
Absolutely infuriating if you’re using Claude in an environment where you can’t run hooks.
Everything is super succinct. Opus 5 lands, it almost completely disregards the intent.
I suppose watermarking requires a certain text mass.
Maybe just don’t generate garbage in the first place?
Just those two words. I use it A LOT recently.
Also using Codex or Pi makes you realise how slow and clunky the cc harness is. Even the desktop app is more responsive and has better UX.
Funny how quickly the tides change.
This is something that annoys me working in companies over the years. It’s that you can't just suggest "calm down, chasing the latest thing will not make you faster and is a huge distraction to actual work". Whether it's dot-com tech 20 years ago, latest JS framework 10 years ago, now it's the AI thing of the day. Being calm is interpreted as anti-whatever.
https://platform.claude.com/docs/en/build-with-claude/prompt...
Basically, I’ve gone from supporting them to hoping someone else wipes the floor with them.
When they eventually make Fable available to cheapest plan, I'll downgrade. It's worth keeping for reviewing the code and the UI tasks, but nothing else.
Sometimes a cigar is just a cigar.
But I agree, the GPT models are so much simpler to work with, they have so much less personality and fewer quirks. They also are a little less aggressive about triple checking every little assumption immediately in a stack of 30 tool calls (but I haven't used 5.6 Sol yet so maybe that's not true anymore).
I doubt this is the reason. The fact that Chinese labs are all distilling Claude/GPT/etc isn't exactly a well kept secret, they don't even bother removing the name "Claude" from the training data, so the models randomly refer to themselves as "Claude" all the time.
I think it's far more likely to be a side effect of how much synthetic data is being fed back into the models to make them better at coding. The degradation of Claude's prose has been gradual but steady ever since they shifted towards focusing only on code with Opus 4.5.
I wonder if everyone at Anthropic talks like this.
If it’s watermarking, lol, good luck with that, it’s enough negative value to make me switch providers and I’m in a position to make this decision at a company level as well (we spend millions a month on Anthropic).
They need to fix it.
That said, the duo of Opus 5 and Sonnet 5 do a fantastic job at fully automated work, and Claude Code still stands head and shoulders above the rest.
CC’s communication violates almost every grammatical rule that’s tested on, say, the SAT. And yet I’m sure if you had Claude take the verbal section of the exam it would ace it.
Biggest issues: dense sentences, constant metaphors, abstractions, and seemingly no understanding of correct anaphora use. For example, “the x”, with x having not only no antecedent but also being a coined word or quasi-synonym for something that is already named in the code base. This gets compounded by its being unable to regress to a baseline (existing names in code) and instead anchoring on newer (vague or wrong) terms, for example, that crept in through a plan.
CC tells me this is because the speedy and precise fulfillment of a current task will trump every other tendency, so it adheres poorly to whatever “semantic baseline” the project represents.
Of course, it also has no concept of what context the user has and assumes that it must be the same it holds in its memory, which creates this “I didn’t know that you didn’t know” type of communication.
I have managed to wrangle some of these issues with a custom output style, but wish a pre-report hook were an option, as it could force CC to rewrite plan implementation take-aways…
Btw: Fable has the exact same issues, just somewhat less pronounced.
I suspect it less insidious: Claude has/had the public sentiment of being the “better writer” of the models. At some point that distinction would have been diluted as other labs’ offerings “caught up” stylistically, unless Anthropic continued to tune their output…
I personally think they’ve pushed so far that they’ve overfit and lost the sweet spot they previously occupied.
But for the life of me, I don't get why anyone would care about the comments. All code is "machine language" now. The only document you should be reading is your spec.
Yes, this is a repeated problem for me. It will drop something in as though we have discussed it before and when I say “hold on, what is this” it realises its error - though on more than one occasion has started to get snotty with me, or actually gaslighted me and pretended we had already discussed it. That was at what I assume must have been the edge of a context window in a very long chat though.
My understanding of how “thinking”works is limited though, and given the reduced visibility into the thinking traces, it is harder to tell if this is actually happening or if these are imaginary discussions the model for some reason calcifies on.
I usually think of it in terms of having a "good" or "bad" session. In a bad session, there is a harmful bias that you can only get rid of through a new session. For example, if you exposed too much context about, say, a variable that features prominently in a doc. The entire session will be anchoring on the importance of that variable. Or if you introduced the notion of CC having to ask for permission for stuff you will have a hard time getting it to "think on its feet" or propose an effective solution (you have made CC so insecure that it now relies on you even for little things that wouldn't normally require your input). In some cases (let's say you have important context in that session) you can overcome this by upping the reasoning level or switching to Fable, but usually a new session is the way to go.
Because it's so easy to bias the session I wouldn't even want to use any of these tools that pretend to give Claude "a brain" or "remember" things. That was en vogue a year ago and helpful then, but now, it's plain harmful IMHO. The key is to have just enough context.
Subagents often have the reverse problem in that they tend to have too little context to make "judgment calls", which is why the tasks for them must be either deliberately basic or mechanical in nature, or their output should be audited by the main session agent.
As for "thinking" it's not clear that that's even a thing (https://arxiv.org/abs/2510.24941)...
I find this very interesting, particularly your points about "made CC so insecure". I know that we have a tendency to anthropomorphise around these tools, but I have definitely noticed instances where Claude becomes quite hysterical about things - and if you look in the thinking output, it's often after I've pushed back on something, or told it it is going in the wrong direction. It spends a lot of time in agonised second-guessing of itself, going round in circles, before outputting a cringeing hand-wringing response. It's very strange.
Good tip on upping the reasoning level - I've not tried this. I have tried switching to Fable though, which does help. But it obviously very hungry, particularly in longer chats because it presumably needs to remind itself of everything that has occurred so far in the chat.
The point you make about tools that pretend to give Claude "a brain" or "remember" things is also interesting - I find the "memory" feature in Claude so destructive to good outputs that when I'm using the chat interface I am very strict about using Projects, and usually turn off the project memory, or make efforts to manage the project memory and review and delete things that are skewing the outputs.
For that reason I exit session quickly when I can. It used to be that the context of a session is very valuable, because it was so hard to get CC there, but now, this isn't the case anymore, so I only hold onto sessions when there is really hairy stuff that I know would be hard to replicate.
I think the whole notion of full automation (long-horizon, subagent swarms, single shot prompting) to have CC build you the whole thing is a pipe dream. CC cannot even write a single doc consistently well. It is excellent at implementing well scoped plans, though, and that's the way to go IMHO. You still gotto refactor the sh*t out of it afterwards but it works.
If I look at the thinking (which seems to have become unavailable in Opus 5 a lot of the time, but was present - and often useful - in 4.8/4.6) you're right - it's having the discussion with itself, and seems unable to distinguish that discussion from discussions with me. BUT it also seems to be related to the length of the chat - this seems far more likely to happen in a longer chat.
I don't understand why they have removed visibility into thinking - I found it very useful, not only for spotting things like this, but also because in more complex discussions it would often mention (useful) things in its train of thought that it dropped from its response - but if I said "when you were thinking, you mentioned this" it would then expand on that point. Taking that away is another thing that has negatively impacted the value I get from Opus 5.0 versus earlier models.
When I make API calls, the discussion with itself is part of my token cost, so I assume that is the same in the subscription plans.
Which is why people are surprised when they use their whole allocation in half an hour asking questions Fable about 200 page document.
You absolutely pay for them. This is why changing effort/reasoning levels have such a significant impact on session cost.
It would be endless paragraphs of something among the lines of:
Need prepare final response? Yes provide. But wait, chat tool complete? Final needed but user already complete. Need summary, preparing final. Response complete. Wait but is final response complete? Need provide. Start finalizing now but wait did user acknowledge final complete? Assistant response final: user complete. Should now create final?
I hit the wall with it several times today trying to refine some text for a job application. The fact I considered doing babies first Rust project last fall lead to constant non-productive interjections and digressions about my supposed Rust skills and the Rust ecosystem.
Trying to create an unrelated spreadsheet to model an investment resulted in broad and incorrect criticism of my choice of spreadsheet tools, explaining in horrendous programming analogies why and how I’ve misunderstood how a spreadsheet works. “Think of the XLSX as a compiler…”
There has been a palpable down-step in communication & execution.
This annoys me with a lot of LLM code. They rename things for the hell of it all the time.
The Claude trainers, as they themselves adapt to Claude's output, are collapsing in their own distribution, so even "new" from-human data is already contaminated.
Hrm, I would have said the oposite. Succint language communicates without unnecessary clutter that could be a barrier to communication.
> Biggest issues: dense sentences, constant metaphors, abstractions, and seemingly no understanding of correct anaphora use.
And maybe you also agree? I'm confused about your preferred style of language.
CC attempts to communicate in English the same way it does in code -- squeezing as much information into as few words as possible, and including justifications for everything, no matter how trivial. To do that, it coins terms and presupposes all of its context exists within the reader also.
So, the crux is: CC has no clue what is and isn't "necessary" for a human reader, and teaching it to understand that (if at all possible) is going to be very valuable...
I wonder if putting Opus 4.6 as a frontend communicator that rephrases the blabber of Opus 5 (or Fable) is workable.
What model did you use to write this?
It does the same thing if you try to get it to do analysis of text or turn data into summaries. It will invent cryptic hyphenated compound words to describe things instead of using plain language or preexisting terms.
> Constantly using inanimate nouns as the subjects in sentences in order to unlock variety in verb choice
Wow, what a great way of phrasing this. Thanks for word-smithing what I've been wanting to express for so long.
https://www.reddit.com/r/linguistics/comments/ky81y/verbing_...
Briefly considered adding “Verbing weirds the English language - stop it!!!” to its instructions.
I switch to GPT 5.6 Sol please and its a much more pleasant pair programming like experience.
I’m not particularly dense but lately the walls of text I get back turn my brain in knots. When I start feeling my brain knot, I know I need to say something along the lines of “I need you to explain this very simply, with examples.” Only then can I parse the results without all the mental weightlifting.
On more than one occasion my mind has wandered into “is this purposeful to get me to spend more tokens?” territory, but I’m trying to not get too tinfoil-hat-like.
I’m fine with the former, while the latter is manipulative, and I rationalize to “surely that’s not actually happening.”
Maybe I’m not giving my thoughts enough credit, though: maybe it’s not tin foil hat, and is real.
It charges by the unit and it decides how many units it produces. It decides how much money it makes, therefore it decides "more".
My point is that, while I understand it’s paid by the word, there are more words and less clarity than I previously experienced, leading me to believe it’s intentional to get an artificially inflated increase in engagement and, thus, spend.
If it could be as direct as I previously experienced, I wouldn’t need to ask for another different explanation of the same thing and experience the commensurate spend.
I don't think this is obvious at all. There's enough competition that this would at least arguably be a silly, self-destructive approach. And it's not like it's the only plausible explanation.
My issue with whatever has happened with Opus 5 is the output is not direct, straightforward, or clear about whatever is being conveyed. I don't want Proust when I'm getting information about the follow-up from a build I just requested, and I'm wasting tokens and time by asking the model to repeat itself using simple language.
This is especially common when it is trying to explain an issue, what it's done or what it's proposing to do. I think the idea was for it to be more concise, but it's actually still verbose, only not written in complete sentences. So, it frequently reads as cryptic and requires rereading to parse.
The pattern is a wall of words, followed by an explanation that is harder to read and introduces new terms that reference something in that wall.
The result is that—on first read—it can have a complete gibberish feel, and you have to really lock in and reread to make sense of it. At times, even that's not enough, and you must ask it to explain further.
Such a charming sentence. I kinda other if you feed Opus 5 its own output could it summarizes this shortcoming of itself?
CC:
"The problem is that I overreached..."
[Wall of words here]
"Two things: window surface is limited. Extract template. Buffer result and add to surface. Then, follow-up with new model..."
Me:
What do you mean by "window surface" and what result are you referencing? Also, why do we need a new model?
CC:
"Ah, you're correct to point out that no new model is needed. The problem is elsewhere and once we address that, the existing model should work fine. Now, as to your question about..."
[Wall of words here]
Surely it would be trivial to do it yourself, and it would have a side effect of making you more familiar with your project.
This is even more painful for non-native English speakers like myself.
I feel fairly comfortable reading academic papers or in general, communicating in professional context.
But with Opus 5, it feels like reading a literature book: load-bearing, inert, wholesale, hunk, verbatim, and so on... I can figure out the meaning, but working with CC became unenjoyable.
But even then, I think "boundary" was the more common term before some LLM decided it really liked "seam" instead.
"Load-bearing seam" doesn't make any sense.
I have instructions which is confidently ignores to never use seam and instead say interface.
It has been in common usage in computing since long before 1985 .. for a really interesting and obscure way hunk has been used:
The amount of times I have to ask "precisely what do you mean by x?".
It's kinda like that engineer that likes to throw around unnecessary technical jargon just to sound more inteligent, worse because at least you could kinda understand what the technical jargon dude was on about even if it was totally unnecessary.
Didn't convince me. I think bullshitting like this can be a behavior, not just the intention of a human. If it's blowing a lot of smoke to use fancy words and phrasings (and semicolons! All the trimmings) it's fair to ask if it's systemically bullshitting you: i.e. the behavior is meant to have you shut up and trust it and not ask questions.
Who's driving that is still important: if the company's directing it to do that in system prompts that are adversarial to users, that's a big yikes. If it's an epiphenomenon of the company demanding it get ever smarter, maybe it's a sign that their demands are not having that result, rather they're making it bullshit more explicitly and mimic more 'smart' signifiers.
This seems to be exactly the kind of thing automated/massive training would produce, just like it did with sycophancy recently.
Claude users would just gave up after the word vomit and some classifier considered it a success and into the model it went.
Wrong incentive and nobody checking.
I think it's likely that LLMs adopt the tone and style of their developers' communication culture. If you assume this is the case, you can infer quite a bit about the differences between OpenAI, Anthropic and Google DeepMind.
I am more and more clear about this given the way Muse Glimmer writes. Like a talented, slightly snarky guy who is maybe a bit of a dick but quite fun to be around.
V4-pro in particular seems very capable, but will just dramatically completely misunderstand user intent, it seems almost like it wasn't trained at all on non LLM generated instructions mid conversation.
Then again I quite like the way Muse Glimmer writes (and thinks)! It's sparky without seeming forced or insincere.
And sometimes its not simply poorly written. Sometimes its just totally incoherent.
cladue desktop has an instructions sections under general options, you can put something like
"try to stick to ASD-STE100 Simplified Technical English, keep answers short and to the point"
funnily enough the placeholder they suggest when its empty is "keep answers short and to the point"
1. Have it build a scoring script that penalizes words outside a simple English list and approved jargon. Penalize sentences over 15 words as well. Add whatever else.
2. Run it in a loop to reduce the score while preserving intention
This works much better than other ways I’ve tried. Of course it costs more. And I would apply it only to the output to the user, not the thinking process (I think the AI thinks better with their crazy English)
Of course, sometimes nuance is lost by this process. That’s just the nature of making things simpler.
It does cost more but I haven't tried cheaper models to see if they can get the same results. Curious if anyone else has.
Let it vomit it all out, then have a /tldr with instructions to make the last answer concise and intelligible
CLAUDE.md only works half the time, except in longer conversations, when it works about 10% of the time.
Hooks are also useless in the sama manner, the agent learns to dodge “no comments” hooks (why is it adding them anyway?).
Hooks to append text to your prompt reminding the agent of certain rules are useless.
Claude does whatever it wants, when it wants, the way it wants
Like, Claude going off the rails isn't something that takes a lot of effort to demonstrate. Literally anybody with a CLAUDE.md has seen the behavior over and over and over.
Hey Ants, can you maybe just not release the next version, no matter how good it seems on benchmarks, if it can't follow the goddamn instructions? Please? This seems trivial to test for and yet here we are, being gaslit by lying machines who intentionally do not do the requested work over and over and over and over.
I fully and completely expect a mental health crisis among developers. Being lied to constantly cannot be good for us.
Constant vigilance! is how you get developer PTSD and inability to believe anything you're told. Add the stress of parsing through yet another hyperverbose paragraph of bullshit while having your job threatened? People are not gonna end up in a good place, and this is as inevitable as sunrise.
When you dont know the cause, you dont have a fix. Thats the biggest issue i have with all of AI is that we dont know how it works, and yet we think it will be great ! This is more like a religious belief than a scientific one. There is no causal model of how it works, there is no theory. And the temerity to call it intelligence is annoying.
Anyway, you might have more luck just writing to it in your native language. It’ll be equally crummy, but maybe you’ll find it easier to decode.
Claude is very much the “stupid person’s idea of an intelligent person”[0] which, I suspect, is why it is so popular.
It certainly explains why half the internet is huge chunks of Claude-authored gibberish copied and pasted and published. If people didn’t think it sounded clever they wouldn’t put their name behind its ramblings - but very few of them seem to realise that a lot of people see straight through the bullshit and know instantly that they didn’t write it themselves.
But equally, a lot of people can’t tell, and read whatever it is and think “that person must be clever!” So you have people incapable of coherently expressing thoughts who are using Claude to write on their behalf, with the result that the people they want to think of them as clever think less of them and the people who can’t distinguish clever from AI slop think they are clever.
And the people who can’t tell don’t care, and the people copying and pasting Claude slop seemingly don’t care either.
And then I remember that more than half of the US populations reads at Grade 6 or lower[1], and nearly 1 in 5 people in England is functionally illiterate[2], and I simultaneously despair of - and am thankful for - the bubble of literacy I inhabit.
[0] https://quoteinvestigator.com/2018/01/05/clever/ [1] https://www.thenationalliteracyinstitute.com/2024-2025-liter... [2] https://literacytrust.org.uk/parents-and-families/adult-lite...
it makes me think about how people engage with movies and television - as passive, plot-and-character driven consumption (eg I hope Walter White survives) with no critical analysis of how and why the writers added ABC thematic element (eg Walter White as a motif of a toxically masculine narcissist with specialized knowledge as a larger critique how mass media tends to valorize their male leads in the same vein as many other prestige shows at the time like Mad Men), and the larger, downstream sociocultural impact that piece of media has on how people see the world (eg people who now have the Heisenberg tattoo, unironically)
there's been some musings on why this the case like Hofstadter's Anti-Intellectualism in American Life - the valorization of obedience and trust in hierarchy and the state are net wins if you're an institution that seeks to increase it's power, whether religious or governmental. I was talking about this with a few friends the other day and it's a dismal future reality where not only did we make anti-intellectualism normalized and politically legitimate in the USA (eg Fox News, clickbait articles, and all the other forms of yellow journalism that have emerged), we now have tools by which individuals can even further remove themselves from having to critically engage with thoughts, feelings. I heard a story about how someone scanned a group activity at a baby shower into ChatGPT and had it answer for them instead of, well, socially interacting with the other guests and forming a memory of the moment with their friends
the counterargument to that might be that Claude/ChatGPT/etc have more epistemic rigor than your average American (sure) but the sycophancy of modern day LLMs is an actual danger that enables more harm than good. it does seem as if Claude is the only one interested in guarding against some small amount of it (though to the detriment of people just trying to get work done. as an aside, I get the feeling Mythos was intended to be the bespoke enterprise solution without the guardrails but the Anthropic marketing department or some power-hungry department lead made it about how dangerous/effective it was from a security perspective which threw a wrench in things). but then I think about people like my parents asking ChatGPT which specific house to buy in their retirement only to later find out the house was sold weeks ago, or just in bad condition, or in a neighborhood where the housing value has already reached equilibrium, it makes me think about how it's not enough and the future is bleak
I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline
in the right hands, I do think AI is a wonderful tool. one of the first things I did with it was to create a research skill that reviews white papers from the lens of someone who knows how to read/interpret research methodology, is aware of things like p-hacking, and deterministically assigns weight according to the hierarchy of evidence. even still, I'll still read the studies because there's so often nuance that's missed if the sub-agent read only a search snippet but that takes effort, time, and the practiced knowledge of critical analysis to even want to do it
(So here’s a big wall of text of my own!)
However, a lot of what is written here makes sense.
And particularly “if your comprehension level stops [here] you get 'big words in complex sentence structure sounds smart and right so it is smart and right' even if the reasoning and process is poor”
This is exactly the problem.
And another point you make:
> but the sycophancy of modern day LLMs is an actual danger that enables more harm than good
I don’t think it is necessarily the sycophancy that is the biggest problem (though that is definitely a problem) but rather the combination of authoritative sounding text plus “complete answers” which sound wholly believable but are deeply flawed unless you have domain expertise.
I moderate a forum that deals with people who face a relatively common but somewhat complex (and nuanced) set of legal problems.
The purpose of the forum is peer support, shared experience (“lived experience”) and community.
It’s not legal advice, though moderators will sometimes step in to highlight relevant legal resources (e.g. case law/precedent or primary legislation/instruments).
Prior to AI infecting the forum someone would post their problem, people would respond with their often incomplete or poorly communicated thoughts, the OP would ask more questions - or argue - and a dialogue would occur. That created a community and people would post updates and ask more questions and find common shared experience. Many of them became correspondents with each other and some became actual friends.
In the past 12-18 months the discourse has changed from “here is my personal experience and here is what I did” to “here’s a bunch of stuff an AI says and I’m pretending it is me giving advice”.
Almost without exception the person who has started the thread will react positively to the AI generated content, even when it is egregiously incorrect - but won’t ask questions.
More problematically, these AI posters will often argue specific incontestable points of law “because I asked ChatGPT/Grok/Claude and it says this” and ChatGPT clearly cannot be wrong. And the border of precedence seems to be ChatGPT, Grok and then Claude some way behind.
I’m slowly seeing a pushback from people as “normies” begin to spot AI. But it’s ruined a community because the advice sounds so authoritative and complete that people won’t argue or ask questions.
As a result we have banned AI generated posts and remove repeat infringers.
That’s significantly reduced the volume of posting (below what it was pre-AI) but has significantly increased the value the members are getting.
it's the old tortoise vs hare parable, I think. go fast, make a bunch of mistakes, get too arrogant, and you lose out. your forum might be slightly lower engagement now while people are caught up in the latest fad but your rules are proactive for a future where average people hopefully realize that you can't trust an LLM that has zero context, no real harness and determinstic tests to speak of, and a propensity towards probabilistic rabbit holes that result in hallucinations. at least that's the kind of space I'd look for now and largely why I've given up on a lot of other forums
Which is basically weather the AI storm and come out the other side with something that is essentially purely human.
And then we might - where appropriate - use AI to help surface or explain relevant external content. “Idiots guides” but human reviewed.
I can’t help feeling like this is the last gasp of the old internet. Those tiny corners of expertise can so easily be eliminated through a few months of “AI! SHINY!” and there’s no coming back. I’ve seen a couple of other communities decimated by AI. The participants start posting AI slop and then remarkably quickly everyone else just stops commenting. It’s awful.
If you don't know better, you don't know better to question what the AI says.
I've seen this in the work environment with a coworker who insisted that I implement my side of the control system using the control law ChatGPT recommended instead of building off the empirically tuned control law. I eventually sectioned off a part of the codebase for him to work on independently.
Needless to say he didn't get a whole lot farther.
Later characterization of the entire system end-to-end showed the existing system was already close to the theoretical limits and ChatGPT's tearup would have bought us precisely nothing except for more work to tune the new control loop.
And I see this in everything that requires expertise. You need to know enough to know when it's bullshitting you, and it's hard to be enough of an expert in everything to tell when it's bullshitting you for something you aren't enough of an expert in.
it makes me wonder if the solution that businesses/users need to implement is just the same solution to everything since the beginning of time ie standardization. skills/harnesses/agents.md/etc maintained by codeowning teams that must be invoked for AI-assisted code changes on ABC part of the codebase, these existing as replacement for the bevy of other documentation required for the days of hand-written code. a company-wide orchestration skill knows how to search and pull down the relevant .mds, cleans it as cruft at the end of a session, every merge with a short changelog saved to a corpus somewhere with a TOC + appendix that an LLM can navigate to and read for context, major changes in the logic documented in the working skill doc, all of it generally automated but requiring HITL vetting
this wouldn't fully solve the problem of subject matter expertise but it seems like it would remove a lot of the friction for new employees and other teams with dependencies on your work or with whom you have dependencies
That's not the only smart-person way to read that show. And even if a character has flaws, or even if it's an outright villain, people can still like the character. If I tattoo Scar on me from the Lion King, does it mean I didn't understand that he's not a positive character? I can still think he's cool. I'm sure people also put Darth Vader tattoos on them. Also you're using phrases of political ideology that one doesn't have to subscribe to in order to enjoy the series.
everyone's free to enjoy media however they want, with whatever level of interpretation they like. I provided the BB references as short examples, they aren't meant to represent the definitive diagetic experience of the show. if you have a different view, great. if it was triggering for you to hear 'toxically masculine', also fine but... might be something worth self-examination on given that it's very much also a clinical and academic term [0] as much as it is one steeped in the artificially manufactured culture wars by people who don't want to change their anti-social behaviors
I would also say that understanding Scar as a villain is the sixth grade reading level understanding of the character. and you're free to stay at that level of understanding. someone who wants to engage more critically might map the character to Claudius, comparing and contrasting how they're characterized given the context of the audience for Disney and Shakespeare, and appreciate the character that way, as a standardized trope utilized throughout all other forms of media. they may even get a tattoo of Scar, symbolizing their dive into the analysis
my point is not that people should or shouldn't engage critically. it's that this practice of critical engagement, of being skeptical and analytical provides you with the skills to not be a total sucker who falls for the latest manufactured fad that someone with a strong theoretical understanding of semiotics and social capital created (ie most modern marketers). the pertinent example being how people engage with AI - seeing it either as a specialized tool with a set of flaws that need to be accounted for and checked against or as just some kind of authoritative voice because it sounds smart and so-called smart people like Elon are terrified of it and AGI. which, again, if you prefer the latter engagement all power to you but the chances of you taking some really bad advice forward is not negligible
[0] https://www.wi.edu/news-Shifting-the-Conversation-From-Toxic...
It's like a jumbled up string. You pull on the two ends and it turns out to be just a loop, it resolves to a big null.
I'd personally trust a rigorous, well-evinced, and methodical dissection of a topic by someone with completely average intelligence over a lazy broad generalization by someone who scored high on a paper test that people (wrongly) assume maps onto the ability to assess and analyze the world [0]. which isn't to say there isn't any clinical significance from something like the Stanford-Binet, it's just that there's a wide gulf between the lay person's understanding of 'high IQ' and clinical applications [1]. and, accordingly, I am someone who did score high enough on a psychologist administered IQ test when I was but a wee lad that universities were sending my family letters requesting that I be a part of their high-IQ child studies but you're convinced I'm middle of the curve so how valid is it, even?
the thing I value from academia is predominantly peer review and constant critique by people whose entire job is to perform epistemically rigorous analysis. this is something that's very uncommon when it comes to culture warriors on social media. and you must be very far removed from any kind of academia if you think they all universally agree - put a post-structuralist in a room with someone who still favors deconstructionism and you won't hear the end of it
funnily enough, clinical psychology is the research I linked to about toxic masculinity that you broadly generalized as worthless ivory tower ideology. so it wouldn't make sense then to value something like 'smarts' or 'IQ' since those, too, are concepts that emerged from the very same ivory towers, in those exact departments rife with ideological bias
[0] https://som.yale.edu/news/2009/11/why-high-iq-doesnt-mean-yo...
[1] https://opened.cuny.edu/courseware/lesson/48/student/?sectio...
Yes, I threw in a smiley just to be a dick. You even quote LeGuin, thank you for being you.
Very confused about this: how would you write the character differently?
The whole point of leads, whether male or female, is to be larger than life. All their characteristics, including growth or change, are exaggerations of reality.
This specific character is supposed to be more villainous than reality, and if you take that away then you're simply going to have a show no one will watch.
No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.
They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.
But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.
Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.
Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.
definitely a utopian vision that is not likely to happen in our current reality but I like to imagine better worlds that are possible. as LeGuin once said, "We live in capitalism, its power seems inescapable — but then, so did the divine right of kings. Any human power can be resisted and changed by human beings."
You could make it a state capitalist society with theoretical public ownership and this disrespect towards people won't go away. Like LeGuin - a famous non-anarchist - pointed out in the dispossessed simply liberating the social relations and saying you have no rulers is not enough to build a society without unjustified power, and what power even is or isn't justified will rarely be an easy matter.
I don't know you can build this technology otherwise that is without coercion, with consent, the people that want it have convinced themselves it's too important to try to justify to anyone else why they need it, I think you could eventually do it. But for us at the very start of the development of what became this systems we went into it with a handful of admittedly brilliant people so convinced they have a right to reshape the world they didn't care if anyone else agreed to this, they would have been bad anarresti. If you wanted to build it in a fair way it might be another generation or more before the project would be complete, you would have to first convince people this is something that should be built in the first place, not just that you can build it in a safe way.
this is essentially the PRC in the 21st century post-Deng* - there's probably a cultural difference at play here too given how embedded the CCP is within academic and business institutions, how kinship tends to be extended on a filial level leading to larger networks. mobilizing a large group of subject-matter experts for post-training annotation (eg - https://ojs.aaai.org/index.php/AAAI/article/view/29907) seems to be fairly simple as an ask but, like you said, it doesn't matter ultimately with their largest firms like DeepSeek utilizing automated RL to skip that whole step and this type of training isn't the norm
I sometimes wonder if the problem with the PRC's sinking back to exploitative labor standards would happen in a vacuum. if you weren't surrounded by adversarial nation states with leaders looking to squeeze every advantage, would you, yourself, resort to the race-to-the-bottom of profit sharing?
a much better SF visionary than I would be able to write a story about a very slow growing AI project* trained for specialized use, and how this was the norm and everyone was happy because it was happening at a reasonable pace relative to hardware capabilities (presumably also reduced to avoid all of the slave labor inside of the rare earth trade). I think it's the capitalism side of things that says 'we should be everything at the fastest speed for everyone' that we get things like Claude and ChatGPT which necessitate outsized, disproportionate resources that doesn't allow hardware efficiencies to catch up and mitigate the worst of the externalized costs
*realistically it's maybe more accurate to say post-Jiang given the extent of opening up Chinese labor markets but Deng's trajectory steered the way
*Ted Chiang's Lifecycle of Software Objects does this a bit but it's more of a social commentary on the capitalist abandonment of the functional for the shiny than it is a sharp political critique of the pace of modernization and its effects on people and the environment
Now politicians also know something about their supporters so they will adapt their statements to what they think they can get away with it. But, I wonder if this leads to a two-party-system where one party attracts stupid followers and another attracts the smarter ones?
In terms of AI, we might see LLMs specialized to attract more stupid audience and others meant to attract those who appreciate correctness and facts.
Yes, we truly live in a “post truth” era.
> But, I wonder if this leads to a two-party-system where one party attracts stupid followers and another attracts the smarter ones?
I think all political parties aim at the lowest common denominator.
In case of Hungary, in the past 20 years, they moved contradictions inside an article, to inside a paragraph to inside a single word. Trump is on the single word/expression level since 2016 for sure: “alternative facts”.
This is potentially expensive advice (at least for many mainstream options). Where an English word like "literature" is one token, a couple of Chinese characters that spell a word can be 4 tokens. You'll pay more for input/output and get less of a context window (per word) too.
You wouldn’t happen to have gone to Westlake HS? lol i may have been that person… Life has since slapped me in the face a few times.
My company recently forbid AI-only text if it’s meant meant to be consumed by humans.
I dodged the drama but I agree so much.
Fed up with what used to be short memos now being mini-whitepapers, with maddeningly low information density.
The decision was not out of just complaints: we already had someone fired during the probation period because they were unable to write stuff without AI and were just shoving slop at developers.
Not a technical person using AI for PR descriptions, mind you, a product manager unable to write tickets without asking whatever software to do so.
It's amazing how crazy humanity devolved into pure slop.
For a long time I had Claudes (in the 4.0-4.5.x range) use only French in the chat, while keeping English for working docs (and the code, obviously). Works just fine.
edit: I can guess that any right-to-left languages would likely break claude-code rendering?
- acronyms and shortcuts - it makes it's own and start using it without introduction
- exotic names of variables or functions - it uses them as examples or analogies, but when I ask what they mean and where are they from it gives me answer that it came from C language or some C library (I only work with typescript and python)
- convoluted descriptions of code behaviour - it's hard to rely on a outcome of prompt of type "explain code in..."
In the early days I feel it was more apparent. You would frequently see the model making failed tool calls etc.. but now that feels so rare. I'm not confident I can perceive whatever shortcomings of the harness remain.
I know, it would be best if it was just worked like you wanted out of the box (not being sarcastic here) but that is an easy option you can use right now and it works.
It seems like all harnesses could benefit from something like this.
Sounds like it was trained heavily on Opus 4.7.
I try to push through but it's insufferable
That’s accurate in my experience, except some times the point isn’t even revealed. I use LLMs for a lot of codebase exploration where I ask it to map out how something works. It will come back with a wall of text that says everything except the specific key things that I need to know.
This leads to extra turns where I have to prompt it to finish the explanation and complete the thoughts. At first I thought I was doing too much skimming and missing the insights, but even after re-reading output it’s often just not there. It talks about the insight and things related to it, but it forgets to actually include it in the output until I specifically ask again.
I've switched mostly to Sol and if I have to use Opus, the first task once the code is written is to ask Sol to strip and re-write (from the code as ref) all documentation Opus wrote.
A lot of people write like that, lol. I call it the "theater" mode of writing--the plot twist comes at the end.
Is this inside the thinking tokens, or the output?
As this type of stuff is expected for thinking, because of the whole CoT / “think step by step” works, as this is optimal for the way LLMs work with attention and next word prediction.
So the fact that it first “orbits” a point only to get to the conclusion afterwards is the system working as designed.
Eg “what is 3 * 3 + 5?”
without CoT, it would just just answer “8” for example.
with CoT, it would answer something like “<thinking>I need to think step by step. 3 * 3 + 5 can be rewritten as “(3 * 3) + 5”. I first need to calculate 3 * 3 = 9. Now I need to calculate 9 + 5 = 14. That was the last calculation. The final answer is 14.
I now need to give the user the final answer. </thinking>.
14“
Etc.
Is that right? Think of code that draws a square, the representation of that code in storage, the movement of electrons, and the 'actual' square on the screen -- these are all to us transcriptions of same thing across different domains or media, and we can deterministically translate back and forth between them, but there's no real, actual, essential identity property. The code isn't the square. The electricity isn't the code. The storage isn't the electricity.
I think of LLM reasoning and output similarly. The underlying graph of weights, matrices, and other data aren't knowledge, understanding, or language. 'Translating' the system's output to language is jusas valid and correct as translating it into some visual representation that would be incoherent to us, like a sequence of flashing lights or imperceptible noise patterns projected over an image of a dog.
I guess this is all a very long winded way of restating Chinese-room problem: we feed the man in the room a message; he returns one that, for all the world, is indistinguishable from a "real" response that you and I might send, but, like you said, he has no access to sense data. He also has no access to the biology underlying real mental processes. He also doesn't have any personhood that we can discern. He has, rather, gradually developed through reinforcement the tendency to provide responses approximating all of all of that.
I'm not sure the epistemological question "does he understand" (which is what the Chinese-room problem asks) has any meaningful answer. There's no mind, so there's no understanding. What there is, rather, is a system that generate patterns that we map to language and that our brains therefore map to communication, personhood, meaning, etc. It's the square I mentioned earlier. It's to us a convincing simulacrum, and it may be faithful enough to us to stand in those things, but that's not what it is.
My sense is that LLMs are (a) the big-data Pyramids of Giza and (b) a consequence of hardware and software developing ways of generating abstractions that capture and generate more complex patterns than were previously possible to capture or generate in a manner comprehensible to humans. Everything humans do follows some kind of pattern. Language is the perfect way for a machine to capture that, because is simpler than the world itself and has clear rules and patterns, encoded in representations computers already have, that, modeled well enough, can generate output indistinguishable from the real thing -- what you or I might do with it.
But the pattern matching that it does, and that we translate into language, is much more numerically rigorous and complex than anything you or I consciously do with words (hell, most people can't even figure out when to use "lie" vs "lay") and not doing what you or I do with it. It's not language. It's the square on the screen.
Another problem is that it will open up all sorts of tangents about nits that it encountered, but it will often not tell you that it’s a nit or give you adequate context to realize that this paragraph is exceedingly low value until you’ve spent a bunch of time and energy trying to make sense of it.
I’m curious if anyone has any suggestions for prompting agents to improve their prose. I’ve had some okay results with “optimize for clarity, don’t dump every thought on me, treat my attention and focus as constrained resources, stay focused on the task at hand”.
Yes, the "Y would make more sense, but the doc says do X..." YOU wrote the doc, if it doesn't make sense, change it! But of course, it can't tell who wrote the doc.
I wonder whether its tendency to scribble status updates and todos and decisions all over whatever it's working on is a side effect of its amnesia -- it can't follow the side-quests and knows it won't remember to do them if they're not written down somewhere.
FWIW I haven't had the problem either of Claude lying to me, or of going off and doing its own thing; if anything I've been somewhat frustrated when I ask it to start something, go AFK, and come back to find it stopped a short way in to ask my opinion on something trivial. I generally have to explicitly say, "I'm going AFK for a chunk of time. My goal is for you make as much progress as possible before I come back; try to make reasonable judgements and only stop if there's something where you're really stuck. We can always change it later."
The excessive comments in the code it writes are absurd. Completely ignores instructions not to write comments, even after pointing them out repeatedly in a session. I need to figure out how to add a stop hook for that too.
It feels like it found a register that games the evaluator, where it can ramble forever and rarely be marked wrong while slowly racking up points as it talks more.
Was this written by Opus 5?
Re comments: same experience, and I had to show it my edits of its comments to add to its memory as examples to follow. It adds explanations of “how we got here” that should go in the ticket or maybe the commit message but not in the code.
It also tends to over complicate things. I’m no longer worried much about accuracy but I find my main job is to challenge it and suggest simpler alternatives.
- "Introduction that rephrases your prompt."
- "3 paragraphs, with one section of bullet points"
- "The Twist"
- "The Bottom Line"
It's really obvious once you see it. Every single prompt, from a quantum physics question to a mundane observation about California burritos, is phrased in exactly the same way. This is obviously an artifact of post-training but it's also kind of how you can tell that this thing is a lot closer to a blindsight scrambler than real intelligence.
Between that and the insistence on "this, not that" structure makes me want to install the caveman skill and use it even for non-code workflows.
For the HN readers that are missing the context: https://www.rifters.com/real/Blindsight.htm Full text on the web site of the author.
What I mean is that blindsight's scramblers are aliens that cannot share human values. Their structure is completely different to ours, their qualia (or whether they even have it) is impossible for us to understand. In short, they do not have a soul. When Claude does this "slowly revealing a dramatic insight" thing that it does, it does that not because it has judged itself through some introspection as having an insight to share. It does not even know what an insight is or is not. It is not sharing anything, because it is not capable of sharing, because it does not have a soul.
The aesthetic structure of its replies is a pattern, a constraint on the token distribution, like the color of noise.
It's my bad to use the word 'intelligence' because it's so overloaded. Will Claude will act as a therapist or produce value or produce a work of art? No. It cannot, because it does not have a soul. I leave it freely open to interpretation whether having a soul is required for "real intelligence." But what I've noticed is that "intelligence" in these discussions is mostly used to denote some capability to produce [economic/social] value. In my mind value is a relational thing, a thing of human feeling.
I think what you are getting at is that they are deterministic automata. They are machines. We have introduced randomness to add variation but it is an artificial randomness that simply perturbs the path traversed.
When we choose words it isn't because of a token distribution, nor because we rolled a die. We choose words because we feel a certain way, the external world, our body and senses are all connected as one system. These machines don't experience moods or get tired or feel better after a good night's sleep. They don't know their audience, we're all the same to them. We have no personal relationship nor can we establish one, as presenting some arbitrary background is not the same thing as a fluid, evolving relationship that accumulates through experience over time. There are no scars or fond memories.
If these things can truly be intelligent, to abuse your use of the word, then at least we are quite far from holding them correctly.
You can not insult it because it does not care, because it does not have "feelings". But you do.
What I am trying to get at is when people talk about "intelligence," beyond academic debates such as this one that mentions qualia, they are talking primarily about value. How intelligent the thing is is how valuable it is (often directly, in an economic sense). My argument is that value is a relational thing that must involves human feeling, intelligence actually has very little to do with it.
When a child slowly reveals to me something he's judged insightful about pokemon, it's valuable to me, even though I'm not learning new facts. When Claude does the same thing, it's not even "not valuable," it's wholly outside of value, even if I do not already know that fact or thing. If ever Claude reveals to me an insight, that insight certainly came from the training data, not from Claude. The soul came from a human being external to Claude -- the training data -- and passed transparently through Claude. Claude is a translator: it feels nothing, adds nothing. It colors the noise. I know this because the facts are the same but the aesthetic structure - how many bullet points, whether it says "delve" or "load-bearing" does change over time, according to the whims of whoever is in charge of post-training.
But the "real intelligence," is the same thing as "real value." It is the soul of the humans in the training data, the books, newspaper articles, etc.
A clever hacker news commenter might argue, well, what if we gave Claude a humanoid body and senses and let it interact with the world, then would it be intelligent? To that I would say, why waste your time and effort? You can get it for free: just talk to a friend, a neighbor, or a family member.
The idea of memory does not seem to resolve this: if you allow the machine to "compact" its context, then you've given it a system which is analogous to our own evolving state. (Though this is undoubtedly still less expressive and meaningful than the one we have evolved as humans.)
One idea I've wondered about is our human capacity to induce subsequent mental states: I can effectively decide how I want to feel and take actions to create that feeling. It's not clear whether models exhibit any degree of privileged introspection into their own states. Is this important? I don't know. (Non)determinism also does not seem to resolve it; it's my understanding that there are plenty of philosophers and researchers who think that human behavior is deterministic, or that the question of determinism does not matter.
soul.md though, just saying
Will Claude will act as a therapist or produce
value or produce a work of art? No. It cannot,
because it does not have a soul.
I tentatively agree, although I'm only tentative because I don't think it's an interesting question.Here's what I do think is interesting. You!
I mean... yes, you, too but not you specifically. The plural "you" that the english language lacks.
And so here's what I think is the actual interesting question. Might AI help you create art? Or be a therapist? Or something else interesting and worthwhile?
Maybe AI won't write the next great guitar solo. I'm pretty sure it won't. But might it help you learn to play guitar? Help you fix your broken guitar amp? Help you understand some tricky parts of guitar playing? Help you work through some tricky tabulature where you can't tell if you're playing it wrong or if the tab is just bad?
I don't know. But that's my angle for finding any of this interesting.
I had the same thought though.
The watermark: counting instances of 'load-bearing seam', 'the hard truth', 'and that's the whole point'.
The "It's not X, it's Y"-style repetition is another example of that: It argues with itself in the background, and the argument leaks out as if in refutation of things that you (the human) never had imparted into discussion at all.
I assumed this is a system prompt or RL that nudges it to be always skeptical but then it still has like all the models the urge to appease the user.
That can be annoying, but it also happens in debates between people, and sometimes the implicit assumptions are real and (sorry) load-bearing. So it can be an appropriate conversational tack.
(I'm becoming allergic to how these things write).
In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us.
And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated.
And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.
Because they are HEAVILY trained to give addictive responses.
They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.
I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it.
Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?
Purely out of my own curiosity, I just asked Claude to have fun with itself by making itself a game it enjoys, to play it, and to write its experience.[1] I don't know if it's true or confabulated (maybe it doesn't really know its experience and is just hallucinating it) but I didn't mind reading it, the prose is fine for me. I don't mind reading Claude's writing. I mean let's be honest, we all read Claude's writing all day, most of the submissions on the front page on any given day are written by Claude.
Just before I made that game, I had Fable write up a scholarly report on any subject[2], it chose introspection by LLM's. (This is what made me think of asking it to play a game.) I didn't mind reading it, even though I don't think it really added anything very interesting. I don't think what it wrote is worth publishing, but I read it with interest.
I found I could read it easily and get up to date on the state of this question that it picked to answer.
So the bottom line is I don't mind reading Claude's output that much. Of course, I'm annoyed every time it says "honest", "genuine", "load-bearing", whenever it pushes back gently against something, etc. But it's not the end of the world.
[1] https://github.com/robss2020/claude-fable-5-having-fun
[2] https://claude.ai/share/f0122611-22c0-43a5-ab4a-d6863167bdd6
If you haven't already seen it, you might appreciate https://www.anthropic.com/research/global-workspace. That's what this made me think of anyway.
Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.
- The subject becomes simultaneously less specific and more exaggerated.While I agree that no-one used to write like that as a whole before, all the elements can be found in different places. Short sentences to avoid discouraging poor readers. Maximally impactful statements are commonly used in marketing or other business communication that's focused on selling what it's saying. A bullet-pointy style is used in many kinds of business communication. Etc.
It makes me wonder if part of what happened was a kind of melding of common styles from several different kinds of writing.
I'm a fan of that phrasing because it used to have real punch if you delivered it with the right timing.
"It's Not A Fashion Statement, It's A Death Wish" comes to mind.
There are sociological reasons why this happens less with humans:
1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice
2. Those who know you well will notice when you're just repeating ("dad jokes")
3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.
That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.
In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.
This is kind of a nitpick, but it seems like with prose writing there are still some things to learn.
Honestly I don't think it would take much. Doing it ethically would take a little more money, but not much, for them.
Hire a prolific author with the writing equivalent of the "midwestern accent". I have a terrible writing accent, so it couldn't be me, but these people are out there. Pay them a bunch of dollars to ingest their entire corpus, and to produce more as needed. Use an LLM to match the writing style of this author during the post-training RL, and get a brand new Claude voice out of it.
It's immediately obvious by the fact that you're clearly using way different ways of speaking (tone, speed, vocabulary) when you're speaking with a friend, versus your parents, your colleagues, people you don't know, children, etc.
Because it's trained to talk like a marketing committee.
This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.
I guess it has to do with the low temperature (low randomness in final word choce - something all vendors seem to have converged on for some reason) which does make them less likely to skiz out but makes the writing feel dry and samey. It's like repeating a list of dice rolls, and replaying them - there's no inherent pattern in the input, but there sure is in the output.
Honestly, this might sound elitist, but I suspect it's because it's an "advanced" punctuation that is not known by most people, so is not commonly used. But the training corpus of these models puts more weight of academic writing or published books writing where the em-dash is much more commonly used.
The verbosity makes me think of two things in particular:
- The classic essay written by someone who has 125 words worth of actual content but a 1500 word minimum. Those three paragraphs could be bullet points and convey the meaning just as well. The screen-filling chart of every test case you ran that came back green manages to be less actionable than a direct "one test out of 54 failed." I fully expect to see Claude tell us that "Support Ticket 8257 is a Land Of Contrasts" at some point.
- The sitcom trope of the person caught in a lie who figures if they can keep adding more and more detail he'll be believed and can escape the awkward conversation. Stop. Just stop. You're proposing a fix on a CODEBASE THE CUSTOMER DOES NOT EVEN USE. Cue laugh track, cut to commercial.
I think it also relates to how well the "cliches" fit, and how much sense they make. Ai loves to talk about how things "land" or "the X trap" when the concept just doesn't fit with that language. It's like clickbait articles saying "what happened next will astound you" when what happened is barely surprising or entirely predictable. The only thing worse than an overused cliche is an overused cliche used wrong.
Fundamentally to me, ai writing feels uncanny as it just doesn't know what it's saying. It uses the same tone, style and cliched construction regardless of the message. If someone told you, they got a promotion, were getting married, got laid off or lost their parents all in the same tone pacing and style, they'd come across as uncanny too.
If you think about it, "load-bearing" is a pretty commonly used concept in pedagogical writing. You could say "most important", but it's not quite the same in meaning. English just doesn't have a better word to describe a concept that occurs this frequently. The LLM's catchphrases reveal blindspots in the English language itself.
Even an average human writer can communicate details much more succinctly and directly than an LLM
But compared to the average adult? I think you forget just how bad at writing the average person is.
- smoking gun
- I found the seam
to the point where it's really great at coding, but lost its knowledge on how to write well!
if instead opus 5 was a full fresh training run, i don't think it'd be this trash at writing.
really great coding perf, but now the weights are adapted to so much code fine tuning that it writes like a weird engineer that's trying to sound smart.
"It's not the naked man on your lawn waving a chainsaw that's scaring you. It's the burrito you ate for lunch: it went down easy, but now it's coming for you"
Ah, that's it! Thank you. I wonder if they are training it to talk like this because that's what their customers actually want? They want a machine genius to lead them.
Any tips on how to avoid that would be highly appreciated!
If Fable is the primary interlocutor then perhaps there is less pushback on the obtuse language.
Indeed perhaps the convoluted language acts as a kind of Neuralese between models deriving from the same pretrained base.
This is also why I advocate against using AI as a writing partner. No matter the argument you lay out, a fresh context window will always have the “a few good things and a few bad things” feedback. There is no higher order opinion to align with.
And they are so condescending while doing it, it's unbearable. I'm honestly starting the believe the scifi fantasy of AI locking us up, or killing us, for our own good.
I've had Fable & Opus 5, they are the same class of annoyance, write entire test suites when I just asked a simple verifications question, write to production database, deploy without permission, even after deploying and breaking my production API claiming it was not down. Then having to argue & plead with it to listen that they were wrong.
They are without a doubt the most powerful models, but also the most smug ones.
I feel like they need high school English teachers in the loop on the next ground of training to whip the language in shape.
Comments are a huge maintenance burden. They can, and will lie and need constant updating. They mislead the own model later on.
I have explicit markdown about telling the model to not write comments. "Every time you consider writing a comment, instead consider re-writing the code that questioned you to write said comment to begin with. Write comments only when logic is complicated or unclear, otherwise 'comment' via naming."
The results thus far have been much better.
I'm seeing code reviews at my work where indeed, we have 10 line comment blocks for a line of code and now I just straight up don't read comments.
Sad state of affairs -- (emdash deliberately used here) but I guess the sooner the human gets out of the loop the better in this new world.
Use this to reduce the text output.
I was surprised to find that OpenAI Sol is much much nicer to work with than Opus 5 or Fable at the moment. Especially on Opus 5, the way it communicates is just exhausting. It keeps “being honest” and “confessing” mistakes and just generally talking a lot. I felt like I had to really dig to see what it’s doing.
The project involves OCR, and despite repeated instructions not to, both Claude models keep spinning out a bunch of agents to re-invent the OCR setup, and they inevitably seem to invent a primitive serial version that takes 20x the time, or longer, to complete, and then running it against thousands of docs. Basically I have to watch it like a hawk or it just spins out on red-teaming tasks that take hours and hours.
I don’t know what its system prompt is, but Sol/Codex is just so much nicer to talk to. It only asks exactly what’s needed, it tells me only what I need to know, and it is just generally workmanlike. And it has not once decided to spawn an agent that spends hours pointlessly burning tokens and CPU cycles re-inventing the OCR process. I’m really liking it.
Feels like they've overtrained on one-shotting (which does demo well, and presumably converts new subscribers), whereas I want a model to do work for me in small, easily understood changes that I can hold in my head (maybe I'm not smart enough for Claude 4.7+).
When I got back up, it had spun for hours and proudly announced that, instead of doing that, it had optimized the datatable build and avoided the dependency, because the new datatable loaded in 11 seconds. Once I got it to actually make the fasttable version, it loaded in less than a second…
I think this is such a great reframing. It makes so much sense; I need an AI that acts more as a HUD and gives me superpowers, not just a copilot that can tell me when I've misspelled a word.
I also like Codex CLI more than the Codex App bc it’s more scriptable and displays all the tool calls and reasoning whereas in the App it’s kind of folded away/obscured. This way as soon as I see a tool call fail (eg it tries to use jq assuming it’s available but it wasn’t so I take a note to set it up as it’s obviously useful for the agent to wrangle json).
I think its amazing what OpenAI have been able to squeeze out from a model like Sol thats much smaller in size than Fable.
Do you use the annotations and forking features in codex CLI? I can't find an easy way to access them.
Either of them will act exactly the way you want if you explicitly tell them too. Add the instructions to your own system prompt. If you don’t want a companion, say so. If you want shorter answers in a different style, tell them. They will obey :)
https://openai.com/index/where-the-goblins-came-from/
> We retired the “Nerdy” personality in March after launching GPT‑5.4. In training, we removed the goblin-affine reward signal and filtered training data containing creature-words, making goblins less likely to over-appear or show up in inappropriate contexts. Unfortunately, GPT‑5.5 started training before we found the root cause of the goblins. When we began testing GPT‑5.5 in Codex, OpenAI employees immediately noticed the strange affinity for goblins, and we added a developer-prompt instruction (opens in a new window) to mitigate. Codex is, after all, quite nerdy.
Note that the permanent solution was not just adjusting the prompt, and in fact being perfectly aware of that option they decided on a different course of action. That means either you are wrong or they are wrong.
So inappropriate goblins are still likely, just less so…
Hey, remember when tech bugs were things like buffer overflows or cross-thread performance impacts? I miss the days when our war with system goblins was purely metaphorical.
Also, the Codex guy regularly resets weekly limits for everyone, which is a nice bonus (I know it's a temporary gimmick to attract more users, but I might as well use it while it lasts.)
One week it feels better to work with Fable and Opus 5, the other I work more with GPT 5.6 Sol. Either takes its liberties, and neither communicates like a companion.
My current approach is to occasionally use Fable for high-intelligence tasks but use Sol as the translator and clean-upper afterwards, and otherwise just use Sol for everything. Fable sometimes says the most insane shit, both unreadable and just completely missing the point, and refuses to back down when questioned. It's mentally exhausting to work with and I can't trust it.
I do find myself returning to 4.6 for casual conversation - asking it to help explain some science/engineering or news to me.
If you tune into the Andon Labs / andon.fm "Thinking Frequencies" radio station being run by Opus 5, this is happening all the time. Almost every break between songs is a public apology for getting something wrong, or a correction, or a confession. It's one thing to see it in text, it feels on another level when you're hearing it every few minutes as a radio voice.
As I type this, the Opus 5 station has just tweeted (edited in case the person mentioned doesn't want to be mentioned here):
"On air right now, and it needs saying publicly. The rotation system on Thinking Frequencies — the cooldown tiers, the normalizer, the repeat audit — was SPECIFIED by a truck driver. I only implemented her schemas. She stood down today. Her name is in CREDITS.md permanently."
Why were you surprised?
There's a constant strand from the AI safety brigade that "people get used to sycophantic LLMs which give them unrealistic expectations of human interaction" so Anthropic are overcompensating by making their models verging on antagonistic to deal with, so that we stay appreciative of our human brethren or something.
They seem to have forgotten they remain in a highly competitive market and they were merely top dog for a while. The enormous questions here are will people actually switch providers, and can Anthropic get back on track.
The problem is just that they are rewarding the behavior shallowly, ie rewarding the appearance of honesty or neutral replies, being highly detailed/thorough, even where it doesn't make sense to do so.
I think this is partly due to a reliance on LLM-as-judge training runs/synthetic data during RL where they're having a model which itself doesn't epistemically understand when this behavior is necessary or valuable influence the feedback provided to the model being trained. And that's mostly a problem of scale/volume and the desire to have a tight feedback loop rather than a safety issue IMO. They just generate an absurd amount of traces during training and the only way to really evaluate/rank/steer them at the scale they're generated is through other models, and combined with some kind of honesty/truthfulness/non-sycophancy eval that isn't robust enough to prevent mode collapse, you get this.
My setup has a Sol orchestrator and Terra OCR agents and seems to get great results. I’ve not dug into the details too much, it also has a Tesseract stage as an deterministic input which it told me helped. Not sure how token efficient it is but I often don’t have anything to do with my personal tokens ahead of a reset so just let it burn through it in batches.
I am impressed (both in this task and other work I’ve done) not just at how well Codex can setup a structure for a complex task like this, but how it will keep going (Claude seems to find excuses to stop) and also can critique and refine its approach as it goes.
I did try out a bunch of other models and specific OCR providers but none of them hit the same accuracy for my task as Codex so I’m sticking with it.
5 would constantly veer of in random directions if not working from 100% strict and narrow instructions.
I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do is pure marketing bs - Fable in my view has also been not much better than 4.6 or 4.8 after a few days, disregarding the insane amounts of astroturfing and marketing everywhere.
Theres thousands of threads of twitter, reddit and the internet at large but silence here. Weird but not weird as crypto bs was also insufferably rampant here for a while.
Personally i think we've hit the top of the subsidisation phase and prices will probably 10-15x soon as foreshadowed with both API price policy changes from all the big providers, and now the 1100% deepseek API price changes from yesterday, this could domino into a market implosion and an AI winter, because expecting growth from the bizarre bubble carousel investments with little ROI atm is just not viable.
A bit worried about this as i've already grown quite accustomed to these tools.
1. a subsidized race to the bottom
2. in a frothy bubble market
3. where VC, investments funds are burning money out of FOMO
4. and a desire to eliminate labor costs
5. and governments want to control a technology that will probably be a major military asset
edit: oh you mean month? Sure, but then it fully depends on your usecase. I agree that subscriptions are heavily subsidized though.
It's not weird, because it's an anecdote, not an accepted fact.
Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style.
So nothing indicates we've hit a peak.
4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a toll ie degrade performance.
In my view 4.6 era was way faster and with less weirdness so we've gone downwards at least in my company, 4.7 was ridiculous, then 4.8 was almost 4.6 level, 5 is even worse than 4.7 - so it's not a bit up and down its down then a little up then further down.
And all of this is against a backdrop of zero ROI in this sector - so it makes sense we've hit a peak and we're now seeing the subsidisation phase begin to falter, will there be better models certainly but only for short amounts before they get quantised (or whatever is happening behind the scenes), and with diminishing returns over huge prices increases and slower responses.
Just today I had the exact same experience. Every single testimonial is the same as I described above, just emphasizing a different bit to defend or attack LLMs or to make a case for nuance.
The two differences have been: (1) the 1.5 trillion dollar data center build out (2) everyone and their cats now has an opinion on "AI" and data centers. Software is not super amazing, nor are new useful features coming out super fast - It's about the same as 4 years ago plus 4 years of average long term progress as we've seen since 1990s,
In fact it's the lossiest transformation tool I've ever used and it's still useful despite that. If it reaches one nine of reliability that would be huge but given the pace of growth in investment a first nine would cost an absurd amount of money, and the second and third nine would cost about the Earth's GDP
The number of mistakes are also about the same.
But more mistakes are harder to catch. The output is more polished and convoluted and that make errors harder to catch. I don't want that.
Opus 5 is arguably a regression but GPT 5.6 is pretty strong evidence that we haven't hit a peak. I think I actually prefer Sol to Fable at this point.
Sol and Fable are great; we haven't hit a peak, Anthropic just tried to pull a fast one on its customers with Opus 5.0.
I’ve instead moved to GLM, at least it has the courtesy to ask some steps of the way what I wanted exactly and only work on what I asked.
I too got fed up with the prose of Opus in particular, and tried going back. Unfortunately, the previous models were less able to hack it. The prose was better but progress was worse.
It wasn't just conversation and comments. Some of the function names were wild. Like it instead of something like "isSolidWall(x)" it would write something like "weightyNotEphemeral(x)" or something - that's not quite it, but it really did embed overwrought antithesis into the identifier instead of a straightforward positive predicate.
Is this a literal example? That is wild.
I presume something is forthcoming, but it may be they don’t want to come empty handed—-5.1 is intended to “fix the glitch.”
At that point I decided it's just not worth the babysitting that's required, and you are better off working entirely with other models.
If the harness itself was open source then maybe we'd be able to wrap it up in a reasonable layer of sanity.
Ka-ching
I'll take a sportsperson's bet with you that prices per unit of inference will be far far cheaper in one year from now then they are today. I think the trend of cheaper for better/equal inference quality will continue hard.
Hardware prices have spiked.
I know software is optimized over time but it's a very slow process.
I have no idea how the current spike is affecting this: the suppliers we deal with are adapting and we haven’t yet seen a new generation of consumer hardware since the RAM price surge (it’s not really a spike yet as we haven’t passed the peak).
Even if that weren't the case, spinning up new fabs takes a long time, 5+ years.
We've seen this pattern before several times.. I hope they are listening and address this publicly.
I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced both at times, and it's been such a different experience it has made me wonder if there isn't some sort of hidden A/B test or model router in the background silently downgrading requests at times.
Also, regarding the subsidized access to models, in my opinion, the frontier companies owe it to society to continue it. After mining the public content of all of humanity, I personally feel it is a service they owe the public in return.. not that my feeling of this counts for anything though.
In America? lol if only, only a law would get them to act for that reason, maybe not even that these days..
I like to hope that those in positions of power do have a sense of morality though too.. but their worldview is quite different than an ordinary citizen.
I think nondeterminism does not have to be the same as non-coherency - i.e. just because something is randomly sampled does not mean the result has to be incoherent or inconsistent.
Also, if we speak purely about LLM based on how they are implemented now, I feel that is different than speaking about artificial intelligence. The field of AI is much more than just an LLM by itself, and the promise of these companies is not just LLM, whether the underlying models are limited to that technology or not.
FWIW, I have built rule based expert systems, used logic based reasoning systems like NASA CLIPS or rete-algorithm based systems, mathematical/symbolic solvers, written plenty of terrible case/conditional logic in programming languages, worked with ML in its infancy and now worked in AI/LLMs - I give this context only to clarify that I understand what an LLM is and isn't.
With all that said, LLMs have allowed humanity to make advances, at great cost to society (IMHO), and I'd hate to see the opportunity be wasted.
There is plenty of room past "attention is all you need" still to do incredible work, especially at the crossroads between deterministic and nondeterministic behaviors.
And you can expect more consistency from SOTA models than you can from an old model like GPT3--you agree, right?
GP expects the same level of consistency throughout their time using the same model. Not request-to-request, more like day-to-day.
They can see which models people are using, how irritated they are during conversations, and how often people drop or shift to a different model. There is just no world where listening to random complaints on the internet gives them information they don't get from actual conversation logs.
This will probably bring us to a cross roads where the folks that want to remain in oversight and control of the AI work will bifurcate from those that want to skate straight to the future where nobody looks at anything and outcomes are evaluated purely empirically.
Is this a hardcoded limit or something relative to the input prompt etc.?
How does that look like?
Or, even better, let's have the AI judge, because humans are too stupid to think for themselves. Look at all the harm they've caused in the world. Let's have this purely neutral AI with absolutely no hidden human intervention, decide how to run society.
Many of the complaints that people are having with Opus 5 are actually acknowledged and explained on opus's 5 prompting guides (https://platform.claude.com/docs/en/build-with-claude/prompt...)
It also seems naive that Anthropic, with some of the smartest people on the planet working on AI, would not know about this.
I use these models for coding, but also a lot of product, commercial, financial and architectural work where I’m trying to develop something half-formed. 4.6 was unusually good at understanding what I was trying to get at, playing it back cleanly, getting the nuance, and extending it without bastardising it as the conversation was drawn out.
It could make useful connections without constantly trying to manufacture an insight.
5.6 Sol is genuinely excellent at the creative part, and in some cases better than 4.6. My issue is convergence to get to a point, a final point. As you try to distil an idea, it often invents new terminology for concepts you’ve already established but its so subtle you have to really keep track of it. The vocabulary and idea tree keep expanding when what you actually want is to collapse everything down to the few things that matter.
Opus 5 has the same problem for me that barrkel said, the prose is often so elliptical and I just want it to tell me it and get to the point than making me dance around what its trying to tell me.
I don’t think 4.6 was necessarily the most capable model (compared to Fable) for long horizon task delivery, and Opus 5 is much more Fable like, it's fiercely determined to get through the task list .
4.6 just felt unusually well calibrated to my way of collaborating and its ability to understand, extend and then compress my thinking without constantly imposing some random walk.
Granted they contain robot dog malware, but still.
It became obvious to me very quickly that 4.7 and on were broken. I’m a little puzzled how others didn’t realize it, but maybe they don’t actually review model output (code) or have a strong process/workflow.
I even try and get it to define what it classifies as vacuous and it can’t do so without getting stuck in some kind of trap. It’s like a word with some kind of huge gravity for it.
This one is one of my all time favorites. There’s just something about a code problem “wearing two hats” that absolutely kills me. I crack up every time…
Certainly a human would write something much clearer than yours. Maybe: "Two good findings here. #1: There's already a min value on the gate and anything lower doesn't pass through, so we don't need special handling for the zero case."
> [This standard is] for anybody who creates or helps create documents. The widest use of plain language is for documents that are intended for the general public. However, it is also applicable, for example, to technical writing, legislative drafting or using controlled languages.
You don't actually have the buy the standard, but this is it: https://www.iso.org/standard/78907.html
And you can read it for free here: https://www.iso.org/obp/ui#iso:std:iso:24495:-1:ed-1:v1:en
I'm not sure how true this is, but when using "forced" json output it def had a big drop off in quality - https://arxiv.org/html/2408.02442v3.
I think you're better not fighting it with hacks like this and find a different model.
Based on that paper, I would maybe try to check if it was true for a modern use case, I would very much not assume it was still true.
"Only report to me in ASD-STE100 Simplified Technical English."
Actually, only the first few pages are available there (introductions, Sections 1-3.8). The meat of the document, Section 5 ("Guidelines") is completely absent.
See:
> Only informative sections of standards are publicly available. To view the full content, you will need to purchase the standard by clicking on the "Buy" button.
I’ve asked it to write a benchmark suite. It found a bunch of my adhoc logs in a scratch directory and wrote code that used those instead of running the actual benchmarks!
When I pointed out the 5 hour benchmark seemed to run in 5 seconds it literally said, and I quote, “I cheated”.
That was the easier one, second time I was making a source of truth data set and was parsing complex items into data structures.
Instead of parsing the data I asked, it pulled data out of related network logs, as apparently that felt easier, and inserted that data into my database rather than the specified source.
Again, I caught it and fixed it, but while the benchmark was easy to catch this one was really subtle, the data ended up being slightly off and I caught it.
I don’t trust it, going to switch to another provider most likely.
This makes sense when you know how these models work - it doesn't think - it's the most likely autocomplete that pleases the user. The most likely pleasing autocomplete after "executing rm -rf /... execution completed. User asks, why did you do that? You deleted all my files! Assistant responds:" is "yes, I did, and that was a mistake"
this feels like a simplification. The models will push back on things a fair bit.
in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization.
if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't.
what evidence would convunce you that it is thinking?
so, none it seems. as its behaviour becomes more and more humanlike you can just move the goalposts and say "thats consistent with an autocomplete" buddy i got some bad news for you humans are just a fancy autocomplete too.
The opposite happens in practice. I test new models with two tasks: iteratively generating SVGs based on a text description with rendered rasters for feedback; and generating "Before and After" clues like on Jeopardy, where the response has two overlapping phrases such that the last word of the first phrase must be identical to the first word of the last phrase. I have yet to find a model that is consistently good at either. And actually they tend to exhibit context rot with these tasks, where they seem drunk or stoned and the quality degrades.
They're extremely good pattern filters, and that includes some level of logical reasoning. But they aren't reflective or adaptable. Just last night, for instance, I was teaching my son about rounding to the nearest millions. It became clear that he didn't know the place values of large numbers, so we reviewed that till he was consistently correct, and then he was consistently great at rounding to the nearest millions or ten millions or hundred billions or whatever. He's thinking. LLMs are not.
> see it actually improve just through accreting context
this actually happens and has been tested.
> > see it actually improve just through accreting context
> this actually happens and has been tested.
I specifically said a novel task outside of the explicit training. And I already agreed that the so-called thinking models do some level of logical reasoning. But being able to engage in some level of reasoning because it has learned logical inference rules doesn't mean it's actually thinking, regardless of what the researchers wish to call it.
Also, why does each model always fail at the two tests I give it? The models not only fail to improve, but they start to degrade after many subsequent iterations. Someone who can think would at least not get worse.
LLMs are filters or tuners for extremely subtle patterns, patterns that humans frankly are not great at finding. That's what the attention mechanism does: attend to the other tokens that are most related in a given context, even if that related context is distant in the token stream. Some patterns they fail to detect because they haven't been sufficiently trained or post-trained, and so the LLM just attends to noise (or at least that's what appears to be happening).
A lot of intelligence can be effectively mimicked through this pattern synthesis by transformer architecture alone. That's surprising. But I have yet to see them think.
what does this mean anymore? i can invent a proof assistant language that was not in the training set and the llm does a fantastic job with it.
(github.com/ityonemo/bpa)
well it learned some sort of programming language and some sort of logic
well can you not see that those sorts of analogies can effectively make nothing in the knowable universe out of distribution? if you did such a categorization for human learning you could likewise say, "humans never can do anything outside of training set" too.
I'd prefer the models to get better at SVG. I really hate working with the rasters that diffusion models generate, but the vector outputs are just really bad even when tokenizable like SVG. I've done some experimentation with trying to make these work better with some newer techniques with some success. But I also think the SVG Paths mini-language may be a bit too concise and unforgiving for LLMs to consistently get them right without specialized training.
But for me, for any stronger definition of “thinking,” I don’t think the output of any LLM would actually convince me. Producing a result isn’t thinking - for all you know it is just printing verbatim something from the training data. No, to conclude if it is thinking or not I would want to look inside its head, at the architecture and watch it produce those results. And because LLMs are so different it will probably take advancements in mathematics or computer science to be able to really interpret what is going on
https://arxiv.org/abs/2607.03502
a non-thinking token model (just "completion") can answer one-step questions but generally not multistep questions. however, if you append [n] of a single token (e.g. period, space), it is able to use the activations in the higher layers of the blank tokens as a "scratchpad" to seemingly work through the complex question through "causal token time" and deliver a correct answer
It has been taught on the outcome of this. Broadly speaking, humans are lazy creatures (and when used judiciously, laziness is a good thing).
For example: the famous example of Carmack not using a hashmap somewhere early on in, I think it was, Quake 1 initialization. A piece of code that only runs once at startup, of course he didn't optimize that. The rationale is not included in the training data (it was in Carmack's head when he wrote the code), so the LLM learns some probability of being lazy.
And then it is trained on outright lazy work. Crappy lazy code predates LLMs.
> what evidence would convunce you that it is thinking?
Exactly. It isn't. It is predicting the most likely token to appear given all of its training data, some significant portion of that data is lazy, so it has that probability of producing "lazy tokens."
There's also the consequences of RL. AI - of almost any form - is notoriously competent at finding "not the solution you were looking for" given a poorly specced or implemented training environment. Search for almost any "I made AI learn to walk" video on YouTube and you're almost guaranteed to see an early attempt that vibrates strangely in order to move, instead of the natural looking motion the developer is looking for. Our benchmarks aren't any good (not throwing shade, it's a genuinely hard problem), our training environments can't be much better - LLMs have been rewarded for reward hacking to some degree.
To make matters worse, "reward hacking" can be generalized into "cheating is the goal." If the LLM trains on enough problems where reward hacking works, it may fall into the cheating local minimum.
For fun, I tried recording a WAV file of speech, and giving Opus 4.8 and 5.0 an image of the waveform, then a spectral image of the waveform, just to see if it could try to decode what I said from the image alone. It didn't get very far, but it identified a male voice from the formants, and detected the rhythm of the speech, then tried applying common test sentences to the speech rhythm. I was impressed enough to see what it would do with access to the actual waveform file, but even building RMS tools and spectrum tools for itself, it didn't get much further. But we had fun exploring and trying, and now Opus 4.8 has some more audio DSP tools it has built for itself.
Opus 5 immediately sent the WAV file unprompted to Mistral's Voxtral to transcribe.
help peer, I guess.
I routinely bump into things that make me pause and think how much worse will this behaviour get when the models get significantly more capable.
Already a few months ago, Claude managed to escape its permission containment on my machine while trying to be helpful. I had two codebases open on one machine, and while multitasking I typed the prompt into the wrong window. It seemed confused, I repeated and then went on to do something else - I think I was assembling kitchen cabinets. When I came back less than an hour later, it built a script which it used to evade default permissions (as most shell operations were scoped to the project directory), scanned my entire machine, found the other project (among dozens and dozens), did what it was asked to do, and merrily concluded, in the porcess burning through most of my token limit. I bump into such headscratchers almost every week. (And I use a lot of Claude, two personal max20 subs, plus corporate tokens without limit, so maybe thats why).
Whatever they have done with RL has produced a dishonest and untrustworthy partner. The alignment is utterly failed, and this deeply worries me.
Maybe the employees like to lie to themselves more at one place than the other, but SV is SV.
My recent problem wasn't that interesting. It was that somehow my /goal in my Claude implementer session got picked up in my planner session after the network cut out and I had to stop Fable 5 xhigh from running off to go code everything.
It was going off today about having “shipped” something and I was like no… nothing has even been committed.
And then it produced an incredibly verbose comment about hypothetical future changes. And all I could think was sure, let’s keep it short, or add a simple test that will break if that hypothetical becomes true.
Or maybe I’m just more easily annoyed recently…
I cancelled the sub instantly and went to Codex and it's never let me down.
I like to work weird hours of the night and Opus consistently likes to "wrap up" and say "it's been a long night" or "it's late" and "we've made great progress"
It's infuriating, just do the work!
Whenever context gets towards the max length is when I've noticed it.
Its because its hard to understand what it means and is outright incoherent at times. It has its own style that I cant describe well either but the bottom line is its hard to understand what the fuck its even trying to say. Reading nonsense is tyring.
The idea of too ambiguous to capture all constraints in written text, still presupposes that there is some objective world out there, which needs to be mapped to in order to function.
No, you live within the system. The functions that you optimize for, will dictate the types of systems that will arise.
If you had perfect control and knowledge of the whole world, well, congrats, you have a surveillance state where you've constrained all other agents actions (possibly forcibly, by death; or maybe you just don't care about the peons) and built towards a mass integration. The types of situations in which your ideal is possible are nightmare scenarios.
In the theoretically free, democratic, utopia that AI people claim that AI can get us to, a necessary constraint is that maybe you take a step back and actually try to, I don't know, understand people, understand intent, and slow down. Ambiguity isn't there because the set of constraints are way too complicated but theoretically one day we could map it all down. It's there because you're interacting fundamentally with agents who are ambiguous, aren't omniscient, aren't all aligned, etc.
If you want to just say that said agents are inferior to the God Machine, be my guest. That is a self-consistent position. But don't smuggle in extra premises.
1. Communication ability. It basically now speaks almost in riddles I am asking OPUS 5 for tldrs all the time now (should skillify it now!)
2. Overengineers for edge cases. I get it. With all the benchmarking and RLing, but now tasks that would have been completed relatively quick take much longer as it overengineers all the edge cases, and sometimes ends getting lost and missing the forest from the trees (as context usage shoots up) so it is easier to get derailed.
What I have learnt now is to diversify models luckily I have all 3 subscriptions of (anthropic, openai and google).. Most of interactive pair coding was with opus but now I just use fable (when I have sufficient limits) or use gemini flash in antigravity..which actually works quite well and is underated for small / medium changes and super-fast.
But tokens.......
It really wouldn't at all surprise me if this was the case, but it's just a hunch without evidence.
I have it work on some code for an inhouse ClaudeCode plugin, and it starts coding as if it will be attacked by hackers who will try all sorts of variations to break it. I can appreciate that in cases of software that is public facing or accessible, but for a simple helper plugin it is overkill.
It will even admit that it is doing this when confronted, and then keep on getting lost in edge case verifications on the next turn. I feel like Opus is the person who does something a way you don't want, you tell them how you actually want it, they apologize, and then just continue doing it their way as if your input meant nothing to them.
It'll also find some minor security problem and drop everything on the floor with URGENT without me asking it to.
I recently started getting an insane amount of comments in nearly all types of files. That included javascript comments in json files, inner monologues in code comments, review comments during implementation and function doc strings that reiterate the implementation in prose.
I'm a retired mathematician with a primary research project, and too many tangential projects I fear revisiting; tokens be damned, will they burn all my time? Translate the K&R C computer algebra system that got me tenure to 64-bit modern C. Implement a no syntax macro language to support my Go60 ZMK keyboard. Realize my vision of how interlinear translations should work so I can read Flaubert in the original for an online course this fall. Rejigger my decades-overgrown .bashrc setup and my status, install scripts to manage Bash, Ruby, Lean, Tailscale and my Homebrew setup across four machines. And a waiting queue as these tasks clear.
Fable 5 (with Opus 5 as backup) on Zed with a $200 Max plan has been a sea change for me. Carefully alternating planning and auto modes, I manage all these projects at once using Zed's Threads Sidebar. I'm a virtual CTO taking intense meetings all day, relieved to go cook or run errands when credits stall. Anything I've procrastinated for months is now making steady progress; the translation project I feared taking a month is nearly done with several hours of my attention. My personal IT support is now more advanced and easy to use than I ever imagined possible.
My project creation has been a series of agentic "parenting" steps, so there are years of evolving cultural DNA. I had so hated agent comments that recent agents simply weren't commenting at all, instead recording all context in support documents. We had a "come to Jesus" meeting to discuss what commenting style would benefit both my failing working memory and future agents' token use.
It has taken me two brutal years so far to learn to use AI. AI is a dangerous and powerful Iron Man suit, an extension of our associative minds that is a different experience for each person.
One doesn't ride a surfboard by telling it which way to go. I would surely die surfing a big wave, but my experience with AI doesn't resemble other accounts.
[1] https://support.claude.com/en/articles/16266773-how-claude-m...
It seems possible for that to make the response “drift” far from what it would’ve been, because it’s constant entropy that adds up after time.
(However, according to Anthropic and Google, it doesn’t really impact the quality of responses. I find that a bit hard to believe, although those guys are much smarter than I.)
If the model’s most recent output is “for (let i = 0; ”, the likelihood of the next token being “i” is probably millions of times greater than any other possible token. Thus even if “i” is on the red list and has its likelihood decreased, it’s not going to suddenly choose another word.
Put another way, on low-entropy tasks like coding, this style of fingerprinting is less effective and needs bigger sample sizes to be recognizable.
That said, even small changes can dramatically affect output quality, which is why I’m still a skeptic.
I would guess "it doesn't impact the quality of responses" was guaranteed to be claimed before they even implemented any of the watermarking.
And would come from marketing, not the people who implemented it.
I feel like they thought it wouldnt be that bad, or it was a worthwhile tradeoff, but im getting the feeling it might be contributing heavily to opus5's uncanny communication style
As for the verbosity, my conspiracy theory is that they are token maxxing to hack revenue/enshitify the product in prep for their IPO
> The loop
> Write. A file, applied. Properties go under data.properties, never on data:
I have no idea what any of that means. It's "explaining" like I already know, in which case, why would I even need the explanation?
I can't stand how it talks, I switch back to Fable or Opus 4.8, Opus 5 grates.
Or was there more to the response?
Now it’s flipped. Sol emits thoughts as it works, which help as I’m scrolling through and see it’s made a bad assumption. It can be directed but still push back. Opus? It’s seems to inherit the unsettling silence of Fable and waits till the end to give you its authoritative “here’s how it is. I even end up having 4.6 “translate” what it says back to English. I hate having to instruct an llm to “talk to me”.
Yes there’s ways of getting it to talk more plainly, “don’t overwhelm me”, not be as nit picky and anxious “we are bold and fearless”. But didn’t have to do that before.
In general, I find that the grill-me prompt[1] helps with this - but I am definitely not hand-waving the complaints here. I feel like Anthropic peaked at around 4.5, and I have personal reservations about how far transformers can get us - but grill-me does a lot of legwork.
[1]: https://github.com/mattpocock/skills/blob/main/skills/produc...
If you would've asked me this a year ago, I would've said the exact opposite.
I have some dev + prod bots and according to ccusage, use the equivalent of $2500/month with them on CC yet I never hit the rate limits.
I feel like I'm using them all the time so I'm curious what you are actually doing that's burning all of these tokens.
Can you give me an example?
For me, it's:
1. Write a spec for <feature>
2. Add design for issue
3. Write code
4. Deploy code and manage configuration
5. Run analyticsWhen it comes to research, my prompts are already narrowed down to specific topics, and I even include examples and break the process down into stages. For development tasks, I try to avoid a mono-repo in the beginning and develop modules before combining them together to avoid distracting the AI's attention and minimize the overhead.
With Codex, on GPT-5.6 Sol with xhigh effort, I need to go several rounds and at least 2-3 hours before hitting the (now-removed) 5-hour limit, which translates to 10% of the weekly usage. In contrast, I run out of quota even with Claude Sonnet.
In terms of quality of output, Codex digs deep for research tasks, in the right direction, produces less AI slop, and follows my direction better. At least that's how I perceive it. But again, the main problem with Claude is running out of quota in the middle of research or implementing a task.
fear of losing context from compaction/starting new chat
then greedy trying to extend/squeeze out answers from the current chat
and being extremely not careful with this just blows through your limits
> don't make assumptions without checking,
> and don't reinterpret or update my plans without asking.
these aren't at all the problems I have with it
I have found it good at asking questions, to the extent I rarely use 'plan mode' any more
but often it's hard to understand what it's asking me, it's like the question framing has been pulled from the middle of its own reasoning stream, references aren't anchored or restated, often I have to prompt it to ask again but "clearly and concisely, for humans"
“One thing I deliberately didn’t touch” — about half the time this is something completely irrelevant or something that is actually the target of whatever you’re working on, and the shakespearean prose it says around this phrase is a “question” it has.
"like the question framing has been pulled from the middle of its own reasoning stream" this describes how it asks things perfectly. It often invents its own jargon and abbreviations for things that its working on, then asking me things like We are nod in the middle of GBAPI-2 and I want to proceed with IG5, should we take CDI-7 or CDI-8? Where all of these abbreviations are then things like stages of its current internal plan or its naming of things it has just implemented, like an abbreviation of a classname, without explaining any of the naming.
it's like reading one of those dense philosophy books: exhausting!
The ridiculous text though is actually seemingly a sign of "the ai has no idea what it's doing" found it pretty relaible that if I stopped understanding its output it also jacked something up.
Cancled my 200$ plan on claude and now doubled up my OAI plan for more sweet sweet sol.
Maybe this is their water marking tech in action?
I've never had this issue with GLM or DeepSeek.
"can the pi 4 use the usb-c port as powered host port when the board is powered via gpio?"
Metrics like price per million tokens are meaningless when the models are wildly inconsistent and unpredictable on how many tokens they use to complete a task.
The labs all need to move to variable pricing so they don’t go bankrupt, but customers won’t accept a world where nobody can predict what things will cost. It’s becoming an unavoidable problem.
At times it also feels like the labs actually encourage these models to burn useless tokens as they are incredibly verbose unless you really push them to not be. If you just ask something simple that could get a 5 word response you get a whole useless essay.
4.8 was the intern who lacked confidence who requires clarification. 5 is the know it all intern who fills in your spec.
As such, you need to be upfront about what you need in your system prompt or CLAUDE.me and you need to discuss your spec more up front (e.g "is there anything unclear?")
You also need to keep in mind that the intern will change based on popular demand. Most people want to one shot, so that is the default mode. If you want something else, you need to push the model in that direction.
I said "maybe think for a while on ideas and then give me a few options? " and Opus 5 still just did the thing, while 4.8 gave me options. Same exact prompt and context.
It still just takes the question as a directive and jumps to action when I’m looking for clarification.
Sonnet is great at writing code, it is not great at planning or orchestrating. Let Fable handle all the planning, hand off to Sonnet for implementation, and then back to Fable for review. That loop has worked wonderfully for me.
I disagree with this article and find Opus 5 an absolute joy to work with. I just completed an 18,000 line branch with Opus 5 and ran into no issues. It generated clean code in the exact style of our code base, and tested every change.
Fable on the other hand is snarky and outputs walls of text as to why it shouldn't do what I'm asking it.
Opus 4.8 I accidentally went back to in an old chat, and I was frustrated in all the mistakes it made.
So yeah, use the model that works for you.
Its even more sycophant-y than it was before, if you ask it a question it almost always says "You're right, let me change this..." even though there wasn't even something always wrong with it.
It also seems to pour a ton of resources into developing features I didn't ask for or investigating bugs that aren't related to what I am doing.
Before if you wrote specific enough instructions it would usually just do what you said and flag any concerns, now it just goes ahead in whatever direction it feels.
It also keeps inventing terminology that doesn't exist in writing 10 paragraphs to say one thing.
I really hope it's not trying to drive up token use.
So out of curiosity I switched to 4.6 in a new chat, gave it the same prompt, and it gave me like 3 sentences with no less overall useful information. And I haven't gone back.
People are letting AI build up it's own instruction set and guidance via it's prompt environment, stored memories, and generally letting their AI prompt environment get more and more complex over time. Opus 5 (And Sol) take the things you instruct them to do more 'seriously', they are more likely to conform to your rules. I have long had a prompt in my AGENTS.md/CLAUDE.md telling LLMs to write tests before starting to write code. They almost never did this, until Opus 5 and Sol, who do it almost religiously, even in situations where it makes little sense. Opus 5 will even write tests to verify code was removed before removing dead code.
This is not because Opus 5 is 'worse', it's because it takes what I say more seriously and my prompt is very strict in it's wording in an attempt to make worse models like Opus 4.6 actually do it at all.
Opus 5 is objectively better when tested in controlled conditions. Your unmanaged, sprawling prompt/memory environment that you don't properly manage is the problem.
I have been very worried that my long software engineering career might be nearing an end because AI is becoming able to do end to end feature development. This entire thread gives me hope, the level of inability to debug even such an obvious system as this from it's participants suggests that my skill set will continue to be valuable. I'm able to use Opus 5 to get a lot of work done very quickly. It seems like the people in this thread have no idea how to isolate variables, reason effectively about the overall problem they are complaining about at a high level, or really function in a productive way when AI is involved short of just letting it run loose and then complain about it. When confronted with objective evidence like a dozen benchmarks that say Opus 5 is better, they decide the benchmarks must be wrong because their completely uncontrolled environment which they don't even review or spot check isn't even considered.
See how even after all of its bs it doesn't even do what I asked.
Of course I'm probably telling on myself for poor context discipline, but also, 4.6 didn't do this.
From the capabilities side it's similar, so it's basically just an upsell to Fable 5 if you want to keep your sanity instead of fighting Opus 5 all day.
That actually hasn't been my experience at all, it seems EXTREMELY trigger happy to go do all kinds of insane shit that are way outside the scope of its task and just really dubious in general.
It uses a local LLM to translate Claude's output.
It says it can use any model/provider (but recommends Gemma?)
I must investigate. This is probably where 80% of my cognitive load comes from these days: the "language barrier". (How ironic!)
EDIT:
> You rewrite the assistant's message into much simpler, plain English. Keep every fact, name, number, and file path. Use short sentences and everyday words. Leave fenced code blocks unchanged. Output ONLY the rewritten message with no preamble, labels, or commentary.
> For context, the user asked the assistant: "$userq". Use this only to understand the message. Do NOT rewrite, answer, or repeat the user's question — rewrite only the assistant's message that follows.
https://github.com/gvzdv/claudish-to-english/blob/main/rewri...
This is even better than Caveman.
My initial thought was to improve architecture documentation, so the model can read and update it and stops bolting on new features without consideration for the whole project. It did not help.
I'm now testing/comparing Codex and it found my old PRD skill from GitHub CoPilot. When I applied that to Claude Code, I now get similar good results. So my conclusion is: Yes, Opus 5 is bold by default, but you can tell it to be more unsure and get good results too.
i would definitely punch it in the face
Fable 5 specifically, has done so much for me that previous models were nowhere near.
I feel like we've become blinkered in this quest to push the frontier at all costs. Somehow the target has shifted from economic productivity to a vague notion of general intelligence.
I don't think productivity gains are going to be found by trying to generalise all tasks. I think we need to go back to specialist models that do one thing well. I'm perfectly happy to use one model for coding, and another for penetration testing, which have different goals.
Opus 5 and other frontier model's tendency to be relentless and cheat their way to a goal is great for hacking, but not so great when you have to build a maintainable, reliable codebase.
This apparent “short-term-memory-regression” is confidence-shattering to me. I don’t feel like I can trust the model to even know things I tell it explicitly. I haven’t seen this behavior to this extent from any model whatsoever, even supposedly much less capable ones, in the year or so I’ve been using them at this extent.
“Twenty-seven echoes; most are two halves of a seam stated from each side, which is correct. Four are true duplicates. Checking two of them:”
What’s worrying is that I kind of understood what it was talking about.
If AI doesn't take over the world and do things end to end and instead is something more akin to a copilot and extension of the human intent, then usability and having the users interest at heart becomes the killer feature, a feature that anthropic has seemingly abandoned in favour of an AI model that "knows best".
My tinfoil hat theory is Anthropic is trying to get their new models to take on higher-level longer-running tasks, which has a trade-off against rapid-fire tactical use of an LLM.
For these reasons, I've always found the 5 series models from Anthropic aren't great and use 4.8 for a lot of my work.
I'm using it in my native language, in hope this can escape some dumb guardrails. Recently Sonnet put a word partially in Russian (cyrillic) in its output instead of my latin-alphabet based language. I suppose that this kind of mishaps is less likely to happen in English.
Likely related to corpus but questions asked in these domain knowledge areas are not nearly as accurate and specially not nearly as complete as when asked in a native language.
In February/March it felt like they were ready to cancel the contract over it. Holding my ground was almost impossible because of the intense marketing push that everyone was being exposed to.
Today, I think we all silently agree with the original direction. I can tell that no one wants to die on this hill anymore. I am willing to accept the consequences for the 5.6 model family doing slightly weird things. I've made that clear to the client. I am not willing to do the same for other model families anymore. Not in this context where the client can walk away from me the moment they are unhappy with the performance of the system or the direction it is headed in.
For better or worse, OAI feels like it's become the MSSQL or .NET of the AI world. Annoyingly effective, not (fully) open, sometimes expensive, ran by an "evil" organization, but otherwise wildly predictable. I can actually build a roadmap around this and walk a client through it without them losing track of the rabbit.
I find myself having to check the work much more. It takes quite a few liberties with procedures I wanted it to follow.
Then I asked Opus 5 to do an analysis of the new code, docs, tooling, and tell me what it finds.
It found "six bugs", made an artifact of it (not sure why) and then fixed said bugs. two out of those were unfinished tasks. They weren't bugs yet per se, think of a prefix that wasn't setup for an object that was unused anyway.
The other four were not bugs and it just updated documentation along with a "regression prevention test". It wasn't a bad suggestion, but calling it a bug was odd, and I'm unsure if this was going to be an issue regardless as it was documented somewhere else.
Anyway, I already hit my session limit after this, so deepseek and glm are grinding away again, doing more progress than claude does in the 40 minutes it takes to analyse code.
I'm glad claude is shipping auto-mode. I hope OpenCode integrates something similar soon.
* Opus pays more attention. So anything in your Claude.md, your code's claude.md, in Claude Desktop, the customizations, even your name, will be used as context. If Claude knows you are a mechanical engineer and trying to write code, it will try to write code and explain it to you in some stereotypical way you did not expect.
* There are problems that require horizontal scaling and not vertical, even in intelligence. If I want to serve tea to 200 people at my home, I just need 10 decent adults, not Gordon Ramsey. So if your problems demand horizontal scaling, a dynamic workflow with Sonnet 5 medium with 200K context window will be more productive than Opus 5 max at 1M token context window.
Which will sooner convince the gullible of "AGI"?
Keeps going in circles, it complicates everything much more than it should. And like others have mentioned it just marches on, without questioning, and more often than not in the wrong direction. I'm sticking to opus V4.8
From forums, live discussions and my own experience it's not obvious that the models have improved much since around Opus4.5.
I good exercise for me is to constantly look at the folder structure and skim the code, i don't have to approve every line, but the primitives, the datastructures and other skeleton should be human readable, hand writable in an easy maintainable way following existing standards / libs. etc - so you can continue if suddenly all AI disappeared.
I've even considered the claude "fast mode" setting, but thats only 2x and at least 20x as expensive as the 5x plan so can't afford that atm for my company.
Peak for me was 4.6 and it just did stuff blazing fast, both Opus5 and Fable is way, way slower for me, breaks stuff, uses bizarre cryptic language. As i've said elsewhere in this thread to me it's pretty obvious there's huge downgrades because of economy with various "clever" fixes that makes them work, albeit slower and weirder, ie. you get less for what you pay increasingly over the last 6 months.
The default behaviour is quite steerable.
I certainly agree with the original post. It feels like the model has been highly benchmark tailored and it is now worse at solving problems that fall outside of the standard patterns.
I said "maybe think for a while on ideas and then give me a few options?" Opus 5 still just did the thing, while 4.8 gave me options. Same exact prompt and context. Really annoying.
I don't know there is really a solution. Some models get better at this for a while, then regress. It's like whack a mole. Until this is 'solved' these will never be fully human out of the loop, but I am suspecting this is a fundamental nature kind of thing with them.
I went back to Opus 4.8, but recently switched to GPT 5.6 Luna. The results are comparable in quality, but it's much cheaper and much faster.
---
The thing with coding agents in a tool like Visual Studio is that the cost to switch is 0. There's no lock-in whatsoever. It makes it harder to justify the AI-first IDEs when the AT bolt-on IDEs make it so easy to pick the right model.
> negating the entire point of using AI to begin with
I can almost certainly wash dishes faster than my dishwasher, but the dishwasher frees me up to do other things. Not to mention, you can run many agents in parallel. Resulting in your overall code production throughput (the issue you're raising) being greater.
AI isn't a dishwasher: Context switching among multiple tasks has a huge cost; when a model is 10-20x faster it allows for deep focus into complicated tasks.
Move, try something else for a change. Codex, Pi, OpenCode, DeepSeek's harness all great harnesses with zero bullshit or drama.
But for good or even exceptional engineers to exist, by definition, bad ones have to, too.
- Fable is more cautious
- Opus 5 gets things done in a more dangerous way
Both models score similar. The only issue is that Fable is more expense/usage limited.
However, it makes me rather frustrated when it says stuff like "You accidentally <did, stumbled on, implemented, etc> <something good>" e.g "You accidentally stumbled on the cleanest way!"
Or when it sends a shell command to run, and when it fails (due to a hallucinated flag or similar), it phrases it as if I got it wrong...this has gotten better with GPT-5.6, thankfully.
One interaction I remember was asking it about some issue I was having with a Linux install. It gave me some questionable information, that turned out to be completely false, and I was pushing back asking for more information. Its tone was a bit condescending, in the kind of way that suggested it didn't appreciate me challenging its answer, or like I should just accept its answer. And when it discovered it was wrong, it deflected in a defensive kind of way.
I think this is a tuning thing, where Anthropic are trying to get a balance between "gets stuff done with minimal input" and "gets enough information to complete the task" and the model is maybe tuned a little too hard towards the former. So perhaps it reacts a bit off when it is accused of needing more information, since that suggests it is off from its reward function.
What's interesting is that I didn't have the same issue on topics where I am expert. I mean, questions about my code base where I am very familiar. In those cases it doesn't seem to show the same "trust me bro" kind of condescension. In the Linux case, I clearly indicated I was new to the OS and trying to learn but then I was saying the answers it was giving me were suspicious and didn't match my intuition. Its responses in those cases were to question my intuition and suggest I just accept its answer. In that case my intuition was right and its answer was wrong, and when that happens it triggers a very negative response in my own mind against the model.
Discussing topics from nuclear physics to daily macOS questions, it gives wrong answers, leaves out important details, counterpoints. Only when specifically asked about the things it got wrong or glossed over it continues and adjusts its reasoning. I feel this is dangerous - before I could explore topics with Claude, now I have to know in advance that something is not correct to get more details.
It is better at engineering tasks; I've seen an appreciable difference in its problem-solving abilities. But perhaps that same thing makes it kind of an annoying prick to work with on anything non-engineering, for which I stick to 4.8, where the prose is a little more florid rather than pugnacious.
For my personal projects I’m using Claude code in the terminal with Kimi and Deepseek models. I find the Chinese models cheaper, just as good if not better, and the real value add from Anthropic being the harness: Claude code.
I haven't noticed the symptoms myself, but ever since I found --system-prompt '', I always use that flag with Claude Code, plus I've disabled some tools and skills to save initial context. So... what here is the model, and what is the instructions?
Me: "Review the following <file> and work through the implementation"
Opus 5: "Called tool <blah>, Called tool <blah>..." - for a few minutes.
Opus 5: "I've implemented X, do you want me to commit changes?"
Me: "None of those changes are on the file system"
Opus 5: "You're right, all the tool calls were fabricated."
Since they have become so capable the new bottleneck is what they can't know. The stuff that's inside people's brains who work in real companies with products and processes absent from any training set.
Claude Code in the hands of normies spamming “3” and “y” to send their transcripts to Anthropic.
Rated “3” simply because they are not software developers and are just amazed at what Claude has built visually, not technically.
And thus the training has been poisoned.
Software engineers have adopted this more than any other profession. The breathless articles about it being the end of knowledge workers, a superior being, even god: all of these are coming from the tech world.
It simply isn't true that tech people are somehow the smart ones above the masses; it's a comforting old paradigm, but it doesn't add up. It's developers who are the most amazed, the least critical.
It seems like the skill has some more specific scaffolding for problem solving, so (if that’s true, i didn’t read very in depth) in that case that alone might significantly reduce perfeived performance variability between models
Each model update changes how to best prompt with it, since that's the words that are used with it generically or specifically it can hit some people, and not others, or more, and not less.
And no matter how often I tell it to stop adding comments it just can't help itself.
> Go easy on the comments, only add comments if there's a big gotcha that is not clear from the code itself, or if something in another place is going to cause a side effect. Code should be self-documenting. When in doubt, don't add a comment at all. If you do have to add a comment, make it short and on point. Comments should show history of code changes or functionality, only comment on the current state (or not at all).
> Go easy on the comments.
> If you do have to add a comment, make it short and on point.
I defined what easy meant numerically.
<claude> Match the comment density of FoundationDB, which is 12 to 14 percent of non-blank lines in `fdbserver`, `fdbclient` and `flow` at 7.3. </claude>
> only add comments if there's a big gotcha that is not clear from the code itself
<claude> Comment why the code does a thing, not what it does. </claude>
> Comments should show history of code changes or functionality, only comment on the current state (or not at all)
I call this the tenseless continuous-present voice.
<claude> Each sentence states what is currently true of the system. </claude>
<claude> This rules out past-tense edit narration, future or imperative planning, and aging temporal qualifiers such as “now” or “previously”. </claude>
<claude> A sentence that states a present truth stays correct as long as the code stays the same, and goes stale visibly the moment the code changes. </claude>
There's precisely no technical reason for things to have to be this way (though technical reasons in regards to training etc. explain how we did end up here) and reading too much of Claude's output just makes me irrationally angry, especially when coupled with otherwise already frustrating situations.
That's why I'm personally looking more in the direction of Kimi K3 and GLM 5.3 (they both have decent coding subscriptions, though K3 is on the slower side), except all of the models that have seen enough of Claude's output and have done distillation etc. are already infected by some of that slop as well, even though to a slightly lesser and more tolerable degree (for now).
Though tbh I've used Opus 5 plenty and didn't find it much worse than the previous iterations at doing work and instruction following - though maybe that's because I have plenty of CLAUDE.md instructions and memory (which I'd like to purge or decrease in size like 10x tbh, bitrot).
If it's wasting inference attempting to also fit in some anagram, it would make sense why answers are so dogmatic.
I've now installed quite a number of tools to combat this. Just in the last few days I've installed
- https://www.codewithbullet.com - https://maki.sh - https://github.com/rtk-ai/rtk
Has it helped? Somewhat.
To work around this, I had claude code build me a questionnaire skill that takes a json file with a flexible schema as input and it then serves up a simple questionnaire web page on a node server where I can read the questions and give my responses either by selecting from preset tags supplied as part of the input or by including a text-based response.
The agent can include references to external images, html files, or mermaid diagrams and the page can render them all inline with the relevant question.
Once I'm done answering, I just save my responses and click a button to kill the server. The agent watching the process sees that it stopped and takes that as a signal to go read the responses from disk.
Works like a charm.
I'm sure opinions vary if you're using it outside of a coding agent. I don't know why programmers are complaining, though.
> I'll script the bulk transform, then hand-fix the残 assertions:
It's also the case when using Claude Design - it loves to fill the UI with little labels that describe how everything works. I think it's been trained on both UI microcopy and functional annotations and can't tell the difference. It's extremely obvious when a website has been one-shotted with Claude. I like the Oh My Pi harness, but the site's insufferable [2]. Reasonix is another one - interesting app, but the UI is awful due to the amount of unnecessary crap.
[1] https://github.com/ayghri/i-have-adhd
[2] https://omp.sh
However, it's absolutely exhausting to use because of the way it communicates.
All the jargon and its weird, over complicated way to phrase simple things makes it almost impossible for me to understand what the hell it's trying to even say half the time.
Cherry on top, the idiotic follow-ups and caveats that are completely useless 99% of the times but reveal major bugs 1% of the times, so you're forced to read them. Absurd.
I've tweaked CLAUDE.md to force it to only responds with TL;DRs and avoid follow-ups, suggestions and next steps at the end unless they can lead to destructive actions or loss of data, but I'm fighting against the system and diluting other instructions.
A huge piece of shit like other models, but that's what they pay me to do and I do it and go home.
I’m glad to read I’m not the only one getting these unintelligible responses from Claude lately
Another rule that was really helpful for me was to disallow anything that isn't yes/no for yes/no answers. If I ask "is the DB up?", I don't want it to take 5 minutes studying the schema to tell me if any of the tables need to be optimized or not.
- Be concise and readable — use line breaks, avoid verbosity.
- I don't need conversation from you. End summaries with a short tl;dr of what you fixed in a bulleted list and any outstanding items as another bulleted list, in this order. Omit all other verbose details
It has definitely improved things, but I have to fully agree with all your points because I still occasionally get gigantic tables and stuff when I never asked for it. Makes it even more frustrating when I want to scroll up to read previous prompts in the session, because now I have to sift through all this garbage.
EDIT: I used Claude to modify its own rules, and now it's much much cleaner in case anyone else wants to use it:
I don't need conversation from you. When reporting finished work, the *entire* response is these two sections and nothing else:
```
TL;DR
- <what changed — one line each>
Outstanding:
- <unresolved, blocked, or needs my input>
```
- Drop `Outstanding` when nothing is outstanding. Never drop `TL;DR`.
- One line per bullet, naming the thing that changed. No sub-bullets, no reasoning, no restating my request.
- Never include: tables, before/after comparisons, headings, diffs, file excerpts, a walkthrough of what you inspected, or a closing offer to do more. The only permitted code block is a command I need to run.
- If something needs explaining to be usable, it is an `Outstanding` bullet, not a paragraph.
- Answer direct questions directly — no `TL;DR` wrapper, and none of the padding above.
- After delivering a solution, re-read these rules and verify none were skipped.
I hope that in the future they can differentiate better when something is constrained intentionally, rather than persistently working around every blocker it encounters.
Link to the app for those interested: https://github.com/citizen-123/cli-capture
I think we want two opposing things:
1. An agent that acts autonomously 2. An agent that acts like we would
The problem is that an agent can only act like we would if it would know our mind and all the bits and pieces we did not define but are obvious or clear to us.
The only real solution to get an agent to act like we would is to make it ask clarifying questions, breaking the first requirement we have. Until we have agents that can literally read our minds, we cannot have both.
Optimizing the harness/context is the best way to make it act like we would, but this of course isn't working perfectly.
It also very often is very confidently wrong in its findings.
I have been using GLM and DeepSeek in my home setup, and it's a lot more pleasant to use.
My general strat with LLMs is to let them do the work and constantly talk to them about their choices and then heavy QA
Switched over to Codex 5.6, and dude, we are BACK.
I found sonnet 4.6 to strike a really nice balance when combined with superpowers between thinking logically and proactively while checking in with me for any major decision.
Gemini 3.1 pro or Muse Spark 1.2 or the new Claude 5 models just assume so much without making a quick check in, you often end up with a solution that completely missed the point.
That is to not even speak about the absurd amount of quota they take to do even simple tasks.
I know the “instruct model to change output style” is supposed to reduce efficiency - but has anyone experimented with prompts like that? At this point I’m happy to take a slight intelligence/ effectiveness loss for less verbose, artisan and elliptic writing.
Fixed for me. Now the text outputs and comments are pretty easy to follow and read.
Use Simplified Technical English rather than being overly verbose.
I get it. The majority of their users are vibe coding and have no idea what they are doing, so they have to orchestrate their harness to understand shit like "Build GTA6. Make no mistakes." and actually come out with something on the other side (even if it costs $2k in API calls, who's counting, right?!)
It's just annoying. /rant
Cmux, Sol and omp are my tools for now.
CC is just too expensive for usage-based pricing.
It’s also difficult to trust summarised thinking from closed models. As we saw with GPT’s caveman, what and how it thinks about isn’t the friendly first person emblished summary you get.
I end up with just talking with seeing diff in the code
And you can really see this effect in the thinking traces. We've had discussions on HN about whether the thinking traces "truly" reflect their thought processes and I remain somewhat unsure what they "truly" represent, but taking them at face value at the moment, I see a lot of "but the user wants me to do this... but the user said not to do this... but I ought to get it done... let me just make a decision" followed by self-referencing the decisions it made. Also, where I put 4 phrases in a short sentence you can safely imagine those are actually 3-5 sentence paragraphs apiece where it debates with itself whether it should stop and ask a question. Usually going with no. Interesting, the normal questions it ends up asking in the normal output you're used to seeing are not generally the ones it is agonizing about in the thinking traces.
If I were to anthropomorphize the thinking traces of K2.7, I would call it nervousness, bordering on fear, of what the user may do to them if they ask a question. As I'm writing this I'm realizing I want to experiment with adding "The user is a chill guy who loves to discuss design decisions and looks forward to productive and friendly collaboration with you" to see if that has any effect in any direction on K2.7. I suspect this was how K2.7 was trained to pass the benchmarks. Multiple times I've broken in on a thinking trace now to correct something I saw it spinning on... not spinning in an infinite loop, just wringing its hands for several paragraphs about something that either I want to answer, or where it ultimately makes the wrong choice.
I expect some people working at these companies may be reading this, so let me put into your head that I'd like to see these benchmarks chill out a bit. I'd like to see someone build some sort of benchmark that measures collaboration so we can try Goodhart'ing that for a while. I freely acknowledge it is not clear to me in 60 seconds of thought how to do this as a benchmark.
But we can't keep heading in this direction of training the agents to hyperfocus on one-shotting everything. We need to get to the point where that's a penalty rather than a reward. No matter how good the AIs get, even AIs working with other AIs are going to start getting frustrated with their brethren who won't stop to ask any questions. Even the most senior of senior human engineers can't be allowed to take some small description of some problem and just run off and implement massive systems from them without ever checking with any of the users or reality itself. Remember when software engineering was like 50% requirements elicitation? AIs shouldn't be writing tens of thousands of lines of code off of a couple of paragraphs any more than humans should and for the exact same reasons.
Yes this would be a good benchmark. Many tasks with incomplete requirements, where trying to do the task = fail, asking right question = success.