The roulette wheel isn't rigged, sometimes you're just unlucky. Try another spin, maybe you'll do better. Or just write your own code.
The roulette wheel isn't rigged, sometimes you're just unlucky. Try another spin, maybe you'll do better. Or just write your own code.
Also, another difference is the stochastic nature of the LLMs. With table saws, CNC machines, and modern 3D printers, you kind of know what you are getting out. With LLMs, there is a whole chance aspect; sometimes, what it spits out is plainly incorrect, sometimes, it is exactly what you are thinking, but when you hit the jackpot, and get the nugget of info that elegantly solves the problem, you get the rush. Then, you start the whole bikeshedding of your prompt/models/parameters to try and hit the jackpot again.
This is a guy with 10+ years experience as a dev. It was a watershed moment for me, many people really have stopped thinking for themselves.
The way humans are depicted in Wall-E springs to mind as being quite prescient, it wasn't meant to be a doco
I think part of the problem is that our brains are wired to look for the path of least resistance, and so shoving everything into an LLM prompt becomes an easy escape hatch. I'm trying to combat this myself, but finding it not trivial, to be honest. All these tools are kind of just making me lazier week over week.
I know I know you're going to say (or simonw will) that effective and responsible use of LLM coding agents also requires those things, but in the real world that just isn't what's happening.
I am witnessing first hand people on my team pasting in a jira story, pressing the button and hoping for the best. And since it does sometimes do a somewhat decent job, they are addicted.
I literally heard my team lead say to someone "just use copilot so you don't have to use your brain". He's got all the tools- windsurf, antigravity, codex, copilot- just keeps firing off vibe coded pull requests.
Our manager has AI psychosis, says the teams that keep their jobs will be the ones that move fastest using AI, doesn't matter what mess the code base ends up in because those fast moving teams get to move on to other projects while the loser slow teams inherit and maintain the mess.
Absolutely, not understanding why you even ask. Humans are creatures of habits that often dip a bit or more into outright addictions, in one of its many forms.
But it's also a tool that (can) save(s) you time.
This scenario obviously does not apply to folks who run their own benches with the same inputs between models. I'm just discussing a possible and unintentional human behavioral bias.
Even if this isn't the root cause, humans are really bad at perceiving reality. Like, really really bad. LLMs are also really difficult to objectively measure. I'm sure the coupling of these two facts play a part, possibly significant, in our perception of LLM quality over time.
When I fixed this, it was like magic, working how I wanted again. I now have a skill to periodically audit MEMORY.md and CLAUDE.md according to the conventions I've learned work best for me - which I suppose /dream is supposed to handle eventually, but you're kind of trusting it to audit its own memories, which have, at least to me, already proven to be unreliable.
With so many factors like this, not even to mention context exhaustion, window size, effort, etc. - anecdotal evidence is almost worthless without examining someone's entire local state.
A lot of it, to me, feels like user error, I haven't really noticed much behavioral difference between 4.5, 4.6, and 4.7, at least in my own workflow. I will note though that constantly managing these things is a lot of work that I hope one day becomes less necessary. It's more than I can expect people on my team to manage on their own, and unless I sit down with them 1 on 1 and review their issues, or write some clever agent to help them, I don't really know how I can help people reporting things that I hear posted here a lot.
I've cancelled my subscriptions to both Codex and Claude and am going to go back to writing my own code.
When the merry-go-round of cheap high quality inference truly ends, I don't want to be caught out.
"I think we can postpone this to phase 2 and start with the basics".
Meanwhile using more tokens to make a silly plan to divide tasks among those phases, complicated analysis of dependency chains, deliverables, all that jazz. All unprompted.
Don't use these technologies if you can't recognize this, like a person shouldn't gamble unless they understand concretely the house has a statistical edge and you will lose if you play long enough. You will lose if you play with llms long enough too, they are also statistical machines like casino games.
This stuff is bad for your brain for a lot of people, if not all.
On the upside, there wasnt much to atrophy in the first place
Some day maybe they will converge into approximately the same thing but then training will stop making economic sense (why spend millions to have ~the same thing?)
And it does seem likely to me that there were intermittent bugs in adaptive reasoning, based on posts here by Boris.
So all told, in this case it seems correct to say that Opus has been very flaky in its reasoning performance.
I think both of these changes were good faith and in isolation reasonable, ie most users don’t need high effort reasoning. But for the users that do need high effort, they really notice the difference.
Though I reckon even if the HN crowd is a loud minority Anthropic has no problem with traction, and even if eventually it will the enterprise market doesn't care much about HN threads.
We aren't superstitious, you are just ignorant.