It's horrendous. It constantly scope creeps, will attempt to use admin overrides, assume you are incompetent, speak in half thoughts, and just blatantly ignore instructions. Opus 4.6 was peak for Anthropic. Its a lot better outside of Claude Code but it's still annoying. I vibed out a CLI tool to do packet capture for TUI applications to get an idea of why Claude Code makes it worse. The amount of additional unnecessary context and tool bloat that goes with your sessions is crazy. The memory system ships so much extra info about what you did yesterday that I think it's misguiding the model.
Pebble is the ultimate smartwatch because it knows what it is. Its a watch first and foremost, readable in any lighting conditions, waterproof, battery lasts for weeks. They made a watch first then added only the QOL features that most people need.
As someone who regularly attempts to give up a cell phone, it is not optional. A not insignificant number of things require it. It is the expectation society is now built around and therefore it is not an option. A car without telemetry is for the time being still optional. I take very good care of my old vehicles hoping I never need a new one.
Most large orgs do not need to train end users. They just need to add glm-5.2 to their router and their in house harness will pick it up. Then slowly limit usage on anthropic models and people will swap willingly. It's a simple /model command in every harness.
It's in YouTube's best interest to only show users content they're interested in. Replace the word algorithm with users and you'll have a more accurate representation of how YouTube actually works. The reason those videos didn't get the love they deserved is because they're niche content, the 10 minute review videos appeal to a wider audience and therefore gain more traction.
It's crazy to release a model that just swaps you to another model when you ask it hard questions. Fable changes to Opus 4.8 when you talk about cybersecurity, biology, and a couple other categories. You still pay Fable input token cost though. Frontier models are stalling, this is anthropic trying to hype the market up. Now they're talking about stopping frontier model research. It's kind of strange how the moment they become the highest valued AI company, all of a sudden they're talking about everyone stopping frontier model development for "safety". They're just as corrupt as the rest.
is:unread -is:starred <-- go through your inbox and star what you want to keep, this filter will help you delete everything else unread. Add something like older_than:1y to also prune your Gmail from time to time.
This is the same gripe I have over any LLM vulnerability tooling. 95% of what gets flagged is something that if taken by itself could be a vulnerability. However, the path to execute that specific vuln, in that specific function, is impossible in that particular code base and it just makes noise.
Can someone give a tldr on why this happens so much with npm ? I can't recall seeing this with any other package manager. Is npm just the default used these days and therefore sees this more often?
I've been saying for a while that given a proper harness, small local models can perform incredibly well. When you have a system that can try everything, it will eventually get it right as long as you can prevent it from getting it wrong in the meantime.
Am I correct in my understanding that they are not actually able to 100% know what Claude is thinking? They have trained a new model to make a guess about what Claude is thinking, but we cannot validate that the guess is 100% valid, right? They are basically saying "we have trained a model to reaffirm what we believe Claude is thinking" ? Hoping I'm wrong in my understanding of this because this does not appear to be good research to me.
I am in the same boat. Reading is a transaction and lately everyone wants to put 60 seconds of effort into writing an article and expect me to put 10 minutes into reading it, and I just can't. The writing feels dead, soulless even. Every sentence or phrase is structured like a mongering, click baity headline and it's insufferable.