Anthropic Risk August 2026 [pdf]
www-cdn.anthropic.com
www-cdn.anthropic.com
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
They can’t measure even measure it, it’s just vibes. They may not even be more productive.
Maybe they should ask an AI to create one!
slop
bubble
I find it hard to imagine launching this criticism at a new technology.
Seems like that would be easily recupable even without future growth.
While it is true that investors only got a fraction of the company for that money, and their EV can be debated, I think the clear the value is there from a net cash in to value produced.
Last valuation was like 1 trillion. Company could be worth like 1/20 and still justify the cash.
I'm not betting on whether it will or won't. I'm absolutely 100% suggesting it hasn't happened and isn't guaranteed.
I agree returns on actual spend are not guaranteed (nothing is), but think them pretty likely. Valuations are are more challenging, but risk is known there.
A performance improvement is relative to some baseline, and that baseline for some may be a lot lower than others, and if they adopt tech effectively it really could be a big boost. Across the board though I don’t think it’s sensible.
I work in an industry that very intentionally tries to be inefficient and I can tell you having certain tasks automated that before had a person barrier intentionally acting inefficiently that you can now sidestep by outsourcing their tasks to something like Claude gives me a massive performance increase because I’m not blocked as much anymore. I can literally just replace some external tasks that were intentionally slowing processes down for their own benefits with a few prompts and move along. I could have done the tasks before but then people would ask why I’m spending my time doing it, now I can just say “oh, I was blocked so I had Claude take care of that blocker” and move along.
It can further be true that some developers are getting 5x or 10x while as a whole their organization is sped up less than 2x. I'm sure many tasks at Anthropic are sped up 5x or 10x or more.
It can further be the case that many people overestimate their gains as well. That's fine, and I think what you're saying - but it's still wild to shake your head and go "pfft, they have not even doubled their productivity". Double is a lot!
Is that correct?
They've probably already settled on most of the architecture and the big ideas, so they're details in big things instead of how to make complete small things.
The thing LLMs really speed up is how some ordinary person-- a PhD student, or similar, can whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where LLMs shine: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.
Also it occurs to me that they're somewhat incentivized to downplay cyber risks after what happened last time...
> Same model weights as Mythos 5, deployed with higher-coverage safeguards (see Section 4.5.2.2)
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
Seems like a strange expectation.
Expectations that fail to match reality are a sign of confusion or mental illness.
27B local model just dropped, it's 6/8-month old SOTA. General ROI of AI investment is expected on a baseline of >10y.
Why would it be important to explain things to you? Are you one of their major stakeholders?
If Chinese model hacks US government... free marketing?
By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.
This sounds like it might be a Mythos finetune for some specific task.
EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench
It does feel they are trying to ask the government to lock the market for us.
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
What am I missing?
LOLOLOL
>6.3 [Appendix redacted] > This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.
interesting...
EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.