195 karma · joined July 4, 2011
For intelligence/size only OpenAI and Anthropic are the frontier. Google has more compute so it can compensate for that with size of the models...
FF works for me.
That could explain the token usage difference because larger models usually use less tokens per the same unit of intelligence.
This means that I, as the owner of my data, can refuse to provide it for some use cases, request its deletion, etc. It’s my data after all.
Article: https://alignment.anthropic.com/2025/subliminal-learning/
Doesn't this only apply to subagents, which don't have much long-time context anyway?
> Opus 4.7 is substantially better at following instructions. Interestingly, this means that prompts written for earlier models can sometimes now produce unexpected results: where previous models interpreted instructions loosely or skipped parts entirely, Opus 4.7 takes the instructions literally. Users should re-tune their prompts and harnesses accordingly.
See e.g. https://epoch.ai/data-insights/llm-inference-price-trends/
Also, there are other measurements like inequality, healthcare cost, social securities...
So the closer the two related pieces of information are to each other in the input context, the larger the chance their relationship will be preserved.
It seems that government wanted to use Claude for mass analysis of commercially obtained data on American people and Anthropic wouldn't let them (source: https://www.theatlantic.com/technology/2026/03/inside-anthro... ).
DoD kept asking for changes of contract where at least the legalese would be changed to somewhat more permissive but Anthropic stayed their ground.
Sam Altman probably let them do that, while using language like "we have technical means of oversight and the same red lines as Anthropic". But in reality they will allow DoD to do what Anthropic didn't.
See this for more information: https://www.lesswrong.com/posts/PBrggrw4mhgbksoYY/a-tale-of-...
I've experienced that too. Usually when I request correction, I add something like "Include only production level comments, (not changes)". Recently I also added special instruction for this to CLAUDE.md.
I sometimes reference some of them to build context, e.g. after few unsuccessful tries to implement something, so that Claude doesn't try the same thing again.
Largest production capacity maybe?
Also, market demand will be so high that every player's chips will be sold out.
Models from Anthropic have always been excellent at this. See e.g. https://imgur.com/a/EwW9H6q (top-left Opus 4.6 is without thinking).
Dollar/watt are not public and time has confounders like hardware.
We already know that intelligence scales with the log of tokens used for reasoning, but Anthropic seems to have much more powerful non-reasoning models than its competitors.
I read somewhere that they have a policy of not advancing capabilities too much, so could it be that they are sandbagging and releasing models with artificially capped reasoning to be at a similar level to their competitors?
How do you read this?
Care to elaborate more?
For some tasks where detecting hallucinations is easy I can see it being beneficial.
In general case not so much...
So they could have paid a price in “model welfare” and released an LLM very eager to deliver.
It also shows in AA-Omniscience Hallucination Rate benchmark where Gemini has 88%, the worst from frontier models.