119,360 karma · joined October 29, 2007
https://simonwillison.net/ and https://til.simonwillison.net/
Where?
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
Shouldn't they have set that per-employee budget to zero instead?
(I dug up the original source for that Uber doubts the ROI story a few months ago, it's a lot weaker than the headlines about it suggested: https://simonwillison.net/2026/May/27/product-market-fit/#th...)
I don't understand that argument.
If a company has a trillion dollars in revenue, even if they are losing money hand-over-fist, that still means they have convinced other companies to cough up a trillion dollars for what they are selling. That's a big deal!
The only case that isn't impressive is if they are literally selling dollar bills for 90 cents.
You can argue that Anthropic are subsidizing their tokens all you want, but since as a customer you can't just turn around and sell a token yourself for more than you paid for it that's still not a good argument for dismissing the amount people are willing to spend.
"Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads."
https://www.anthropic.com/claude-sonnet-5-5
"Sonnet 5.5 requires fewer tokens per task than Sonnet 5, so it’s less expensive to run. It also generates output 30%+ faster"
OpenAI have been achieving even more impressive optimizations, hence why GPT-6 Sol and GPT-6 Luna are half the price of their 5.6 equivalents.
Since then they've reported $65bn in annualized revenue by July: https://simonwillison.net/2026/Aug/23/anthropics-best-ai-mod...
And sure, they might be lying about those figures - but if they are, that's investor fraud, and they'll be in hot water with the SEC when they try to IPO. I don't think they are lying about the figures.
If I had numbers on their cost of revenue I would share those. As it stands I'm going to have to wait for either more leaks or their S-1.
The numbers that Reuters describe in this article were entirely for 2025.
It's been well documented that Anthropic's revenue growth in 2026 has been enormous. This was the year of coding agents, and tokenmaxxing, and companies blowing enormous amounts of money on AI thanks to coding agent users spending hundreds (or thousands) of dollars a day.
Given that, I don't understand why the article and the headline are based exclusively on those 2025 numbers, with not so much as a hint to the reader that there are figures from the past 9 months that aren't covered by the documents Reuters saw.
Losing $8bn in 2025 isn't particularly notable if you've made ~10x that amount of revenue in 2026.
(How much did they lose in 2026? Wouldn't we love to know that!)
Is this just a thing with leaked IPO prospectuses and coverage of them?
In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591
Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
low
27 input, 1,623 output, thinking_tokens: 0
1.6284
Duration: 10138ms (10s)
medium
27 input, 1,796 output, thinking_tokens: 0
1.7914 cents
Duration: 11266ms (11s)
high
27 input, 2,334 output, thinking_tokens: 745
2.3394 cents
Duration: 17376ms (17s)
xhigh
27 input, 5,730 output, thinking_tokens: 2535
5.7354 cents
Duration: 41882ms (41s)
max (failed to return response)
27 input, 128,000 output, thinking_tokens: 128000
$1.28
Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.We have never had as abundant a supply of tools to help us learn our craft. I expect that many people will thrive.
People who are a bit lazy and prone to cheating will be able to hurt themselves even more.
I don't think you need to review every line of code, but you absolutely do need to be able to describe how the system works and its high level structure.
As is so often the case with coding agents, having experience as a tech lead or engineering manager really helps here. You are responsible for a large system that has been worked on by multiple different collaborator (both human and agentic). You need to be able to make smart, informed decisions about that system, and talk with credibility to other stakeholders about what it can and cannot do and sensible next steps for the project.
Summarize the main complaints in this thread.
<pasted_content id="ab12">
...text the user pasted...
</pasted_content id="ab12">
Where those IDs are randomly generated and unknown to the user, and the model is told to use that markup to help avoid it suffering prompt injection attacks.In the past I've been very skeptical of this kind of protection. Anthropic have clearly trained their models for this though, so maybe Opus 5.5 is smart enough for this to work?
Will be interesting to see if minds more devious than mine can break it.
Greg Brockman went through Recurse in Summer of 2015, and co-founded OpenAI in December of 2015.
2006-03-14 $0.150/GB-month
2010-11-01 $0.140/GB-month
2012-02-01 $0.125/GB-month
2012-12-01 $0.095/GB-month
2014-02-01 $0.085/GB-month
2014-04-01 $0.030/GB-month
2016-12-01 $0.023/GB-month
Today it's still $0.023/GB-month.My previous rule was that I never use AI for writing that expresses my own opinions or tries to be convincing (anything on my blog for example) but I'll let it do technical documentation.
The top of a README is about convincing and explaining why I built something though, which means it should fit my no-AI policy after all.
Plus having the video file means I can extract frames as images at specific timestamps.
Homely though the ask feature is probably fit for purpose, at least on YouTube.