Big AI labs are not software companies where payroll dominates expenses. They're capex-heavy industrial entities; it just so happens that the "machines" (whose output they sell) are nominally the same category as the devices that their knowledge worker employees use on their desks.
Anthropic was profitable last quarter.
If Anthropic can block distillations somehow (which are fair game imo given that Anthropic et al did the same with the written works of mankind), then they might stop or slow down the chinese from catching up.
Chinese also have like 40% of the AI researchers of the world, plus they have access to a lot of cheap labour for writing training data. I'm sure an hour of training data creation from one of China's 162 million university educated people is much cheaper than an hour of work from one of US's 97 million. Probably still cheaper than someone from the grand area.
China is behind in AI chips/GPUs but they are catching up. One thing where they have a hard dependence on outside is their energy imports: they have to import a lot of stuff from third party countries. The US on the other hand is energy self sufficient.
It may be that US labs use Chinese models for distillation but we'd ofc never know because they can host the models themselves
If you feed the most recent Github repos into the training, most of that code will be written by frontier LLMs. Training on that is distillation.
It doesn't make sense to compare OpenAIs or Anthropics compute spend to that of our average software company, because different products require different raw materials. Dropbox also use way more storage than Snapchat, that's an equally silly comparison.
If you want to take the DDG LLM summary at fate value, apples are lower in calories and sugar but higher in fiber compared to potatoes, which are richer in vitamins and minerals like potassium and vitamin B6. Overall, apples provide more dietary fiber, while potatoes offer more protein and essential nutrients.
Comparison rarely lead to one obvious all superior option that discard every other considerations.
The comparison no longer starts with the goal to assess distinct objects in the frame of a given more or less established framework, and instead our attention is framed toward challenging ourself. That is, anchored toward finding what frameworks would allow to assess anything meaningful. And latter on, what does frameworks and framework creation reveals about ourself.
Not sure it succeeds in that, but I think that's the intent.
As an aside, I observed my stance on proper English use changing in real time over the past 2-3 years. I used to be proud of writing clinically correct prose, and found mistakes in grammar and vocabulary grating. Now, I kind of welcome them and have stopped caring at all about committing such language crimes. I used to cringe when someone didn't capitalize the first words in their sentences - not anymore. I think we're years away from LLMs convincingly faking human-like mistakes (since all the work currently goes into avoiding them), so it's going to remain a useful signal for a while.
Bandwidth isn't free, and all my life I've been told that piracy is theft.
You should, but with two important caveats. First, you don't know what their amortization schedule is like so you don't know what the impact on the pricing will be (are they going to pass the cost on over 5 years or over 20 years?), and second they may go bust before paying the cost down so they may not get a chance to pass it all on. If someone buys the company then they'll get a discount on the value, which means the training costs are just eaten by the investors.
Assuming their investors win the bet they placed on them. Which isn't given.
Anthropic spends [...] about $2m of compute per employee per year against a likely all-in comp of $500k+.
The rest of the software market trails. The top 1% of companies spend $89k per engineer per year on AI
This framing makes no sense. The reason Anthropic spends so much on compute per employee is that they are building models. Anthropic employees aren't opening Claude Code and spending $2m in inference every year, so comparing it to other software companies, where AI expense is mostly inference, is completely incoherent.
Yes, the cost has to be passed down eventually, but it's not passed down to one company; it's passed down to all of Anthropic's customers, so the actual share of that money will be distributed among Anthropic's clients.
Look, I 100% agree with the idea that OpenAI and Anthropic are both unsustainable companies that have dug themselves so far into a debt hole that, most likely, the only way they'll be rescued is with government intervention, but this is still a terrible article.
Though I agree it might be informative to split it by industry sector.
Compare AI costs per-engineer-salary-dollar, because more expensive engineers probably need more expensive AI.
Let's see how this works out in the long run. For a historical analog, more expensive engineers don't use more expensive computers (by and large).
Your most expensive engineer's time is most valuable, so if you give them standard issue which is half the speed, you are throttling the value you can get from your engineer. Not to mention the mental drain of your cursor barely being able to move due to all the bloated virtual networking systemization.
It would seem to make sense to give more valuable employees faster equipment, so that their time isn't spent toiling with the slow machine, but rather actually producing value.
(Well, for this to make sense, you need good computers to be cheap, but the returns to diminish quickly after that. Which used to be the case until fairly recently: single core performance was pretty much capped, and after your whole compile is parallelised adding more cores to a machine didn't really help your best engineer either. Similarly for RAM.)
My 10 year old desktop sitting next to me is a lot faster, but alas, I can't BYOD.
They don't? If you give your best engineers substandard hardware to work on, you're going to get worse output from them compared to if you give them more expensive computers to work with.
Not completely true. Giving developers hardware that is too beefy is the main reason why so much software breaks down when run on users' machines, which are generally old, on spotty Internet connections, and RAM-starved. Devs just don't need to think about performance unless it's really asymptotically bad, while the users bear the full brunt of inefficiencies.