I’m a little confused as to the setup. It was asking each model to one-shot a script and then the scripts faced off? Were the models given a computer environment? Or a test server to iterate against?
You won’t be taken seriously if you push OpenSCAD - it’s simply not a tool that professionals will adopt due to not doing Breps. I think the recent progress of FreeCAD and its spin-off libraries cadquery and build123d will be better to push.
That’s based on Anthropic’s retail price right? Not a fair comparison, like saying that Netflix must be losing money because every movie rental is $4 and a Netflix subscriber can watch 20 movies in a month.
I don’t think anyone has to sell inference below cost. If Anthropic is GPU-constrained, then it makes sense for them to charge much much more on API users and push subscribers towards extra billing, because that’s the only knob they can turn. OpenAI has much more capacity based on news reports.
> Do you think a human is capable of providing assistance with defense but not offense, over a textual communication channel with another human?
> If no, how does a cybersec firm train its employees?
In general, no, humans can’t be sure they are only helping with defensive and not offensive work unless they have more context. IRL, a security engineer would know who they’re working for. If they’re advising Apple, then they’d feel pretty confident that Apple is not turning around and hacking people.
I agree the homepage is a weak sell, but an independent operator in Europe IS value. If it doesn’t really make a difference otherwise, why not choose a home server that is governed by and supports your home region of Europe? (obviously there are other things that would make a difference, but you gotta start somewhere)
With Opus it’s hard to tell what was due to the tokenizer changes. Maybe using more tokens for the same prompt means the model effectively thinks more?
People subscribed to chatgpt before there was codex. Why wouldnt a Claude subscription stand on its own without Claude Code? In fact it’s probably a smart move for Anthropic to split it out.
Their enterprise API customers are literally competing to see who can throw the most money at Anthropic. Anthropic has very little reason to focus on a $20/month user, and with their current momentum (especially since enterprise deals are long-lived) they could remove Claude Code from the Pro plan without any revenue hit. In fact, it may be a huge revenue boost given the strength of the Anthropic brand.
Agreed that it’s 100% marketing. In some ways Anthropic is more of a for-profit corporation than OpenAI which is at least partially owned by a non-profit.
Having too much usage and not enough GPUs is a form of financial pressure, no? since you want to replace your less profitable customers with more profitable customers
Similar to how damaging it is that people believe airplanes can “fly” when they in fact do nothing of the sort. After more than a hundred years of effort we have only managed to mimic flying yet billions of dollars continue to get poured into airplanes.
I think on HN we always underestimate how much momentum matters. Anthropic has so much clout and mindshare that even if they continue burning goodwill and everyone on HN ditches Claude Code and stops recommending it, they will still be revenue leader for years to come. Those enterprise contracts aren’t month-to-month.
My experience has been that this isn’t generally true, mainly because worse models pursue red herrings or get confused and stuck. a better model will get to the correct solution in fewer tokens, and my surface-level understanding of how RL works supports this.
The real question I see nobody asking is how GPT-5.4 beats Opus at a fraction of the price. I doubt it’s only a question of subsidization. My impression from the past is that GPT-5 was around a Sonnet-sized model, and 5-mini was Haiku-sized. At least on my codebase anyways, Codex one-shots tricky things that Opus needs several tries to fully get right.
People make a hobby out of tricking chat apps to leak their system prompt. But I doubt there’s much gain to be had by using this one vs coming up with a custom prompt.