18,544 karma · joined March 11, 2012
Don't worry, my anger is limitless and I can direct it at both corporations and corporations at the same time.
Managed infra is, by its very definition, not your own infra. It's infrastructure someone else sets up and manages for you.
The only open models that are "almost as good or better in some cases" require massive amounts of RAM. I posit that most people cannot afford a decked out Mac Studio, and therefore run the smaller "flash" variants on more normal devices. The issue is that those are nowhere near frontier-level in terms of capability.
That is not what they are doing. They are calling out specific providers who release powerful models without safeguards.
In addition, said providers are not "building models for free for the public." They are doing it to hamstring America's dominance in AI, primarily by undercutting the frontier labs.
It's worth noting that the overwhelming majority of people who use Chinese models don't do this. Yes, it is nice to have the option, and there are US-based inference providers that claim to not send your data to China and maybe indeed don't, but in the grand scheme of things, we need to remember the adage that became popular during the social media era: if something is free (or, in this case, close to free), you are the product.
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
The issue of course is that if you do invest the time, then you're no longer saving time by using AI. You're just spending it reading and trying to understand something you didn't write. And that can be unpleasant in its own way.
My hot take is that for parts of a system that can be considered its core, forming a deep understanding is almost always important, and so is knowing how the different business domains integrate and where the connection points are. For many others, a high level understanding is sufficient. The difference is that now, with AI, you can make that choice. Before, you had to write everything yourself, and for any sufficiently complex and long-lived system it became impossible to hold all of it in your head.
No, I really am not. I'm doing maybe 10% of what I did before. The rest is filled up by other, usually higher order tasks like planning, product management and work orchestration.
The chatbot we added to our B2B product is by far the most popular addition we've made this year. Our users are not tech-savvy, they use a lot of apps everyday and don't want to have to learn and keep up with just another UI. So they like being able to type their wants and needs in plain language (or speak it into their phone, if they are in the field) and get a plain language response back with embedded images and charts.
YMMV of course.
>> There's a reason why the last known wolf in Denmark didn't just wander off, but was shot dead in 1813.
>> When wolves get out of control, you shoot them.
In the case of Stripe, users want one thing, which is to receive correct answers to their questions nearly instantly. And from the sounds of it, that is indeed the case.
OpenAI's display of incompetence and negligence is absolutely stunning.
>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.