Because he should know better? Because it’s obviously a shit show but he keeps on being very vocal about his shit show? Because it’s annoying to have to see yet another delusional vibe coded project being hyped up instead of this forum being used to discuss actually industry relevant information?
It seems pretty transparent that they are heavily resource constrained, (training run for Claude 5.x, higher usage / growth than anticipated). I don’t disagree that their long play is monopolistic pricing, but what we’re observing seems better explained by the fact they have a very tight compute budget they are trying to optimize over to put as much as they can into next gen experiments / training to make sure they stay competitive over the next 6-months / year.
The limitation is efficiency and efficacy. If you have to add an additional layer of inference to any request you’re negatively impacting your bottom line so the companies, which are compute bound, have a strong incentive to squeeze everything into a single forward pass. It’s also not clear that a separate model that is smaller than the main model will perform better than just training the main model to detect prompt injection. They are both probabilistic models that have no structural way of distinguishing user input from malicious instructions.
If you’re getting prompt injected, you have skipped right passed thinking critically about what you’re doing and into the same level of intellectual dishonesty as cheating, ie, not learning the thing and then attempting to still attain a grade for work not done.
But debugging mutable state is much easier than debugging a distributed system. Even in C if some global gets mishandled I can just use: gdb, dtrace, strace or even just look at a core dump and know that whatever caused the problem will be discoverable. I have no such guarantee debugging an issue across a distributed systems service boundary.
It seems entirely logical that if an llm allows each ic to do the work of a 2-3 person team (debatable, but assume it’s true for the sake of argument) then you’ve effectively just added a layer to the org chart, meaning any tool that was effective for the next scale up of org becomes a requirement for managing a smaller team.
What should give anyone pause about this notion is that historically, by far the most effective teams have been small teams of experts focusing on their key competencies and points of comparative advantage. Large organizations tend to be slower, more bureaucratic and less effective at executing because of the added weight of communication and disconnect between execution and intent.
If you want to be effective with llms, it seems like there are a lot of lessons to learn about what makes human teams effective before we turn ourselves into an industry filled with clueless middle managers.
3.5-plus was also only available via api. I don’t know what the long term business model for open weights is, I hope there is one, but it seems foolish to assume that companies will be willing to spend millions of dollars of compute on an asset worth zero in perpetuity.
To a certain extent, I do wonder if just letting claude do everything and then using the bug reports and CVE’s they find as training data for an RL environment might be part of the plan. “Here’s what you did, here’s what fixed it, don’t fuck up like that again"
I’ve heard this take before, but if you’ve spent any time with llm’s I don’t understand how your take can be: “I should just let this thing that makes mistakes all the time and seems oblivious to the complexity it’s creating because it only observes small snippets out of context make it’s own decisions about architecture, this is just how it does things and I shouldn’t question it.”
I don’t know that those people were exactly out of a job though, they didn’t do that job, but I find it hard to believe that any of the people solving orbital mechanics by hand wound up with nothing to do but twiddle their thumbs for the remainder of their lives. Similarly, I don’t know that there’s any realistic prospect, even if ai winds up writing all the software, that there wont also be incentive to have people that also understand it.
Once upon a time, I heard someone tell me a fairytale about this thing called a ‘law' and they said that laws could be used to enforce compliance with standards across an entire country. Pure fantasy I know, but a man can dream.
Presumably, there must be some point in time where the bill is made public in some form before going to a vote. If you could get the right tool in the hands of a journalist to turn whatever obscure format it’s in into something legible by an ordinary person there’s probably value there.
I think a lot of people will do this, it remains to be seen how the actual economics of this shake out in the long run, especially considering, it’s not like the existing vendors are going to remain static.
Once the hardware to run inference for something like the vision understanding module of this can be run on a low / medium power asic drones are going to be absolutely horrifying weapons.
I don’t agree, what people want is very consequential, because those people are paying customers of a service, if they aren’t happy with it they have every right to complain.
People should be vocal about what they do and do not think is reasonable behavior by corporations and then act based on those opinions with their wallets. Lord knows we have precious few other ways of influencing corporate behavior.
I think this is disingenuous, people want to be able to use a tool that they pay for to do useful work on their own terms because they payed for it and don’t see the differential pricing model offered by Anthropic as legitimate.
Good idea, the problem is that LEAN only proves what you tell it to prove. Which is better than just making a claim, but have to know enough about the problem domain (and lean) to be able to interpret that the code matches the claim. Otherwise you can be proving something only tangentially related. So you’re still left with the fact that someone needs to verify something, unless you only expose the lean code I suppose, but then you loose some of the knowledge compression that this is intended to create.
You can guide humans, but ultimately the reason senior software developers have been payed large sums of money is that even with specs mostly we have found it works better to have someone with good judgement actually doing the work, otherwise we would have just been using specifications. The question remains open if llm’s can show good judgement, often my experience with claude is that it doesn’t if the problem domain is non-trivial but it’s possible that won’t always be true.
You might be interested in this paper: https://arxiv.org/abs/2505.05522, in essence they demonstrate that a novel architecture that incorporates bounded convolution at the block level that has some variable time horizon on the number of iterations through the convolutional loop can be very effective and solves problems in a way that is much more similar to how humans do.
> but simultaneously massively reduce the cost of that code complexity.
Citation needed. Until proven otherwise complexity is still public enemy #1. Particularly given that system complexity almost always starts causing most of its problems once a project is further along I don’t think we will know anything meaningful about that statement for at least a year.
A reasonable estimate for the Russo Ukrainian war is that there have been half a million casualties due to drones. I would not recommend looking for the videos, many tens of thousands of those have live footage of them occurring.
I don’t understand why more people aren’t focused on how to get the benefits of ai but on your own machine. If the last 20 years of software transitioning off of our desktops and into the cloud has taught us anything, it’s that letting corporate entities run the software you rely on end to end gives you: worse software with more bugs, surveillance and subscriptions. Why on earth would you want that for everything you do.