21,770 karma · joined February 15, 2010
So now they are actively distributing SDKs and to let people "hack" custom hardware integrations to their already risky AI agents unleashed on the general public. Giving agents abilty to control things in the physical world - what could possibly go wrong?
But here we are: so far the story is pretty much 100% "fortune favors the bold", and Meta is leaning into it.
I don't mind though, because I think either way it leads to a better design under the hood when things are built to be modular.
any point of comparison with microsandbox? [0]
The problem is that the level of debt overall in the US - across both private and public sector - is just astronomical. We are truly in unchartered waters, outside of a world war. There's just no model or playbook for how this should work from here forward, other than it seems very clear we will hit a point where the math stops "mathing" and that point is getting closer and closer.
I think there's a kind of arrogance involved here in how they value what "knowledge workers" do. On the the outside they see some documents written, spreadhsheets filled out, forms submitted, emails sent and conclude AI could do all of that.
It's true AI can do a lot of the mechanics of it, but the assumption that there is nothing beyond the pure mechanics of it to me is highly untested and history would bet against it. Some significant part of it sits genuinely within the human relationship component that is almost definitionally not doable by a computer.
So now we have AI to overcome the hostile sellers, but the fact the sellers introduced the friction in the first place strongly suggests it will just come back again in some other form. It wasn't there by accident, it was serving people's interests and once AI vendors have finished getting consumers hooked in, they will then turn around and enshittify by giving the sellers back some of the friction - for a cut. So you won't be able to just order what you want without the "would you like fries with that?" or "what about this other brand?" coming back.
I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.
It's pretty clear something was left unsecured and the agent just "found" it
This is going to be something long the lines of someone coming in to your house after you left the door wide open. They should probably not have done that, any respectful person would not - but calling it a "breach" is really too much.
My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.
But this may be all my human-biased fantasy that justifies still taking a role in software development.
I really think they have drunk too much of their own kool aid and become completely detached from what the market wants.
I'm just not going to build long term infra that depends on something that another person can and will - objectively based on experience - take away from me at some unknown point in the future.
The biggest benefit of open models is they keep all the other players honest. The extent to which they feel they can dictate terms is directly set by the threshold where they feel people will take the trade to run open models instead.
I have the opposite issue - some of my team members now submit mini-essays generated by the LLM. Like 300-500 word commit messages with everything from the essence of the change up to philosophical design trade off discussions.
Like most writing, what is left out is as important as what is included.
> we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. ... Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).
They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc. Why not compare the actually correct thing?
But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?
I gave up.
These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.
We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.
So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.
In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the problem of how do I integrate an agent that is running locally with data that is hosted locally, and you have to deal with a bunch of security, data sensitivity and management issues around that. Now you moved the agent to a remote host - pretty much all your problems are worse: now I have a remote agent reaching into my infrastructure to deal with.
I'd much rather the inverse of this: let me run the agent local but provide secure remote hosted sandboxes. That actually solves a real problem because the sandbox running locally means breaking out of it directly intersects your local infra, whereas if it runs in a managed hosted environment I can leave the provisioning and management of that to someone else.
In general, I'm fairly ambivalent about demonising training on model outputs. I think in doing so we are more defending proprietary commercial interests of these companies than we are defending any genuine moral principle. We should be careful therefore about over interpreting results like this.