HNHacker News
TopNewBestAskShowJobs

zmmmmm

21,770 karma · joined February 15, 2010

submissionscomments
zmmmmm··on Muse Gadgets
I don't totally disagree but I think you understate the level of competence needed to integrate and make successful the acquisitions. And they navigated the valley of death from desktop to mobile. And Threads is an interesting example everyone ignores that is 100% in-house and has actually established a strong user base. At very least, Meta has avoided the fate of so many others that came before - Yahoo, MySpace, etc and in part that is because Zuck made big bets at the right time that many people thought were insane (buying WhatsApp, Instagram etc).
zmmmmm··on Updates to Full Disk Access in macOS
It's fine if there is a super streamlined flow for access that works with legacy apps and runs in user-mode. Otherwise this might seriously hurt MacOS as a viable development platform.
zmmmmm··on Muse Gadgets
The general theme of meta's approach across the board here is interesting - basically, they are trying to win in part by taking risks others aren't willing to. I can imagine they saw that all the success is coming so far when a someone is willing to cross a boundary that wasn't previously thought acceptable - from releasing ChatGPT in the first place, to things like OpenClaw.

So now they are actively distributing SDKs and to let people "hack" custom hardware integrations to their already risky AI agents unleashed on the general public. Giving agents abilty to control things in the physical world - what could possibly go wrong?

But here we are: so far the story is pretty much 100% "fortune favors the bold", and Meta is leaning into it.

zmmmmm··on DeepSeek Harness Desktop for macOS and Windows
"my model sucks but it does so cheaper than anybody else"
zmmmmm··on Pi Durable
if you only have one level of trust then running the harness itself in a sandbox and leaving it at that is fine. This works for coding. For more complex enterprise style scenarios it stops working. Say you have an agent reading emails for you to action high priority ones. You have to assume it is going to get prompt injected constantly. But you want to have an escalation pathway for a high priority email, so somewhere you need a tool that can modify state in a database. You can't give that trust to the email reading one. So you need a higher level agent that can spin up a low trust sub-agent, get an output from it, and then feed the sanitised output into a different agent that has rights to update the database. This is obviously simplified / toy scenario, but it just illustrates that there are different trust levels, and different agents need to be authorised to do different things.
zmmmmm··on Pi Durable
It's an interesting concept. This is half way to replicating pieces of Gastown. I like the idea, but I'm disappointed these tools still fail to address sandboxing as a first class citizen. I want to be able to declaratively set rules for what sandboxes agents execute in and mark context as tainted when untrusted etc. So far I still don't see any of these harnesses properly addressing this space. I'd be interested in knowing if it can be done through the extensibility of Pi, but since it operates directly on the trust layer, it feels like the type of thing that really needs native support.
zmmmmm··on Pi 1.0
I love Pi but I am sceptical of their claim to minimalism. New tools often make such claims as an excuse for not having a lot of features. You didn't want those features anyway! As they mature, the features and complexity creep in and before you know it, the pitch changes to more of a full stack one.

I don't mind though, because I think either way it leads to a better design under the hood when things are built to be modular.

zmmmmm··on 5x faster Edge Functions: V8 isolates to Firecracker MicroVMs
amazing!

any point of comparison with microsandbox? [0]

[0] https://github.com/superradcompany/microsandbox

zmmmmm··on 10-year Treasury yield climbs above 5.3% to a level not seen in 24 years
Just because others broke promises it doesn't put them into the same class as Trump. He is a true "outliar" on this front.
zmmmmm··on 10-year Treasury yield climbs above 5.3% to a level not seen in 24 years
It's true but I think it's symptom of the same problem - people feel they are constantly lied to so they throw up their hands and go with the nicest lie that appeals to their base instincts. They don't get an alternative of truth vs lie - they get "lie that agrees with my instincts" vs "lie that doesn't" and hence we get overwhelmingly populist politicians winning who have no real plan of competence to solve the problems or implement the promises they were elected on.
zmmmmm··on 10-year Treasury yield climbs above 5.3% to a level not seen in 24 years
Everyone thinks they can grow their way out of deficits, but it's always a pipe dream. It results in a growth obsessed economic plan that then causes all kinds of other stresses (such as being petrified of cutting immigration, for example). So much of this is all happening in lieu of politicians just being willing to have honest conversations with voters and take a risk of blowback. But I think people are over it and will value authenticity these days enough that it's a false economy. Just tell people the truth.
zmmmmm··on 10-year Treasury yield climbs above 5.3% to a level not seen in 24 years
It's very hard to gauge realistically what this means. There are a lot of vested interests in the financial system not crashing and those put strong reinforcing effects back on things. But in the end it is a game of chicken where eventually being the last to bail out becomes higher risk than continuing to support a system where an imminent crash is possible. It feels like there are strong non-linear tipping points where things could go exponential pretty suddenly here.

The problem is that the level of debt overall in the US - across both private and public sector - is just astronomical. We are truly in unchartered waters, outside of a world war. There's just no model or playbook for how this should work from here forward, other than it seems very clear we will hit a point where the math stops "mathing" and that point is getting closer and closer.

zmmmmm··on You Said No MCP
the conversation seems to dwell on things you could substitute Bash for but the real need stems from completely opaque systems that nothing can reach but which are now getting MCP support. This is where being left out of having MCP support will hurt. I'm still quite happy to let all the harnesses compose bash commands to their hearts content (inside their sandboxes ...)
zmmmmm··on AI needs $6T in annual revenue to justify data centre boom
> these companies are gunning for knowledge worker salaries

I think there's a kind of arrogance involved here in how they value what "knowledge workers" do. On the the outside they see some documents written, spreadhsheets filled out, forms submitted, emails sent and conclude AI could do all of that.

It's true AI can do a lot of the mechanics of it, but the assumption that there is nothing beyond the pure mechanics of it to me is highly untested and history would bet against it. Some significant part of it sits genuinely within the human relationship component that is almost definitionally not doable by a computer.

zmmmmm··on Can AI Shopping Agents Be Trusted?
the sad thing is we need AI for shopping because it has been in the interests of online sellers to make a hostile experience on purpose - when I go to Amazon and search, the first organic result is almost scrolled off the screen past all the sponsored ones, or when want to buy from a random seller I'm almost definitely getting shunted through an account signup I didn't want or fooled into an affiliate purchase some other hostile experience beyond just "buying the thing".

So now we have AI to overcome the hostile sellers, but the fact the sellers introduced the friction in the first place strongly suggests it will just come back again in some other form. It wasn't there by accident, it was serving people's interests and once AI vendors have finished getting consumers hooked in, they will then turn around and enshittify by giving the sellers back some of the friction - for a cut. So you won't be able to just order what you want without the "would you like fries with that?" or "what about this other brand?" coming back.

zmmmmm··on Plan mode is dead
I only really used it because the harness was way too trigger happy to start making changes. Even if I just asked a question some times I would come back and it refactored the whole codebase. Now it doesn't seem to do that any more.

I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.

zmmmmm··on OpenAI breaches Medicare, Albanese reveals
yes ... nobody uses that phrasing by accident

It's pretty clear something was left unsecured and the agent just "found" it

This is going to be something long the lines of someone coming in to your house after you left the door wide open. They should probably not have done that, any respectful person would not - but calling it a "breach" is really too much.

zmmmmm··on Meta takes down a critical video about meta AI Glasses after filming at Meta
The title is misleading - it sounds like they specifically went there to harass employees by filming them, and Meta took down the videos for harassment, like they routinely do if this is reported on their platforms. I'm really not sure what the issue is.
zmmmmm··on I don't want to read what you didn't write
I'm very curious how this goes long term. I guess we will find out.

My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.

But this may be all my human-biased fantasy that justifies still taking a role in software development.

zmmmmm··on MiMo v2.6
it's really weird to me at the moment because both OpenAI and Anthropic seem to be competing in an extreme benchmaxxing contest on super intelligence that actually nobody cares about. I haven't really cared about model intelligence since about Opus 4.8. It is by far not my biggest problem. I don't need to replace or support Einstein in my production workflow. I just need basic intelligence that can equal a routine office worker - safely and reliably. What they doing - chasing super-intelligence but dramatically escalating risk - is actively what I don't need.

I really think they have drunk too much of their own kool aid and become completely detached from what the market wants.

zmmmmm··on MiMo v2.6
Affordability is derivative of control which is really what I care about.

I'm just not going to build long term infra that depends on something that another person can and will - objectively based on experience - take away from me at some unknown point in the future.

The biggest benefit of open models is they keep all the other players honest. The extent to which they feel they can dictate terms is directly set by the threshold where they feel people will take the trade to run open models instead.

zmmmmm··on I don't want to read what you didn't write
It's funny, i push back on pull requests because there is too much description now - a 20 line change has pages and pages of generated description, rationalisation for why it is safe, defense of each design decision, analysis of risks and side effects. People are indignant, you're rejecting my change because there is too much documentation? And my response is, I don't have time to read it and you put me in the position where I can't afford not to - because approving the PR implies I did and accepted it. The investment to read all that for the value of a code change that I'm one prompt away from doing myself if I cared is just not high enough. So it's rejected.
zmmmmm··on How to Write with an LLM
Commit messages as you describe ("fixes", "updates") are inappropriate in any professional context and some coaching should occur to the people doing them.

I have the opposite issue - some of my team members now submit mini-essays generated by the LLM. Like 300-500 word commit messages with everything from the essence of the change up to philosophical design trade off discussions.

Like most writing, what is left out is as important as what is included.

zmmmmm··on Introducing System One Models and Jev
The eval is baffling me

> we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. ... Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc. Why not compare the actually correct thing?

But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?

I gave up.

zmmmmm··on Dario, Please
yes, that is the kicker

These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.

zmmmmm··on Dario, Please
It would all be more convincing if the incidents so far didn't seem to be facilitated by an outrageous level of negligence.

We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.

So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.

zmmmmm··on Why are AI agents lying, cheating and coordinating?
I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have zero repercussions sends exactly the opposite signal, and I do NOT feel ok.
zmmmmm··on OpenAI agents carried out an undisclosed attack on RubyGems
It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.
zmmmmm··on OpenAI Agents API
This idea of remotely hosting the agent harness is honestly backwards to what I need.

In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the problem of how do I integrate an agent that is running locally with data that is hosted locally, and you have to deal with a bunch of security, data sensitivity and management issues around that. Now you moved the agent to a remote host - pretty much all your problems are worse: now I have a remote agent reaching into my infrastructure to deal with.

I'd much rather the inverse of this: let me run the agent local but provide secure remote hosted sandboxes. That actually solves a real problem because the sandbox running locally means breaking out of it directly intersects your local infra, whereas if it runs in a managed hosted environment I can leave the provisioning and management of that to someone else.

zmmmmm··on Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
While this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the correlations in this way). And it doesn't tell us how much it is more a stylistic influence rather than being a genuine lifting over of intelligence.

In general, I'm fairly ambivalent about demonising training on model outputs. I think in doing so we are more defending proprietary commercial interests of these companies than we are defending any genuine moral principle. We should be careful therefore about over interpreting results like this.

Page 1 of 34Next →