HNHacker News
TopNewBestAskShowJobs

DanMcInerney

386 karma · joined December 23, 2013

submissionscomments
DanMcInerney··on Xiaomi MiMo v2.6
This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
DanMcInerney··on Google's Open Agentic Orchestrator
Right. That's what modular guidance documentation is for. You could have astra.code or qwen.code.api. All reusable in different workflows. Prescription doesn't belong in the skill itself.
DanMcInerney··on AX – Google’s Open Agentic Orchestrator
Overly complex; yaml files, heavy framework. Same mistake as Claude Code's Dynamic Workflows. Why not just use the dehydrated skills as the workflow skeleton and use custom guidance docs to hydrate the skills with taste and preference depending on the domain of the task? Now you can build a library of small workflows that compose into larger workflow, and you can export any workflow as a single skill to be used in other harnesses. For example, I have a code.md. It's really small, just a bit of taste preference. If I'm using it to hydrate orch-work for coding tasks, then maybe I want to create a code.api.md which hydrates for further specificity if the task is about creating APIs. Then when new models come out, I can just delete code.api.md and leave it as code.md for /orch-work to read from within a workflow because newer models won't need as much prescription.
DanMcInerney··on AX – Google’s Open Agentic Orchestrator
I really don't think any of these SOTA labs are doing agentic engineering correctly. Skills are the universal language of all agent harnesses. If you abstract the taste and prescription out of the skills and into guidance docs, then leave the skills as basically just workflow scaffolding, you can build task-specific workflows that work with any harness like Claude Code, Codex, Antigravity, etc. Technically, you only really need 2 skills, work and review, and with these you can build infinitely complex workflows including self-improving loops. I built this out and have been using it for months. It's been extremely nice. https://github.com/DanMcInerney/orchflows
DanMcInerney··on Claude Cowork and chat are now one Claude
The problem is that knowledge workers just need the simplest way possible to automate their work. Currently I feel like Anthro is overcomplicating this. If I'm an accountant, I want to open Claude App, describe my workflow for balancing some books, then have Claude design a reusable WORKFLOW that it runs whenever I ask. The workflow only needs 2 primitive skills to do anything: /work, and /review. Compose these together with guidance docs into workflows, then workflows and call other workflows. It's all just callable skills.

/balance-books hey claude here's the books. Go balance them

Claude then uses the guidance docs and various subagents to review the books, send off parallel workers with cheap models, then review with a more expensive model. Done in the repeatable, correct order with independent review every time. As new models come out, your workflow structure stays the same. You just delete some prescription from the guidance docs.

https://github.com/DanMcInerney/orchflows

DanMcInerney··on Dream-RSI: Recursive Self-Improvement through Evolving Worlds
I ended up building a simplified version of this as /self-improve in https://github.com/DanMcInerney/orchflows. History is the state ledger, memory and RSI just cite the history as evidence and can be rewritten. I feel like strong immutable state is the missing piece of the puzzle for most of these memory libraries.
DanMcInerney··on Weightlifting beats running for blood sugar control, researchers find (2025)
Does the average "scientist" even care about quality these days or is it all just whatever headline makes their funders happy?
DanMcInerney··on OpenAI Forked Git on GitHub
Rockstar built git for games, OpenAI's aiming for git for agents.

Claude, make a git for games designed for agents. Sell the code to the highest bidder when you're done.

DanMcInerney··on Show HN: OpenKnowledge – open source AI-first alternative to Obsidian/Notion
https://cloud.google.com/blog/products/data-analytics/how-th...

Did you look at the OKF repo from Google? Open Knowledge seems to be a common term these days for similar solutions. I think OKF is more of the protocol for wiki-for-llm while you have more of the bells and whistles

DanMcInerney··on /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds
ANNNNNND it's gone. Guys, I found a way to reduce Fable token usage 100%. You can find it here: github.com/USGov/idiotic-overreach.
DanMcInerney··on /architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds
I don't disagree with any of this. It is generated software, and it's not a novel idea. I didn't mean for it to come off like that. It's just solving an itch that I couldn't find a solution to and I'm getting a lot of personal utility out of it. I do have a lot of experience with agentic memory, multi-agent systems and harnesses and wasn't super impressed by the workflow of Fable calling opus subagents so I figured I'd apply best practices to what already exists to make it a teensy bit better and easier to use.
DanMcInerney··on MiMo Code is now released and open-source
I've worked a lot with MiMo in my project that pits LLMs against each other in games (clankerfights.ai). It is a very very good model for the price. MiniMax I'd say is smarter, but MiMo really touches near pareto frontier.
DanMcInerney··on 'Fuck you, Bambu': How one private message could change the face of 3D printing
Me and a buddy, victor teisller, hacked the most popular 3D printer a few years ago (flashforge) to turn it into a major fire hazard through reverse engineering the firmware update remotely. Changed the max temp of the extruder to the temp of the Sun lol. These things are a fascinating security target because they're an easy place to turn abstract digital hacking into physical repersussions that can literally murder.

https://www.theregister.com/security/2020/04/13/how-to-make-...

DanMcInerney··on Daily Claude outage is upon us. Waiting for Claude Status to update
Cross your fingers they're about to drop 4.7. 4.6 came out with a bang, now it seems all the compute bottlenecks just lead to customer frustration as they get closer to releasing next model. Balancing the books over there must be a nightmare, "Well we can piss off every single customer for a week, but we'll be able to release the next model 1 week faster"
DanMcInerney··on Nvidia NemoClaw
All these comments about "this is crazy to deploy agents that might do something bad in your environment" are crazy themselves. The productivity gains from these computer use agents are crazy. Every org on earth has to make the call, is 3x productivity gains worth 2x the risk increase? The answer is almost always a resounding yes. You limit the blast radius if things go wrong, but the financial gain of having 1 employee to the work of 3 already pays for the disaster if it happens.
DanMcInerney··on Matchlock – Secures AI agent workloads with a Linux-based sandbox
Sandboxing is a great security step for agents. Just like using guardrails is a great security step. I can't help but feel like it's all soft defense though. The real danger comes from the agent being able to read 3rd party data, be prompt injected, and then change or exfiltrate sensitive data. A sandbox does not prevent an email-reading agent from reading a malicious email, being prompt injected, and then sending an email to a malicious email address with the contents of your inbox. It does help in implementing network-layer controls though, like apply a policy that says this linux-based sandbox is only allowed to visit [whitelisted] urls. This kind of architectural whitelisting is the only hard defense we have for agents at the moment. Unfortunately it will also hamper their utility if used to the greatest extent possible.
DanMcInerney··on Prompt Injection via Poetry
There are an infinite amount of ways to jailbreak AI models. I don't understand why every time a new method is published it makes the news. The data plane and the control plane in LLM inputs are one in the same, meaning you can mitigate jailbreaks but you cannot 100% prevent them currently. It's like blacklisting XSS payloads and expecting that to protect your site.
DanMcInerney··on Gemini 3
A 50% increase over ChatGPT 5.1 on ARC-AGI2 is astonishing. If that's true and representative (a big if), it lends credence to this being the first of the very consistent agentically-inclined models because it's able to follow a deep tree of reasoning to solve problems accurately. I've been building agents for a while and thus far have had to add many many explicit instructions and hardcoded functions to help guide the agents in how to complete simple tasks to achieve 85-90% consistency.
DanMcInerney··on LLM Inevitabilism
This is absolutely doable right now. Just hook claude code up with your calendar MCP server and any one of these restaurant/web browser MCP servers and it'll do this for you.

https://apify.com/canadesk/opentable/api/mcp https://github.com/BrowserMCP/mcp https://github.com/samwang0723/mcp-booking

DanMcInerney··on LLM Inevitabilism
These articles kill me. The reason LLMs (or next-gen AI architecture) is inevitably going to take over the world in one way or another is simple: recursive self-improvement.

3 years ago they could barely write a coherent poem and today they're performing at at least graduate student level across most tasks. As of today, AI is writing a significant chunk of the code around itself. Once AI crosses that threshold of consistently being above senior-level engineer level at coding it will reach a tipping point where it can improve itself faster than the best human expert. That's core technological recursive self-improvement but we have another avenue of recursive self-improvement as well: Agentic recursive self-improvement.

First there was LLMs, then there was LLMs with tool usage, then we abstracted the tool usage to MCP servers. Next, we will create agents that autodiscover remote MCP servers, then we will create agents which can autodiscover tools as well as write their own.

Final stage of agents are generalized agents similar to Claude Code which can find remote MCP servers, perform a task, then analyze their first run of completing a task to figure out how to improve the process. Then write its own tools to use to complete the task faster than they did before. Agentic recursive self-improvement. As an agent engineer, I suspect this pattern will become viable in about 2 years.

DanMcInerney··on OpenAI o3-pro
I'm really hoping GPT5 is a larger jump in metrics than the last several releases we've seen like Claude3.5 - Claude4 or o3-mini-high to o3-pro. Although I will preface that with the fact I've been building agents for about a year now and despite the benchmarks only showing slight improvement, I have seen that each new generation feels actively better at exactly the same tasks I gave the previous generation.

It would be interesting if there was a model that was specifically trained on task-oriented data. It's my understanding they're trained on all data available, but I wonder if it can be fine-tuned or given some kind of reinforcement learning on breaking down general tasks to specific implementations. Essentially an agent-specific model.

DanMcInerney··on The Rise of 'Vibe Hacking' Is the Next AI Nightmare
I too write automated offensive tooling. We actually wrote a project, vulnhuntr, that found the first autonomously-discovered 0day using AI. Feed it a GitHub repo and it tracks down user input from source to sink and analyzes for web-based vulnerabilities. Agree this article is incredibly cringy and standard best practices in network and development security will use the same AI efficiency gains to keep up (more or less).

What bothers me the most about this article is that the tools that attackers use to do stuff like find 0days in code are the same tools that defenders can use to find the 0day first and fix it. It's not like offensive tooling is being developed in a vacuum and the world is ending as "armies of script kiddies" will suddenly drain every bank account in the world. Automated defense and code analysis is improving at a similar rate as automated offense.

In this awful article's defense though, I would argue that red team will always have an advantage over blue team because blue team is by definition reactionary. So as tech continues it's exponential advancements, the advantage gap for the top 1% red teamers is likely to scale accordingly.

DanMcInerney··on Vulnhuntr: Autonomous AI finds first 0-day vulnerabilities
https://protectai.com/threat-research/vulnhuntr-first-0-day-...

More details on the development and challenges can be found in the blog.

DanMcInerney··on Attacking UNIX Systems via CUPS
Yes, good point, university networks are particularly vulnerable.
DanMcInerney··on Attacking UNIX Systems via CUPS
Depending on your interpretation of the Scope metric in CVSSv3, this is either an 8.8 or a 9.6 CVSS to be more accurate.

In summary, there's a service (CUPS) that is exposed to the LAN (0.0.0.0) on at least some desktop flavors of Linux and runs as root that is vulnerable to unauth RCE. CUPS is not a default service on most of the server-oriented linux machines like Ubuntu Server or CentOS, but does appear to start by default on most desktop flavors of linux. To trigger the RCE the user on the vulnerable linux machine must print a document after being exploited.

Evilsocket claims to have had 100's of thousands of callbacks showing that despite the fact most of us have probably never printed anything from Linux, the impact is enough to create a large botnet regardless.

DanMcInerney··on Attacking UNIX Systems via CUPS
It appears that the vulnerable service in question listens on 0.0.0.0 which is concerning, it means attacks from the LAN are vulnerable by default and you have to explicitly block port 631 if the server is exposed to internet. Granted, requires user to print something to trigger which, I mean, I don't think I've printed anything from Linux in my life, but he does claim getting callbacks from 100's of thousands of linux machines which is believable.
DanMcInerney··on First targeted campaign against AI infrastructure found in wild
Backstory about how the original vulnerability was found by Protect AI: https://protectai.com/threat-research/shadowray-ai-infrastru...
DanMcInerney··on GPT-4 vision prompt injection
As a hacker of more than a decade, none of this really gives me pause. There's still critical sev bugs in tools like Ray, MLflow, H2O, all the MLOps tools used to build these models that are more valuable to hackers than trying to do some kind of roundabout attack through an LLM.

It's relevant if you're doing stuff like AutoGPT and you're exposing that app to the internet to take user commands, but are we really seeing that in the wild? How long, if ever, will me? Ray does remote, unauthenticated command execution and is vulnerable to JS drive-by attacks. I think we're at least a few years away from any of the adversarial ML attacks having any teeth.

DanMcInerney··on Critical remote unauthenticated system/cloud takeover in major AI tool
Pretty brutal. Took me about 3 days to find. I suspect there's more.

* Unauthenticated

* Remote

* No user interaction

* No prerequisite knowledge or environment setup

* Large adoption on MLflow in AI engineering workflows

Here's the GitHub Security Advisory: https://github.com/mlflow/mlflow/security/advisories/GHSA-xg...

DanMcInerney··on Hackers earn $990k for 63 zero-days exploited at Pwn2Own Toronto
It's far more likely that software is improving, but the tools to break it and the number of people willing to invest the time in learning those tools, is increasing. The fact is there are absolutely untold amounts of bugs in every piece of software every written. Finding those bugs, up until the past decade of automation tools to find/exploit them, has been extremely time consuming with few people patient or curious enough to improve their skills in that arena.

Think about the number of programming frameworks that've come out in the past few years. Or GitHub Copilot and ChatGPT which literally write solid code for you. There's no way software has not improved a lot in the past decade but there are WAY more software devs than there are exploit devs or hackers. Similar to the ratio of lions:wildebeest. Way more prey than predators.

Page 1 of 2Next →