There are only 3 companies doing this to date: Google, Sakana AI and Autohand AI.
12 karma · joined April 19, 2017
There are only 3 companies doing this to date: Google, Sakana AI and Autohand AI.
This made me laugh https://www.forbes.com/sites/annatong/2026/03/05/cursor-goes...
If that's what winning looks like AI Summer is coming!
You can see the turns, tools called, outputs, changes. Similar to what he's trying to achieve.
My view is that software engineering is splitting into two paths. One is deep craftsmanship. The other is guided operators who know how to steer tools, systems, and LLMs toward the right result.
I am keen to hear where others agree or disagree, and how you see this playing out in real teams.
Machine orchestration: stateless execution, structured outputs, designed for CI/CD and batch runs, not just interactive use Auto mode: autohand -p "fix the tests" --yes --auto-commit runs the full task without prompts. Three permission levels plus granular command whitelist/blacklist Skills system: modular instruction packages that activate on demand. Run --auto-skill and it generates skills tailored to your project
One thing that's been surprisingly useful: because it's provider-agnostic, you can prototype with a fast cheap model, then swap to something heavier for the actual run. No code changes, just config. It's TypeScript + Bun, 40+ tools (file ops, full git, semantic search, multi-file edits), sessions persist and resume.
Today we are sharing a deep technical guide on how we built Git Flow Automation in our Evolve platform. It solves a real engineering problem: how to run hundreds of agent tasks in parallel across a codebase without conflicts, without slowing down, and without breaking CI.
Rather than simple scripts or CI tricks, we use Git worktrees to give each agent its own isolated branch and working directory. This lets each agent run tests, detect conflicts, and even roll back individual tasks. We also built conflict detection, AI assisted resolution, and merge strategies that keep history clean and safe for teams.
Here is my take based on 20+ years of experience in Devtools and dealing with lots of large code bases.
The guide walks through:
Why sequential execution fails at scale
How parallel worktree orchestration works
Task lifecycle and dependency ordering
Conflict detection and automatic resolution
Testing per task and rollback controls
Merge strategies and commit hygiene
Safety limits and observability tooling
This is a technical system design share, not a product announcement. I would love feedback from builders and maintainers who work on large codebases or autonomous tooling.
Super interested in read different ideas.
Read the guide here:
A few weeks ago I shared some thoughts on intent weaving in AI coding agents. Today I would like to invite the community to our private beta of Evolve.
Evolve is an AI software engineer designed for long horizon workloads. It uses reinforcement learning style feedback loops to learn from past mistakes, improve future decisions, and reduce repeated human intervention when working on real codebases.
We are early and actively looking for feedback from people who deal with complex systems, legacy code, or long running engineering tasks.
Private beta: https://autohand.ai/products/evolve/
Happy to answer questions and hear what does not make sense.
Today we are warning the entire community, stop using companies providing APIs for coding tasks on your company source code.
I’d love to have different opinions here and create a better version of the future together.
If you want to know more about what I’m doing next www.autohand.ai
Nice job, the scores are superb.
It's not fully functional now as we're building this in the open. Fully supports the 3 major ones, Codex, Claude, Gemini. But I'm also working on a version that doesn't use git worktree, some enterprise folks don't like it. So we're solving this by adding an isolated environment, similar to what others are trying to establish.
We’re building an open stack that lets AI coding agents deliver work with the discipline senior engineers expect. Our latest write-up, “Intent Weaving for AI Coding Agents,” breaks down how we encode strategy, policy, and telemetry into machine-executable intent, plus an honest inventory of where current agents fail (reasoning, repo awareness, testing, etc.).
Highlights: - Mission compiler that turns business objectives into guardrail-rich plans for agents. - Knowledge graph + policy DSL so automation stays inside governance envelopes. - Pain-points matrix from real deployments; new benchmarks that punish regressions, not just pass unit tests. - Open-source pieces as we release them; Commander is already MIT-licensed.
We’d love feedback from folks shipping agentic workflows or wrestling with AI codegen drift. Where should we push harder? What failure modes have we missed?
Link to our manifesto: https://autohand.ai/manifesto
Thanks for reading, and be kind. Creating a new category means stretching before the skills are perfect.