26 karma · joined October 30, 2025
I only leave it truly unattended if it's working on a very tight improvement loop, for everything else I'm still checking in on it between working on other things.
I still find agents need a lot of guidance and steering to produce the kind of work I want, but auto mode in a strong sandbox is very useful to me to take a bite out of that. It's particularly good for exploring problems experimentally - where most of the exploration might be thrown away after settling on a solution.
Having used it sandboxed, I wouldn't dream of letting auto mode run outside it. It's very creative at trying to work around the constraints of the sandbox (legitimately, not to escape it) and the classifier for auto mode seems very permissive with the right context.
For context, I'm using per-project VMs with very limited egress and restrictive mounts. Self-built tool to glue it all together, currently unreleased. There has been an explosion of sandboxing tools recently, none of which was quite what I wanted. Heavily inspired by Gondolin <https://earendil-works.github.io/gondolin/>, but fat long-lived VMs.