HNHacker News
TopNewBestAskShowJobs

alanfuNZ

2 karma · joined October 27, 2024

submissionscomments
alanfuNZ··on Ask HN: How do you gate an autonomous coding agent's shell access?
Agree on the sandbox side. dagger or a microVM is the right blast-radius layer, and moving creds out of the environment entirely (service accounts, WIF, an egress proxy that injects them) is the real fix for the secret-then-outbound-call case. Obscuring the string is weaker than it looks: the agent can base64 it, split it, or just describe it, and the sandbox has no way to know that the request body was derived from what it read ten minutes ago.

That's why I ended up gating at the tool-call level with session state instead of at the network level. Same curl gets a different verdict depending on whether the session already read something secret-shaped. Deterministic rules only, and the approval prompt times out to deny, logged as a timeout rather than a denial so I can tell the two apart later.

On the policy-model approach: I'd be careful about what a model's truthy value is allowed to do. A classifier crossing a threshold is a guess with a confidence attached, and a guess that hard-blocks real work gets the whole guardrail disabled by the end of the day. The split I settled on is that a probability can ask (escalate to a human) but only a predicate can refuse. Curious whether you're letting the model produce the deny directly, or routing its output through a human when it's uncertain.

For what it's worth this is what I've been building: https://github.com/DobermanCore/Doberman-Core. Apache-2.0, 100% open source

alanfuNZ··on Show HN: Doberman: The AI watchdog that stops Claude from deleting your database
This only works in the case of working with databases, also a lot of the times people are not careful with the permissions that are given to these agents and they oftentimes touch areas that they were not originally meant to touch. For example I use claude code from Frontend UI, to backend development to even marketing. The whole point of agents is that it is active rather than passive, making it read only essentially (imo) defeats the purpose of using an agent in the first place.

I'm not saying we shouldn't have built in precautions and permissions, this should be standard practices, but Doberman is a system built to handle the inevitable case when these precautions fail.

alanfuNZ··on Show HN: Doberman: The AI watchdog that stops Claude from deleting your database
Also the intro of the readme was written by me, the rest was meant to just be something agents ingest and use to set up Doberman hahaha, didn't think anyone would read the full thing.
alanfuNZ··on Show HN: Doberman: The AI watchdog that stops Claude from deleting your database
Most of the other frameworks sit on the network peremiter and only inspects the prompts. The biggest competition right now that I see in the same lane Cisco's Defenseclaw which focuses on breadth and has relatively loose security guardrails.

Each piece of tech of the framework is not novel, since I'm not trying to reinvent the wheel here. My focus is on the depth of security and a set of invariants that the system is built around. It's designed to be overprotective sitting on the execution path fail closed and raise only. The adaptive learning lowers that security based on your needs so eventually it can become a silent protector that doesn't constantly bother you like CC asking you for permission to do a git push.