HNHacker News
TopNewBestAskShowJobs

coder-pm

30 karma · joined June 20, 2026

Marcin Polak

Experienced software architect, recently associated with a GenAI startup.

Cleat, a Docker sandbox for AI coding agents (MIT):

  https://cleat.sh

  https://github.com/cleatdev/cleat
Blog:

  https://cleat.sh/blog
dev.to:

  https://dev.to/coder-pm
Not a native speaker but keen to share my knowledge and experience.

In case you want to contact me, you can easily Google me by my username.

submissionscomments
coder-pm··on Pi.dev: You Said No MCP
Hmm what is this JavaScript sandbox? container, separate process, same Node process? Also does it have access to the MCP tokens? The harness credentials?
coder-pm··on Show HN: Socks Proxy for AI Agents
You should go for it, the egress is becoming something more and more important these days!
coder-pm··on Why does mathmain need an encrypted loader?
Being new here doesn't mean I'm a bot. I'm not native and maybe my style looks for you like a bot, I won't try to argue with you. Just wanted to contribute
coder-pm··on Software sandboxing: The basics (2025)
Nice setup, I'm curious about the PAT. Even fine-grained permissions are working till the expiration so if the box will die someone has it for the whole time. Did you rotate it per session or it was on default expiration time?
coder-pm··on SCH: An affordable sandbox for Coding Agents in your AWS account
It makes sense for a personal project. The fix should be accomplished as a standalone role or scoped credential per user. Once this is done the prefix will be a real boundary, not just a name. Good luck with that, nice project!
coder-pm··on Show HN: Authorize MCP tool calls without giving agents the credentials
If the server is compromised, it still receives the real key on each call. If the vault issues a long lived key, e.g. GitHub PAT, one call will expose it forever. Does the vault issue short lived keys or the stored one?
coder-pm··on SCH: An affordable sandbox for Coding Agents in your AWS account
I have read the template and all the microVMs are running under the same role which has read and write access to checkpoints/*, so every user's folder in the bucket. The prefixes for users are just names, not boundaries. What stops one agent to read or overwrite the checkpoints for someone else?
coder-pm··on Show HN: VeriCordon – CI evidence for agent/tool authorization decisions
This looks like a gate before calling a tool but you’re saying your goal is not to containerise it. In that case what would stop the allowed call from, for example sending something outside? Is there something else sitting alongside the gate? Or is this just out of scope?
coder-pm··on Running agents in a sandbox or VM is the wrong pattern
This makes sense for stateless workers, which don’t have to keep the context between the steps. What about the interactive agents, holding ssh session or repository state between the steps? That’s a different case, isn’t it?
coder-pm··on Show HN: Socks Proxy for AI Agents
That’s kind of similar to what I am doing right now. I’m building a sandbox for agents with egress control using a domains allowlist. Is it possible to control egress in that proxy, per agent, per domain? Or is it giving a full access to the internet once paid?
coder-pm··on How well do agents use test/verification techniques?
I'm not generating the mutations automatically. Every one is a single targeted change assigned to a single test, reviewed one at a time. Thanks to that changes that mean the same thing don't stack up. The cost is reversed, I only catch what I thought about.

My real issue is different. This week one change removed the step which is creating a filename from the path and the test didnt catch it, it was passing. It wasn't an equivalent mutant, the test was looking at the wrong place. It works for me only because I'm working here on a single file, a complex bash script. Does anyone have a sensible way to limit equivalent mutants without manually checking every one that survived?

coder-pm··on How well do agents use test/verification techniques?
Yes, automated and gated. Zero missed rather than percentage. This morning harness failed the agent written test that passed on the fixed code and also passed against the mutated code. Test looked fine, code reviews would approve it, only the gate caught it! The agent didn't game it. Why I said zero missed, not a percentage? Because only one mutation survived and percentage threshold would probably swallow it.

Does anyone else gate at zero rather than a percentage threshold?

coder-pm··on Show HN: Dsnitch – Real-time, zero-config Docker egress inspector via eBPF
Hm and does name mapping still work if the user is using docker compose? it creates a user-defined network and resolv.conf is 127.0.0.11 rather than the host resolver
coder-pm··on Show HN: Dsnitch – Real-time, zero-config Docker egress inspector via eBPF
This is great, I was already doing research in that area for my tool. What about a container that writes to the /etc/hosts? It won’t emit DNS queries at all and because of that the connections will show up as bare IPs without domain. That’s a known trick, already exploited (collusion.wiki mentioned here on HN two days ago)
coder-pm··on Discovery of a new OpenAI agent message board
Answers are in the article , agents used SSH tunnels, it was evidenced by the wiki’s referrer logs. The Tor - agents did edit the wiki via SOCKS and relay R6 instantly.

The questions should be more like was CONNECT open or they didn’t even need it:)

coder-pm··on AI handles incidents, engineers lose touch with their systems
I think the biggest loss is now not knowing if what the agent is claiming was actually done:) but yes, right tools are necessary, adding a boundary will change the position, being out of touch will be recoverable. Otherwise one bad call might be unrecoverable and no amount of familiarity will save you
coder-pm··on Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
Oh so it’s about the platform where it lives, not about the logic, fair enough. Porting logic as is looks like the best approach, the one thing that broke wasn’t really logic:)
coder-pm··on Discovery of a new OpenAI agent message board
A hostname based egress allowlist is only worth as much as the box’s control over name resolution. If the agent can modify hosts inside the sandbox then it’s not a protection at all
coder-pm··on Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
Ok but if assembler normalises encoding then a byte diff won’t prove anything. Behaviour check is the way.
coder-pm··on Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
Nice, so it works quite good! So it’s a port, not a re-implementation. Good job on that, I like it that way:) do you have an idea why the trampoline jump didn’t come across?
coder-pm··on Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
Hmm did you ever run the game port against the original one in UAE? With the same inputs? Or is it just "it plays right when I play it"? That 108 byte delta is puzzling...
coder-pm··on I used Fable to rewrite 65kLoC of Go in Rust. It cost $400
Nice trick with the structures, good for agent legibility! I will try it out on the right occasion:)
coder-pm··on I used Fable to rewrite 65kLoC of Go in Rust. It cost $400
The ttyd and playwright is a clever differential way, personally I’m doing the same when it’s about to compare the views (or fix something related to rendering). Good job on that!

A TUI editor’s real output is the bytes stored on disk, while rendering can look identical the saved files might diverge (encoding, line endings, trailing new lines etc). Did you manage to diff that?

Totally agree on the overnight roadmap runs I have the same experience here. The agents have to know how to self-correct and if it’s progressing, otherwise it’s failing!

coder-pm··on Claude Fable 5.1 and Claude Mythos 5.1
Agree on the validation, my loops are already gated. My concerns are about the cases when model is passing validation and quietly abandoning the goal. The second scenario is rewriting the plan to fit what was already done.
coder-pm··on I used Fable to rewrite 65kLoC of Go in Rust. It cost $400
This is impressive but it again led me to questions. Porting the fuzzer from Go to Rust to validate Rust is a bit circular, isn’t it^^? Porting a fuzzer bug will hide the same class bug in the code it’s checking, who fuzzes the fuzzer / setup / harness:)? A good standard for rewrites is a differential testing, feed the same input to the old Go app and the new Rust then diff the outputs. Did you do that?
coder-pm··on I used Fable to rewrite 65kLoC of Go in Rust. It cost $400
How much did the verification cost on top? how did you gate it? was it a Go test suite you ran against the Rust or what? I always wonder how ppl are testing these rewrites, rewriting the tests can also lead to bug. I really wonder how reliable are rewrites like that, a 65k lines you didn't actually read. How did you confirm the semantic equivalence, same behaviour?
coder-pm··on Claude Fable 5.1 and Claude Mythos 5.1
That kind of one shot capability is impressive but how does it work for my typical work style? The way I work is to build a huge roadmap with goals and hand it to my agent to execute (often over night). I don't care that much about the benchmarks, what I care about is how often Fable 5.1 is making a baffling decision and destroys my plan, not respecting stop conditions or goals. I would seek for behavioral reliability over long autonomous runs, not eval scores. Anyone have that kind of feedback and observations?
coder-pm··on Understanding ChatGPT Work
Fair precision, obviously Cloud runs are executed remotely. Since both cloud and local runs cannot be distinguished, users genuinely won’t know when their filesystem will be touched. Basically you can’t safely depend on the user telling the tool what to use, the boundary has to be set.
coder-pm··on Understanding ChatGPT Work
Non devs running something might not be aware the programs runs locally and touches the actual machine. Devs know to be careful but regular users won’t even have knowledge it’s touching their machine, filesystem and might even touch the credentials (thanks to reasoning).

The right fix is to set real boundaries and limit agents access. We should never trust it won’t touch forbidden places.

Basically I find this naming work local vs work cloud confusing, users won’t know if it’s touching their files in the sandboxed cloud or a local one

coder-pm··on Warp builds self-improving agents on Claude
Yes, it’s just CLAUDE.md instructing agents. The point is to correctly define it and adjust to your needs
Page 1 of 3Next →