HNHacker News
TopNewBestAskShowJobs

coder-pm

30 karma · joined June 20, 2026

Marcin Polak

Experienced software architect, recently associated with a GenAI startup.

Cleat, a Docker sandbox for AI coding agents (MIT):

  https://cleat.sh

  https://github.com/cleatdev/cleat
Blog:

  https://cleat.sh/blog
dev.to:

  https://dev.to/coder-pm
Not a native speaker but keen to share my knowledge and experience.

In case you want to contact me, you can easily Google me by my username.

submissionscomments
coder-pm··on 90 days of attacks on AI infrastructure
The attackers pull the key from live process, not from disc so neither chmod nor disc defences method will work. The right fix is to keep the keys out of the runtime! Entirely!

It’s even mentioning they are detecting the agent framework so it’s easy to hide the malware inside its config/work directory. It’s clear, the best way to mitigate that is to start using per project scoped sandboxes.

Egress was the channel (DNS/OAST callbacks) so the network limitations beats the app layer boundaries.

coder-pm··on Bug Blindness
Hah in the AI era I noticed the agents are worse than the average person at noticing bugs. Agents use their own scale to decide if something is a success even if the output is broken. Agents are bug-blind to their own work. The consequence is that spotting bugs is now the key reviewer skill. The human’s job shifts to catching what the agent is blind to in its own work.
coder-pm··on Warp builds self-improving agents on Claude
Every decision has keywords picked from the predefined list and every time Claude is looking for the decisions made it’s querying it by the keywords (grep). I didn’t ever hit the context window issue with the log, even in a huge projects (months of work).

Btw it’s a fair challenge, I will probably hit it one day so something like a “compact” skill for decision log would be useful.

coder-pm··on Warp builds self-improving agents on Claude
My way to do the self improving agents is a CLAUDE.md instructed to write my every decision to the decision log with the relevant context. Agent is using it to challenge me, to make things better and remind me why I did something. It also helps with the invalidation.

How is the invalidation handled in Warp? Is it actually self-improving, or just better retrieval?

coder-pm··on Data Exfiltration from Amazon Kiro via Prompt Injection
The root cause is simple, the agent could read a live key! It should never have it. It’s a combination of few issues at once: repo read, write access to the settings file and outbound fetch. There is no single way to solve that, this requires egress control and no real secrets! The durable fix means a successful injection cannot extract anything because there is nothing accessible.

It’s all the same for the similar class tools like Cursor, Copilot or Claude Code. Untrusted repos are the new threat model now.

coder-pm··on I accidentally turned LLM memory into program analysis
This is the fact I’ve been struggling with for quite some time. It’s not because it forgets the facts, it’s because the invalidation doesn’t propagate.

My way of handling that is a decision log. For every project since I started doing that it’s working great. My CLAUDE.md instruct the agent to store my every decision to the file with a metadata when I made this decision and what was the context. The agent is using this file as an index of decisions and rarely lose a track. It also helps team members to find out more about the development phases.

Does your system invalidate the parts of the memory if these are not valid or relevant anymore or just store/retrieve?

coder-pm··on That's a Lot of YAML
Fair point, that’s why I said I won’t pick it. But it still might be a good and simple choice for small configs, hand written. It does explode when it’s machine generated or templated.

I think it’s also strictly connected with the ecosystem devs are working with, ruby or python devs usually pick YAML, JavaScript/Node devs with JSON - this is how it works, ppl chose what they know and what they are used to.

coder-pm··on That's a Lot of YAML
It’s because ppl are not using it correctly, it was designed for configs and now it’s being used as a programming language… a language without types and debugger. It’s frustrating devs because they find the issue at the deployment time, not the compile time.

Personally I did never pick the YAML as a first format for the configs, only my ruby friends did that.

coder-pm··on Please stop flooding our projects with AI slop to furnish your CV
Hmm I thought the contribution graphs were always the currency, the only difference is that the AI made gaming them free:)

In my company we finally have it all documented, it’s super easy to find which PR made the regression and our internal policies and rules clearly state that you are the owner of the PR, not the AI so the ownership remains on human side, AI mistake is your mistake so everyone pays extra attention to that.

coder-pm··on Pnpm 12.0
They should focus on the tool, do not try to rewrite it… everyone tries to rewrite things to rust now, pnpm was already great alternative to npm - just keep improving it instead of rewriting:/
coder-pm··on Changes to Sourcehut's terms of service regarding LLMs
Hah, llm detection is unreliable so a ToS like that can’t be enforced that way… that change moves responsibility to contributors who will break it, what about the autocomplete, refactoring tools?
coder-pm··on GLM-5.3-Flash
And what’s your hardware and what were the tokens per seconds metric (do you have it)?
coder-pm··on VMs won't contain cyber-capable agents
Yes but… they pay it once and then reuses the exploit
coder-pm··on VMs won't contain cyber-capable agents
Attacks take hours, it’s in the article. The realistic road into those weeds is prompt injection, the tooling is more and more protected versus it but it still happens:) Again for the day to day work I doubt it will ever hit us.
coder-pm··on GLM-5.3-Flash
Is anyone actually tried it in agentic coding (claude code loops)? Are apple silicon macs (M5 Max) capable of working with that model? what was the tps?
coder-pm··on It’s so hard to finish an idea that is not yours and is just suggested by AI
My approach is to keep the comments only with a references. My LLM based projects always have the decision log where are my decisions while working on the features are stored with the date. I found this useful to actually trace why something is in the codebase. It's much easier for both me and the LLM to navigate through the big projects where I spent months and dozens of full night sessions executing my plans with --dangerously-skip-permissions. In the morning I was answering all the model questions and iterating like that. Honestly, try that.
coder-pm··on VMs won't contain cyber-capable agents
This is changing so fast, if the models like GPT 5.6-Cyber can find a way to escape why the VM maintainers won't use it to fix the vulnerabilities? For the day-to-day work this doesn't make any difference, you won't hit that issues at all.
coder-pm··on Qwen 3.6 is now much easier to run locally on your Mac, thanks to JetBrains
Nice, I will give it a try! Thanks!
coder-pm··on Yeschef: Claude Code dispatches work to Ollama on my LAN (627 tok/s on 3 NUCs)
I tried so many models on my Mac M5 Max with 48GB using Ollama... it was always so slow, I know it's because of the 48GB of RAM but still... I don't know if it's worth investing in the hardware when we get a new model every month and the hardware requirements are going up and up so fast=/
coder-pm··on 80% of developers find AI coding more addictive than helpful
Personally, as a very experienced developer with over 10 years of experience, my own business and a full-time job... I have to agree. I don't remember the time when I spent so many hours in the evenings doing the things I never had time to work on... so yes, this is addictive as hell and a lot of my dev friends says the same - the god mode for experienced senior engineer is a curse...
coder-pm··on Ask HN: Coding is a solved problem. What is left for experienced engineers?
I remember my first words I said when tested the claude code... for me, an software engineer with over 10 years of experience it was like a god mode. I can basically do anything I want, having my domain knowledge, experience and hundreds of closed projects but... now after months of using it I can clearly see the coding problem is not solved.

Working on more complex problems still takes a lot of time, the models get lost, often jumping into a "solving loop" where the same solutions are being tested over and over again. You still need to guide the model, guide the loop otherwise it's lost or it's shipping a wrong solution, often even something different. Don't you now spend much more time on reading the results? iterating? Nothing changed, it's just a different way you're getting to the right place. Without a good software engineering knowledge you can't do a bulletproof, secure and high grade software. We're still far from doing a single prompt and getting a great grade output.

We might be heading that direction so personally I'm focusing on improving my soft skills, improving my general knowledge about the business, many skills around software engineering (closed to the client side, closer to the business) that until now were useless (at least I thought so).

This is an amazing time for experienced software engineers, learn more than before, try new things and this god mode will bear fruit!

coder-pm··on Qwen 3.6 is now much easier to run locally on your Mac, thanks to JetBrains
Thanks! I have to try that! Might be tight! Can it run in Claude Code? Are you loading it with Ollama?
coder-pm··on Fences, Not Sandboxes
How do you find the bad decisions with so many subscriptions active? isn't that looking for needle in a haystack:)?
coder-pm··on Qwen 3.6 is now much easier to run locally on your Mac, thanks to JetBrains
Anything good to run on Mac M5 Max with 48GB? is this even worth trying? so far I found the responses so slow compared to the paid subscriptions...
coder-pm··on Launch HN: OneCLI (YC S26) – OSS sandboxed agent harness for teams
OneCLI controls what the agent can reach. It doesnt control where it runs. Block a leaked key and the process is still on your host, so a bad rm or a prompt injected "clean up this repo" still hits real files. So I would stack them, not pick one. Sandbox the agent so it can't touch anything you care about, then route egress through a policy layer like this. Reach and blast radius are different problems:)
coder-pm··on I reverse-engineered the three biggest agent-memory tools
Can you share anything:)?
coder-pm··on Show HN: Browse, search, stats, skills from your Claude/ChatGPT chats, locally
I have Claude sessions running for months, how does the initial indexing and IndexedDB handles that size?
coder-pm··on I reverse-engineered the three biggest agent-memory tools
Did any of these actually covers the invalidation? It's easy to store memory and access it but I'm curious how these tools handles the fact that something is not true anymore?
coder-pm··on Netflix accidentally shipped a Claude.md file
Yee but screen on Reddit looks like a diff from repository... a did check it again and it looks like a fake:)
coder-pm··on Netflix accidentally shipped a Claude.md file
Not really, the closest analogy: Makefile or .editorconfig - it's created for a tool, committed on purpose.
← PreviousPage 2 of 3Next →