HNHacker News
TopNewBestAskShowJobs

GustavHartz

15 karma · joined August 16, 2021

submissionscomments
GustavHartz··on Pi Security – Codex Security without all the bloat
OpenAI recently released Codex-Security. We simplified it based on what we use in our own products and put it in the PI harness. We will release data on DeepSeek v4 and GLM-5.3 performance soon, but so far it's probably the best open-source vulnerability detection tool out there right now.

The idea is the same as Codex Security.

1. Build a threat model 2. Launch a lot of probing agents looking into the security based on the threat model 3. Deduplication 4. Validation 5. Severity and likelihood calibration

We did a deep dive article here on it as well https://x.com/GustavHartz/status/2084926544800035266

GustavHartz··on Codex Security
TLDR on how it works: It's a small stack of skill files and some JS code that starts a large number of Codex sessions. They all get the prompt and the same scope with a limited set of tools. Not one prompt in the repo contains security advice, performance is obtained only through scaling the number of agents looking at the code
GustavHartz··on Building effective pen-testing agents
This started as a response to the recent "you have to post-train a model to pen-test" Show HN — we don't think you need to, just makes life a bit easier.

Across 10K+ of our agent transcripts from benchmarking against OpenAI's EVMBench, we saw zero refusals. In the closed-frontier models, the refusal you hit is mostly a separate content classifier, or a system prompt, not so much the model itself. Breadth (more cheap agents) beats a bigger model, but it puts more requirements on context engineering

https://news.ycombinator.com/item?id=48609231

GustavHartz··on In 92% of DeFi exploits AI security review flags underlying problem
Performance has gotten a lot better the last 6 months, at a level where we almost don't see it anymore at Cecuro.ai. PoC generation and multiple validation agents debating validity is the key differentiator. This is an ok paper on the topic https://arxiv.org/abs/2511.02780
GustavHartz··on AI agents find $4.6M in blockchain smart contract exploits
We've been working on this at cecuro.ai. When we test Sonnet 4.5 against real cyber security audit reports from the major firms on code that came out after the model was trained, it finds around 95% of the same bugs the auditors found. Also catches some medium severity stuff they missed. We find that you can't just point one model at a contract and expect good results though. Need to run multiple models with different prompts because they each have different blind spots. Still tricky to get working well and not cheap. Happy to share more if anyone's curious