HNHacker News
TopNewBestAskShowJobs

schipperai

158 karma · joined February 27, 2026

Find me at schipper.ai
submissionscomments
schipperai··on Claude Opus 5.5
They have for sure improved the safeguard classifier. I have a coding agent guard OSS tool [1] and I use Claude to build it. When Fable first came out, I was simply unable to do any work with it. Nowadays their classifier triggers but very rarely.

[1]: https://github.com/manuelschipper/nah

schipperai··on Claude Opus 5.5
I have used Fable, Astra, Opus 5 and Sol 5.6 in Claude Code and Codex heavily since each of these have been released. Recently, I settled for Sol 5.6 as my daily driver, but with this release Opus 5.5 might be the one.

Fable is on a class of its own when it comes to coding and orchestration, but it runs out pretty quick and is prohibitively expensive and slow. For me, Astra wasn't as good of an upgrade from Sol when it comes to coding and orchestration, and it runs out pretty quick too. Opus 5 had so much potential but was such a pain to talk to, so I only had other agents delegate to it.

Given this, I kept coming back to Sol 5.6 as my daily driver (with consultations from Fable and Astra when available). Sol has an autistic character that smells like RL deepfry, which can be annoying, but it is nonetheless predictable, communicates more plainly than Claude, and is good at code and orchestration.

But this Opus 5.5 might just be it. Fable-level performance that's cheaper and faster, and communicates plainly and briefly. If it pans out in practice, I might just have a new daily driver and it might be time to raise the bar for what can be achieved.

I'll try out the new Sol 6, though Opus 5.5 seems to beat that release fair and square (at least on paper)

I'm also glad we are paying attention now to the experience of using a model, not just how good it is at 'x' class of tasks.

schipperai··on Ask HN: What are you working on? (September 2026)
I’m working on nah, a coding agent guard to block catastrophic tool calls.

It’s easily extensible, has a nice simple TUI, and is resilient to an adversarial agent that tries to remove the guard.

https://github.com/manuelschipper/nah

schipperai··on Mdmanager.ai – Manage your Claude.md and AGENTS.md across machines and runtimes
Point taken, I agree Ansible could have worked nicely for this tool. While I do think there's value in packaging the workflow into a ready-made tool.

I do appreciate the input, thanks!

schipperai··on Mdmanager.ai – Manage your Claude.md and AGENTS.md across machines and runtimes
Ansible can of course handle the templating and deployment. Not arguing that.

mdmanager packages runtime-specific discovery, reusable Sections and Profiles, an instruction-focused TUI, and bundled CLI docs into something you can install and use.

You could build that with Ansible too, but the interface, domain rules, documentation, and maintenance are part of what you’d be building. Packaging all the nuances of 4 runtimes in was a lot of work.

If your existing setup already gives you everything you need, great. You don’t need this. That doesn’t make a dedicated tool redundant for everyone else.

“Ansible can do that” is a reasonable implementation suggestion. Taken it too far then “Python can do that” becomes an objection to everything posted on HN.

And no LLM is required for discovery, rendering, or deployment. Using an agent to edit the configuration is a workflow choice.

schipperai··on Mdmanager.ai – Manage your Claude.md and AGENTS.md across machines and runtimes
The tool is more than templating and automation. It adds a workflow for inspecting and composing the instructions.

The TUI and the docs understand how each coding agent runtime loads instructions into context (they all have nuanced differences) including local overrides, or what gets prioritised in case there is conflict (for instance, Pi accepts both CLAUDE.md and AGENTS.md).

Your agent can load the docs to understand nuances and help you manage how you want instructions to get loaded. You can use the TUI to render the concatenated instructions as they are loaded into your agent.

So yes, templating could be replaced with smth like Ansible, but the tool solves a bit more than just that.

Thanks for the question!

schipperai··on Mdmanager.ai – Manage your Claude.md and AGENTS.md across machines and runtimes
One agent standard that causes me pain: AGENTS.md / CLAUDE.md.

I have multiple machines (personal laptop, work laptop, server) and runtimes (Claude Code, Codex, Pi) which share common instructions, but also have special cases for each runtime or machine.

For instance, I have a section that’s common across all my agents that goes something like this: “Do not co-sign commits with an AI identity; do not add backwards compatibility if you were not asked to” .

For Pi I have a special case: “Use background task tool proactively for in-scope commands expected to run for long”.

In my server I have a section on how to use the credential proxy (which I don’t have in my other profiles).

So I made this little tool ‘mdmanager’ that helps me manage and deploy the .md files with reusable Markdown sections (like “agent-common”, “pi-common”) and profiles (“vps”, “laptop”)

How it works: you tell your coding agent to run “mdmanager docs start” and it guides you through the setup. You can then open the TUI to see the diff and inspect what gets loaded into the context. The TUI is read-only. To modify and apply, you talk to your agent (it has access to the docs via the CLI)

One feature that I use often is the “Context” section in the TUI.

Open the TUI from inside the repo, and you can see what instructions get loaded into context in one Markdown page, including global and project level *.md files that are in scope.

Right now it supports Claude Code, Codex, Pi, and Cursor. Feedback and contributions are most welcome.

schipperai··on Technical leaders should have the largest AI exhaust
Fair - some of the points in these examples will rot. What I think transfers are the failure modes: context rot, agents filling in decisions, whether a codebase is greppable. Those can turn into team principles, and you can only learn this through experimentation.
schipperai··on Technical leaders should have the largest AI exhaust
Author here. Thanks for the comment.

I did refer to some specific models, though a lot of these learnings are from experience over the past 6 months or so, and continue to generalize as frontier models improve.

Curious if there’s a specific point you feel is too detailed?

schipperai··on Technical leaders should have the largest AI exhaust
I’d agree Linus is an exceptional leader. And I don’t think most rules apply to him.
schipperai··on Ask HN: What are you working on? (August 2026)
I’m working on “nah” [1] an agent hook guard that blocks catastrophic agent actions like filesystem destruction, secrets exfiltration, and git disasters.

All the agent hook guards I’ve seen can be easily removed by agents (and they often do so they can get going with the task). I made mine harder to remove, and I’m working on making it impossible.

I also made “nah” easily extensible. Just point your agent to the docs which ship with the CLI, and ask it to build a custom guard.

I’m working on strong Python and Typescript pseudo-interpreters so I can detect disasters in inline code with less false positives. Bash parser is already very strong.

This is a side-project that I want to be proud off. I spend a lot of time designing, not just coding.

[1] nahguard.ai

schipperai··on A way to exclude sensitive files issue still open for OpenAI Codex
Do I understand correctly that you scope least-privilege creds/tokens and pass those to the sandbox? I'd be curious to learn more
schipperai··on The Coming Loop
If an organization decides the engineering team should not be looking at code, that should be coupled with a mandate to figure out what good engineering looks like working that way - what constitutes a good contribution vs what's slop? How do we handle massive PRs? The problem is we are in the "messing around phase" of coding with clankers and have much to learn still
schipperai··on Don't trust large context windows
Working in the era of 200k context window meant I had to narrowly scope tasks to fit in the context window, forcing me to think about how to reduce complexity and naturally resulting in atomic work. 1M context windows and the promise that the latest models are "better at long running tasks" made me lazy in how I scope tasks and quality got worse. I now went back to narrow-scoping one session per task and zero compaction, trying not to go past 400k context window. If I end up with a long session, I was likely too ambitious and should have broken up the task.
schipperai··on Claude Fable 5
Let's hope not all frontier AI assimilates these guardrails. It would be a shame for independent researchers and students.
schipperai··on Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
I get a sense that I was click-baited by article's title with the classic trope of "X is all you need". This research is a solid contribution, but is far from all we need to understand grep vs semantic search in agent retrieval.
schipperai··on Claude Fable 5
Cognition did well in documenting their approach [1].

TL;DR - they worked with OSS project maintainers to build tasks. They score models based on whether a PR is mergeable. All tasks are graded by a human researcher. SoTA models have hill-climbing to do which raises the bar and inspires confidence. I'd say it's legit.

[1]: https://x.com/cognition/status/2064061031912288715

schipperai··on MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
I trust AI to surface general information and best practices on established knowledge domains. For example: best practices for securing my VPS.

For domains whete SoTA is constantly changing like AI, I use LLMs to aggregate and interact with my own research from trusted sources ala Karpathy LLM wiki.

I don’t generally trust everything I read on the internet whether its AI generated or not. I do my own research for the things that matter to me.

schipperai··on MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
You can dig deeper into problems with AI. For me, it supplements my knowledge in domains I don’t fully understand. It also helps me learn. So I can tackle problems I wouldn’t otherwise.

I’m excited for ultrafast AI. It likely means less temptation to multi-thread and deeper flow in single sessions.

schipperai··on Gemma 4 12B: A unified, encoder-free multimodal model
Demis at YCombinator said that they think its best their edge models are open cause once they are put on device they are vulnerable anyways

https://youtu.be/JNyuX1zoOgU?is=PdzCILyi8SP6cfDr

schipperai··on Ask HN: What are you working on? (May 2026)
Yes, you can define sensitive paths and assign 'ask' or 'block' policies to them.

.env, .ssh, and others are treated as a sensitive filenames by default.

Similarly, with hosts and network access - unknown hosts pause, trusted hosts can be configured.

schipperai··on Maryland citizens hit with $2B power grid upgrade for out-of-state AI
This recent article from Semianalysis did a great job explaining part of it: https://newsletter.semianalysis.com/p/are-ai-datacenters-inc...
schipperai··on Ask HN: What are you working on? (May 2026)
Very cool. How do you classify negative signals?
schipperai··on Ask HN: What are you working on? (May 2026)
Which platform have you found is most hackable? I have Garmin atm and like it but there’s no easy way to pipe my data into my agent or server for offline analysis.
schipperai··on Ask HN: What are you working on? (May 2026)
I like the overall premise and would be curious to learn more. The Amazon overview reads like it was written with or by AI though.
schipperai··on Ask HN: What are you working on? (May 2026)
A better permissions layer for coding agents. The tool works like auto-mode for Claude Code, so you can stay in the flow and only get prompted to allow or deny tool calls when it truly matters, but it is fully deterministic. My benchmarks surfaced that most Bash calls don’t need an LLM to be classified as safe, ambiguous, or dangerous. A deterministic classifier can auto-allow or block 95% of Bash tool calls as safe or dangerous, with only the remaining 5% being truly ambiguous or unknown.

Conclusion is permission reviews with LLMs like Claude’s auto mode or Codex auto review are like using a data center to flip a light switch - overkill.

The main benefit is that your agent’s autonomy can be governed deterministically through policies that can be stored at the user and repo level. The bonus is that you save tokens vs using auto modes.

https://nah.build

schipperai··on Mistral Medium 3.5
Thanks, makes sense. I meant Blackwell is explicitly optimized for MoEs.
schipperai··on Mistral Medium 3.5
With most OSS releases being MoEs, and modern GPUs optimized for MoEs, can somebody with knowledge of the topic explain or speculate why Mistral might have opted for a dense model?
schipperai··on I bought Friendster for $30k – Here's what I'm doing with it
100%. The exclusivity of the network is the differentiator here.
schipperai··on An AI agent deleted our production database. The agent's confession is below
Agent permissions layer are broken. We need better a permissions layer that doesn’t get in the way but stops destructive commands. Devs get pushed into running yolo mode cause classifying allow / deny by command is not enough. A sandbox would not have prevented this either.

“nah” is a context aware permission layer that clasifies commands based on what they actually do

nah exposes a type taxonomy: filesystem_delete, network_write, db_write, etc

so commands gets classified contextually:

git push ; Sure. git push --force ; nah?

rm -rf __pycache__ ; Ok, cleaning up. rm ~/.bashrc ; nah.

curl harmless url ; sure. curl destroy_db ; nah.

https://github.com/manuelschipper/nah

Better permissions layers is part of the answer here, and a space that has been only narrowly explored.

Page 1 of 3Next →