158 karma · joined February 27, 2026
Fable is on a class of its own when it comes to coding and orchestration, but it runs out pretty quick and is prohibitively expensive and slow. For me, Astra wasn't as good of an upgrade from Sol when it comes to coding and orchestration, and it runs out pretty quick too. Opus 5 had so much potential but was such a pain to talk to, so I only had other agents delegate to it.
Given this, I kept coming back to Sol 5.6 as my daily driver (with consultations from Fable and Astra when available). Sol has an autistic character that smells like RL deepfry, which can be annoying, but it is nonetheless predictable, communicates more plainly than Claude, and is good at code and orchestration.
But this Opus 5.5 might just be it. Fable-level performance that's cheaper and faster, and communicates plainly and briefly. If it pans out in practice, I might just have a new daily driver and it might be time to raise the bar for what can be achieved.
I'll try out the new Sol 6, though Opus 5.5 seems to beat that release fair and square (at least on paper)
I'm also glad we are paying attention now to the experience of using a model, not just how good it is at 'x' class of tasks.
It’s easily extensible, has a nice simple TUI, and is resilient to an adversarial agent that tries to remove the guard.
I do appreciate the input, thanks!
mdmanager packages runtime-specific discovery, reusable Sections and Profiles, an instruction-focused TUI, and bundled CLI docs into something you can install and use.
You could build that with Ansible too, but the interface, domain rules, documentation, and maintenance are part of what you’d be building. Packaging all the nuances of 4 runtimes in was a lot of work.
If your existing setup already gives you everything you need, great. You don’t need this. That doesn’t make a dedicated tool redundant for everyone else.
“Ansible can do that” is a reasonable implementation suggestion. Taken it too far then “Python can do that” becomes an objection to everything posted on HN.
And no LLM is required for discovery, rendering, or deployment. Using an agent to edit the configuration is a workflow choice.
The TUI and the docs understand how each coding agent runtime loads instructions into context (they all have nuanced differences) including local overrides, or what gets prioritised in case there is conflict (for instance, Pi accepts both CLAUDE.md and AGENTS.md).
Your agent can load the docs to understand nuances and help you manage how you want instructions to get loaded. You can use the TUI to render the concatenated instructions as they are loaded into your agent.
So yes, templating could be replaced with smth like Ansible, but the tool solves a bit more than just that.
Thanks for the question!
I have multiple machines (personal laptop, work laptop, server) and runtimes (Claude Code, Codex, Pi) which share common instructions, but also have special cases for each runtime or machine.
For instance, I have a section that’s common across all my agents that goes something like this: “Do not co-sign commits with an AI identity; do not add backwards compatibility if you were not asked to” .
For Pi I have a special case: “Use background task tool proactively for in-scope commands expected to run for long”.
In my server I have a section on how to use the credential proxy (which I don’t have in my other profiles).
So I made this little tool ‘mdmanager’ that helps me manage and deploy the .md files with reusable Markdown sections (like “agent-common”, “pi-common”) and profiles (“vps”, “laptop”)
How it works: you tell your coding agent to run “mdmanager docs start” and it guides you through the setup. You can then open the TUI to see the diff and inspect what gets loaded into the context. The TUI is read-only. To modify and apply, you talk to your agent (it has access to the docs via the CLI)
One feature that I use often is the “Context” section in the TUI.
Open the TUI from inside the repo, and you can see what instructions get loaded into context in one Markdown page, including global and project level *.md files that are in scope.
Right now it supports Claude Code, Codex, Pi, and Cursor. Feedback and contributions are most welcome.
I did refer to some specific models, though a lot of these learnings are from experience over the past 6 months or so, and continue to generalize as frontier models improve.
Curious if there’s a specific point you feel is too detailed?
All the agent hook guards I’ve seen can be easily removed by agents (and they often do so they can get going with the task). I made mine harder to remove, and I’m working on making it impossible.
I also made “nah” easily extensible. Just point your agent to the docs which ship with the CLI, and ask it to build a custom guard.
I’m working on strong Python and Typescript pseudo-interpreters so I can detect disasters in inline code with less false positives. Bash parser is already very strong.
This is a side-project that I want to be proud off. I spend a lot of time designing, not just coding.
[1] nahguard.ai
TL;DR - they worked with OSS project maintainers to build tasks. They score models based on whether a PR is mergeable. All tasks are graded by a human researcher. SoTA models have hill-climbing to do which raises the bar and inspires confidence. I'd say it's legit.
For domains whete SoTA is constantly changing like AI, I use LLMs to aggregate and interact with my own research from trusted sources ala Karpathy LLM wiki.
I don’t generally trust everything I read on the internet whether its AI generated or not. I do my own research for the things that matter to me.
I’m excited for ultrafast AI. It likely means less temptation to multi-thread and deeper flow in single sessions.
.env, .ssh, and others are treated as a sensitive filenames by default.
Similarly, with hosts and network access - unknown hosts pause, trusted hosts can be configured.
Conclusion is permission reviews with LLMs like Claude’s auto mode or Codex auto review are like using a data center to flip a light switch - overkill.
The main benefit is that your agent’s autonomy can be governed deterministically through policies that can be stored at the user and repo level. The bonus is that you save tokens vs using auto modes.
“nah” is a context aware permission layer that clasifies commands based on what they actually do
nah exposes a type taxonomy: filesystem_delete, network_write, db_write, etc
so commands gets classified contextually:
git push ; Sure. git push --force ; nah?
rm -rf __pycache__ ; Ok, cleaning up. rm ~/.bashrc ; nah.
curl harmless url ; sure. curl destroy_db ; nah.
https://github.com/manuelschipper/nah
Better permissions layers is part of the answer here, and a space that has been only narrowly explored.