HNHacker News
TopNewBestAskShowJobs

dimitrismrtzs

25 karma · joined April 7, 2026

Self-hosting a Proxmox lab. Building OtoDock (AI agent platform) self-hosted, collaborative AI agents. Athens, Greece.
submissionscomments
dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
Glad you like it. If you test it out I would love any feedback!
dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
OtoDock does not verify a report by itself, but you can set it up in your installation. A delegated agent runs as a normal session on the platform, so the whole run is persisted, every tool call and its output, and both the delegating agent and a human can open that session instead of trusting the summary. Delegation can also pin a different model or engine per lane, so the reviewer can run on Codex or a local model while the worker ran on Claude.

For the deterministic part I use scripts, not agents. For example to publish a new version I have a script with all the gates, it runs the tests and the scanners and refuses the export if anything is off. The agent cannot declare a release done, the script does.

dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
I agree with your points, and that is how it is built. The backbone in OtoDock is deterministic code, the model only runs inside a session. Schedules are cron, triggers are webhooks, permissions are a fail closed gate per tool, cost limits are set per user and per agent, and delegation between agents is configured from the admin and the agent managers, an agent cannot change any of that. Agents only see the tools and skills their manager assigned and the integrations are MCPs against existing tools, GitHub, Notion, Home Assistant, Prometheus. Memory is plain markdown files you can read and edit.

So I agree that the company should be tool driven. OtoDock is the home of the agents, and from there they connect to and control the company tools

In addition the agents are distributed at the execution level, the control plane and dashboard run on the server but the agents can also run on whatever machine you pair, with their files synced.

dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
The map is mainly there to see the whole installation at a glance. The departments, the agents in them and who delegates to whom. There is a plain grid view as well.

The actual UX of departments is when you put agents in one, delegation between them is configured automatically, so agents can talk to each other and tranfer files between their workspaces.

dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
Fair comparisons. Paperclip does the org chart and the orchestration but it depends on other runtimes to actually run the agents. OtoDock also runs on Claude Code and Codex, the difference is the agents run inside it, as persistent processes in a sandbox on your server. Or also you can pair one of your machines with one command and you make any agent run on it with its tools and files synced to the platform, with automatic fall back to the server when the machine is offline. Cloudflare OS I only know from their post, so I cannot compare it properly yet.
dimitrismrtzs··on Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
Thank you for checking this out and for your feedback, I will check Firefox now!

The voice mode started from the phone service. I built a full duplex pipeline so agents can connect to frerPBX and use a phone line, and then I ported the same architecture into the dashboard as a voice mode for any agent. Personally I use it for quick things and questions. For real work I use simple dictation or type.

There is no automatic routing. You start sessions with an agent for doing work and if a job needs more than one agent you wire delegation between them so they can talk directly or department head can hand work down and get the result back.

On the 3D map, I was not a fan myself, it came from user feedback but I 've found it useful to see the whole setup at a glance.

dimitrismrtzs··on Incident Report: May 19, 2026 – GCP Account Suspension
stories like this are why i self-host most things on proxmox instead of depending on a single cloud provider. Ok I have to do the maintenance but at least this way no one can suspend my entire stack with not even an explanation.
dimitrismrtzs··on How Claude Code works in large codebases
So if 100k codebase is considered large, this is how i do it. I keep a documentation folder, one for the backend one for the frontend and i try to keep these as up to date as possible. So every time i start a new feature or try to fix a bug, i tell claude to read what it needs from the documentation first and mention the thing we need to focus on. I dont know if the focus on documentation is even needed that much but with this flow i have seen claude sucessfully do some pretty complicated stuff and not a problem with exploring the codebase.
dimitrismrtzs··on Granite 4.1: IBM's 8B Model Matching 32B MoE
The 8B class closing the gap with 32B is the real story of 2026 for anyone running models locally. I've been using smaller models for agent tool-use and the progress this year is real.

The gap that still matters most isn't intelligence — it's consistency on structured output. When you chain 5+ tool calls in sequence, even a small per-call reliability difference compounds fast. Would love to see Granite 4.1 benchmarked specifically on multi-step function calling rather than just general benchmarks.

dimitrismrtzs··on Ask HN: What would you do with an AI model capable of continuous learning?
Honestly with the latest mythos news I am a bit concerned about when this happens. But I would test it as a CEO.
dimitrismrtzs··on Ask HN: Is building a company around an open-source AI agent platform realistic?
Do you have any idea about the right choice for the license?