HNHacker News
TopNewBestAskShowJobs

ubermon

148 karma · joined August 19, 2019

submissionscomments
ubermon··on Plan mode is dead
both points to the same outcome: plan mode within harness is dead.
ubermon··on Everything Is a Stream: runtime composability over compile-time plugins
not yet, but will be included in the distributed version
ubermon··on Everything Is a Stream: runtime composability over compile-time plugins
https://github.com/AntigmaLabs/ante
ubermon··on Everything Is a Stream: runtime composability over compile-time plugins
Hi HN, this Monk from Antigma Labs, a lot of recent agent/harness projects lean toward "everything is a plugin". For the surface, I agree. Our take is that it stops being right at the agent core.

The interface is cheap to verify: you look at it and you know. The agent core (compaction, scheduling, loop heuristics) is not. Whether a change there made the agent a bit worse at multi-file refactors only shows up in eval sweeps across models and seeds. An in-process plugin API over that layer freezes it right where it needs to evolve.

So we borrowed from Unix pipes: plugin the surface, protect the plugged. The engine runs as a daemon (ante serve) and talks over a typed message stream. Clients send operations (prompts, approvals, interrupts); the engine emits events (tokens, diffs, tool progress). The terminal UI you get when you run ante is just one client of that stream.

Some of the examples:

examples/mini-tui: a one-file Rust TUI

ante-acp: Agent Client Protocol adapter for Zed and JetBrains (work in progress; no session resume yet)

ante-gateway: Slack and Discord bot, one isolated session per thread

Love to hear more about your thoughts on the future of agent architecture

ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
the point is you can redirect all the telemetry to your own exporter with nothing leaves your machine and all the benefits.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
it can be opt-out, opt-in, but the telemetry needs to be there to actually improve the agent.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
it is more like an actual harness with a managed and pinned llama the control is inverse.

Harness is compute and Model is data

ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
there are so much content and valuable stuff in the github repo you can pretty much recreated with your own agent.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
yes. Offline mode is choice, should be able to run frontier one first and then figure out how do incorporate local models as real workhorse
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
there are still many bad players in the industry, i wouldn't mind sharing it with trusted group. But I am not yet strong enough with deal and handle all those yet.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
generally the challenge now is that 1. how to deal with PR spams by AI bots 2. how to make the project sustainable especially when one has no distribution. when anyone can insta remix and re-package and re-sell your hard work. For knowledge sharing open source it is ok, but if you are serious about what you built, this is question needs to be answered first before make it a true community effort.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
called out the most asked questions - where is the source - telemetry opt-in/opt-out in README of https://github.com/AntigmaLabs/ante
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
I wonder my self, was notified by a friend, but I am grateful.. Probably because of the Meta's open model release?
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need) So it is here to stay. And it is just our focus as we don't have enough resource (yet) to improve the model and I think prompts belong to the user.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
I love pi and share many vision and value with it. But my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need)

Yes the core part is a simple loop, but we have all built toy compilers, inference engine, browsers (it is just a curl command eth) etc. The core algorithm is supposed to be simple.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
Oh no, definitely not Internet Explorer. Will figure out a way to ship with source first to address the security and concern and then figure out how to do the open source development in agentic era later, I was too carried away by the complexity of the latter and ignorant to the former.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
the public repo README is updated.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
there is one toy version https://github.com/AntigmaLabs/nanochat-rs README updated.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
even with power of AI, we are mere human and still slow. Adding this to backlog.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
feedback received, it was carry over from the preview dev build.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
for now only some of core crates is migrated, will do so progressively
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
so it is more for being self contained and works out of box if being deployed in a bare linux environment. we tried shell out to `rg` it didn't work very well and instead spending time handling the args parsing and jugging string output, we decided to spend time on building Grep natively for agent. It is just a start~

as for why stop here yep, the goal is to be able to find the best sweet spot in being self contained v.s. all-in-one bloat ware. for example, we still use a bundled `tmux` skill for the orchestration.

ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
nice, good to see more contributor in this space
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
we put it in the repo README, will add migrate more into public repo as soon as possible.
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
yes, the goal is to perfect the `ante serve` so it is easy to build gui. We are building one internally to test the protocol version
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
will add those soon!
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
we have a detailed launch thread explaining and show case exactly this! https://x.com/NoCommas/status/2086835536598351955
ubermon··on Show HN: Ante, a coding agent in a single binary that runs offline
Hi HN, I'm Mohan from Antigma Labs. Ante is a coding agent that ships as one self-contained ~15MB binary: the TUI, an embedded ripgrep, local PDF/OCR, and a natively managed llama.cpp engine are all inside. No runtime dependencies, no node_modules, no account.

- Ante installs a pinned, checksum-verified official llama.cpp build matched to your machine (Metal on Apple silicon; CUDA, Vulkan, or CPU on Linux) and handles upgrades when the pin changes. - It discovers GGUF files already on disk (~/.ante/models, the llama.cpp and Hugging Face caches), attaches to llama servers already running on local ports, and estimates RAM/VRAM from model size and context window before anything loads. - `ante --offline-model /path/to/model.gguf "prompt"` boots the server, runs the session, and shuts it down. `/offline-mode` does the same interactively; `ante serve --offline-model` loads a model once for many clients. - No API key, no account. Once the model is on disk, inference needs no network at all; set ANTE_TELEMETRY=off and no telemetry is exported either.

On capability, we'd rather publish the number than oversell: we benchmark local models with the same harness and auditable runs as frontier ones, and Qwen3.6 27B (a 17 GB download) scores 56.2% on Terminal-Bench 2.1 across 445 trials (live results: https://antigma.ai/eval). That's a real gap from frontier models. The design bet is that you mix: hosted providers and local live in the same catalog, `/providers` switches mid-session, so sensitive repos or high-volume work go local and hard problems go frontier.

Hosted models work with your own keys or subscription. But nothing about trying Ante requires signing up for anything: download the binary, point it at a GGUF.

Offline mode is under active development and has rough edges with the overview at https://ante.run/local/overview. I'll be in the comments.

ubermon··on DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness
yeah, that bench has a lot of room to improve, but it is the best we can find now
ubermon··on Claude Code Cut Their System Prompt by 80%. Does That Work for Small Models Too?
it is indeed surprising how much a small model can do! As you said, for specific tasks like coding, the training data quality and new architecture of newer models probably beats model size
Page 1 of 3Next →