148 karma · joined August 19, 2019
The interface is cheap to verify: you look at it and you know. The agent core (compaction, scheduling, loop heuristics) is not. Whether a change there made the agent a bit worse at multi-file refactors only shows up in eval sweeps across models and seeds. An in-process plugin API over that layer freezes it right where it needs to evolve.
So we borrowed from Unix pipes: plugin the surface, protect the plugged. The engine runs as a daemon (ante serve) and talks over a typed message stream. Clients send operations (prompts, approvals, interrupts); the engine emits events (tokens, diffs, tool progress). The terminal UI you get when you run ante is just one client of that stream.
Some of the examples:
examples/mini-tui: a one-file Rust TUI
ante-acp: Agent Client Protocol adapter for Zed and JetBrains (work in progress; no session resume yet)
ante-gateway: Slack and Discord bot, one isolated session per thread
Love to hear more about your thoughts on the future of agent architecture
Harness is compute and Model is data
Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20
Yes the core part is a simple loop, but we have all built toy compilers, inference engine, browsers (it is just a curl command eth) etc. The core algorithm is supposed to be simple.
Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20
as for why stop here yep, the goal is to be able to find the best sweet spot in being self contained v.s. all-in-one bloat ware. for example, we still use a bundled `tmux` skill for the orchestration.
- Ante installs a pinned, checksum-verified official llama.cpp build matched to your machine (Metal on Apple silicon; CUDA, Vulkan, or CPU on Linux) and handles upgrades when the pin changes. - It discovers GGUF files already on disk (~/.ante/models, the llama.cpp and Hugging Face caches), attaches to llama servers already running on local ports, and estimates RAM/VRAM from model size and context window before anything loads. - `ante --offline-model /path/to/model.gguf "prompt"` boots the server, runs the session, and shuts it down. `/offline-mode` does the same interactively; `ante serve --offline-model` loads a model once for many clients. - No API key, no account. Once the model is on disk, inference needs no network at all; set ANTE_TELEMETRY=off and no telemetry is exported either.
On capability, we'd rather publish the number than oversell: we benchmark local models with the same harness and auditable runs as frontier ones, and Qwen3.6 27B (a 17 GB download) scores 56.2% on Terminal-Bench 2.1 across 445 trials (live results: https://antigma.ai/eval). That's a real gap from frontier models. The design bet is that you mix: hosted providers and local live in the same catalog, `/providers` switches mid-session, so sensitive repos or high-volume work go local and hard problems go frontier.
Hosted models work with your own keys or subscription. But nothing about trying Ante requires signing up for anything: download the binary, point it at a GGUF.
Offline mode is under active development and has rough edges with the overview at https://ante.run/local/overview. I'll be in the comments.