FWIW the compaction had failed again, so i used Pi with the same prompt, same model, same config to compare. The prompt was "check the current directory" followed by "analyze the code and give me a report in `CODE_REPORT.md` on how it works (also a brief report here). Make sure to update `CODE_REPORT.md` frequently (i.e. every time you analyze a file) to avoid context loss from context compaction".
Pi ended up hitting a compaction during a tool call and finished without issues, which is what i expected - in fact after giving the instruction i left to visit a relative since i expected it wouldn't need me to babysit it.
FWIW after it finished, i asked it the following:
---
Can you answer me the following questions about how Hax manages the context?
1. How does Hax handle running out of tokens? It does have some form of compaction, but if the compaction fails for some reason (e.g. the summary ends up needing more tokens) what does it do?
2. Can it handle cases where the context runs out of tokens during tool calls and if so, can it recover? How?
3. What happens with compaction if an LLM produces a few large responses or the LLM reads a few large files? Does the process ends up summarizing the entire (or all but one) conversation? Or is the entire conversation lost?
4. Are there any safeguards in place to avoid overwhelming the LLM? For example any file and tool output limits? If there is and any limits are reached, how does it handle them?
---
It went on and checked the code and, briefly the response (it was bigger but i don't want to repeat the entire thing):
1. It doesn't update the session on failure (last "good" session is kept), running out of context is treated like any other error without any automatic recovery and you're expected to fix it by hand (personally i'm not a fan of this).
2. Multiple tool calls are fine (there is a 85% threshold check to trigger compaction and a 50k tool result limit) but if a response and results jumps from below 85% to over 100% despite being under the 50k limit, it doesn't trigger any compaction and the next request is denied by the provider (i.e. llama-server). I have a feeling this is what i hit when i tried Hax with its own code.
3. There is no recovery from a user turn (prompt + response + tool result) that exceeds the window. FWIW this also seems to be the case with Pi.
4. It found a bunch of safeguards (tool output cap, caps in bytes and lines for the read tool, bash writes to a temp file and only a part of it is sent to the LLM, edit size cap, etc). AFAICT it is the same as Pi with one neat addition in that there is a default 2 minute timeout for bash calls (i've seen the LLM more than once run a command in Pi and end up stopping for 30+ minutes because the command wouldn't end).
For 3 i asked it a followup question: "About 3: AFAIK Pi (the harness you're on right now) does a "spit" summarization where the old messages are summarized up to a cutoff/split point and replaced with the summary while the newer messages after the cutoff/split point remain intact, which allow a mostly seamless transition between compactions. Does Hax do the same or something similar?"
The response was that, no, it doesn't, it summarizes the entire context. Which TBH is a bit of a dealbreaker for me since i often rely on this "seamless" continuity in my prompts and feels like the main reason why compactions feel like a non-issue with Pi.
Take the above with a grain of salt, i only checked the code for the compaction not using a split/cutoff point between older and recent messages, the rest are whatever Qwen 3.8 27B understood, but they do match my short empirical test. Also if my own understanding of the code is correct, it seems to be using the same system prompt for the summary as for regular/interactive use while AFAIK Pi uses a dedicated "you're an expert summarizer" (or something like that :-P) prompt. Not sure if it makes much or any difference, with LLMs being what they are, but TBH whatever Pi does works great IME.
On the other hand the idea of a self-contained native AI harness in C/C++ is enticing, especially one that doesn't have any network traffic outside of LLM-related stuff[0] and explicit user requests (Pi does try to autoupdate and has a separate opt-out telemetry beacon - both of which are disabled in different means, one via environment variable and another via a setting, which smells a bit like an dark pattern to me).
Anyway, this is the result of my findings about Hax. It is neat, but TBH the context handling is the main dealbreaker for me, especially since i'm often having the agent do something in the background (using a local LLM isn't exactly the speediest workflow) and do other stuff or leave the computer alone, so the last thing i want is to babysit the agent for errors. Pi's split summarization and context overflow handling seem to work much better.
For now i'll probably stick with Pi (i have autoupdates and telemetry disabled and i hope there isn't any other hidden snitch in place) and perhaps at some point i'll do the NIH thing and make yet another agent myself :-P
[0] well, it does attempt to autoconnect to a potentially running llama-server in localhost without being explicitly told to do so (Pi wants explicit configuration) but meh