How is this different than bwrap or srt and others? Im using bwrap to achieve read only everywhere and and write on pwd. Also pi and other coding agents all have sandboxing that work in similar way
If there are no summaries then when context is full messages need to get evicted. If doing so one by one then it would indeed destroy the cache. Of course... maybe the implementation evicted 50% of messages at once, I didnt verify in code
In Jev you pass options in the input and its output just gives some probability for each. Oai structured output just follows a schema. The exact output is still generated and there is no probability
how do you decide to start a new agent, and when to kill? also... do you use some works-tealing style task board, or otherwise how would the agents get new tasks.
If you use the terminal a lot then TUIs keep you there wo the need to manage yet a other window in the OS. For e.g. you can run nvim in another pane in tmux.
Html has also some of these properties, targeting the browser. Don't get me wrong. I do like TUIs, they are light weight and compose very nicely with the rest of the terminal (eg tmux). Also, i generally prefer just cli commands over tuis when possible.
The downside is that... you don't get to brainstorm with the agent about ways to implement stuff. Architecture, design etc. It would feel like.. you're on your own. Of course you could just a regular chat interface alongside but theb you deal with two interfaces. Maybe this could be a target format for another chat agent that you talk to.
For eg. on a large codebase I do hierarchical review where higher level agents focus on modules and leafs on files. Each reports upwards a summary of its findings incl children summaries.
that's an amazing project! people always say why 6mb vs CC's 250mb even matter when you are calling out to LLMs hosted in the cloud. But... I regularly run hierarchies of agents with say 50-100 on a regular basis. So 650 vs 25050 ... is "can do" vs "cannot"
The x2 instance has 8gb memory. They go up to x8. And if you don't use it say during a weekend, then you stop it and resume on Monday, not paying for that period
I understand that claude-code takes a lot of memory and that's bad. However, harneses are simple loops, in theory should take very little memory even if written in python or typescript. See for e.g. pi agent
Yes! There are linux boxes that stip and resume like a lambda, but really linux. And you pay only for what you use. Look for e.g. at https://shellbox.dev, it starts at $0.02/hr
Try https://shellbox.dev for a less expensive solution, real linux boxes that sping up in seconds and suspend resume as needed. Pay only for what you use