3 karma · joined July 25, 2026
I spent the past decade building infrastructure for production quantitative trading systems. I'm now evolving Kungfu into open-source, local-first infrastructure for long-running agent work: preserving verifiable work state across sessions, handoffs, interruptions, and model changes.
Project: https://kungfu.tech Source: https://github.com/kungfu-systems/kungfu
A long running work could have many attempts, failures and retries should be aggregated into the same result.
An acceptance of a work result should not be inferred as "the model said it is done", but be recorded as a human approval.
To calculate the cost of the whole work, the outcome and verified result of each attempt must all be saved.
I don't have a final method to measure human effort, right now I start by recording recovery, review and explicit acceptance events, the at least we have data to analyze.
I'm exploring this in a work-runtime project: https://github.com/kungfu-systems/kungfu
What is the boundary of the task?
Do we count failure and retry as the same task?
What if the result "looks" good but rejected by the user?
How do we measure the extra human effort for verifying and resuming when using the cheap model?
In my opinion, "cost" means more than the direct token usage of a finished task.
I saw opt-in, file permissions and SSRF in README, but I do not see:
domain allowlist; human approval before submiting/deleting; persistent audit record after operations; how to revoke a previously granted access;
The prompt injection may also induce the agent to perform write operations.
Reuse the real user-login session also delegate the user's full authority to the agent, which obviously has potential security issues.
The point is, the more real authority the agent has, the more important the responsibility the agent must take, which I think should be designed in from the beginning.
If you just want to have an alternative provider, I would recommend OpenRouter, which has provider fallback and model fallback. You can try OpenCode + OpenRouter instead of OpenCode + OpenCode Go; Or you can also try different harness such as Crush and Pi, in this way you have a setup like OpenCode/Crush/Pi + OpenRouter.
If you want to have a even better harness experience, then I would say the native harness from codex or claude code is still the best, you don't have give away the native harness just because it binds to the original models; with cc-switch you can use the native harness with alternative models like DeepSeek, cc-swith does lots of work to support this setup in best effort.
For example, the latest Codex uses Responses API, while lots of cheap models are still using Chat Completions, then with cc-swith, you can have:
Codex Responses API ↓ CC Switch local router ↓ DeepSeek Chat Completions
With this setup you have Codex native harness, and DeepSeek model as the underlying model.
Or even better, you can have:
Codex/Claude Code ↓ CC Switch ↓ DeepSeek + OpenRouter/Z.AI/MiniMax fallback
Then you enjoy the best harness and any model that suits into your workflow.