Happy to answer questions about the eval methodology, the backend findings, or anything in the repo. I'll be around.
Will get this ported over into vLLM work and try to get that released soon.
Thanks to some kind folks who contributed Docker, token counting, and a handful of PRs I haven't gotten to yet.
- Broad slice:
- Full Forge: 48/72 accurate, 72/72 complete, score 66.7%
- Bare: 18/72 accurate, 24/72 complete, score 25.0%
- Lift: +30 correct runs, no paired regressions
- Bare had 42 ToolCallErrors and 6 ToolExecutionErrors; full Forge had none.
- Advanced reasoning:
- Full Forge: 3/24 accurate, 24/24 complete, score 12.5%
- Bare: 3/24 accurate, 9/24 complete, score 12.5%
- Lift: completion improved, but accuracy did not.I run small models at home, so I'm very curious.
Out of curiosity, what models are you running?