36 karma · joined June 27, 2026
https://github.com/kerlenton
But in real applications, does the model reliably drill down from the general summary, or does it often just hang around at the level of the summary?
Is there any information from Herdr about what each agent is up to beyond the output? Or does it just concentrate on orchestration for the time being?
How does Marmot cope with it? Are all of the tools exposed in a flat way, or there is a scoping/search step which allows an agent to select between only a few tools out of the catalog?
Are there any means to create cases from real sessions or everything is written by hand?
How do you deal with replay non-determinism? When I replay a call I captured, I spin up a new server instance, but anything that is stateful, or any time that the model chooses different arguments the second time around makes it difficult to create an accurate repeat of the input and the output. I’m interested to see how Orchid manages that in multi-step execution contexts.