It is way more than 100 lines. Why keep advertising something that is no longer the case?
190 lines with comments included so close enough
Which is the actual agent? Everything else is just glue code to connect to different LLM providers, etc
[1] https://www.induction.ai/docs/context-management [2] https://github.com/ArtificialAnalysis/Stirrup
Because I feel that the technology and space is - so - hyped and fast moving that a lot of cultish feeling rituals seem to pop up, none of which are backed by evidence. Anthropic openly recommend giving the agents.md file an architectural overview of the code, and the one time this was studied they found the opposite - that the agents.md file is best for concrete commands about how to build stuff and such, and - not - huge overviews. This was, and still is, the official recommendation from Anthropic as far as I can tell.
And then there are the benchmarks, how feel vague and not concrete, and everyone kind of knows they're not the best cuz you can't just assign these tools one fixed number ( for multiple reasons ), but everyone still looks at them and compares them.
People share skills and superpowers and plugins and mcps and very, very few of them have and kind of proof they do much at all.
It all feels a bit weird to me, and I've been on the lookout for exactly these kinds of studies more lately, because I think having this research, even if not done on the exact newest models or not the exact, newest thing, are still - vastly - superior to the alternative.
At least 3 times last month I was asked to review a change in a .md file used by agents. And I am like: "yeah I guess it makes sense?"
It feels we need to write unit tests for this stuff, but even how to do so in reasonable time and complexity seems difficult.
I would expect the agent loop and system prompt to be basically the same. Is it the precise semantics of the tools (and how closely they match what a particular agent was trained on) or something else?