Quite a surprising result: “across multiple coding agents and LLMs, we find that context files tend to reduce task success rates compared to providing no repository context, while also increasing inference cost by over 20%.”
Languages == Python only
Libraries (um looks like other LLM generated libraries -- I mean definitely not pure human: like Ragas, FastMCP, etc)
So seems like a highly skewed sample and who knows what can / can't be generalized. Does make for a compelling research paper though!
I mean, it's not that hard to understand why.
If you feel strongly about the topic, you are free to write your own article.
How does this invalidate the result? Aren't AGENTS.md files put exactly into those repos that are partly generated using LLMs?