That's fine. Do what you do. But don't read this article as any sort of science. It's massively opinionated rage-bait.
That's fine. Do what you do. But don't read this article as any sort of science. It's massively opinionated rage-bait.
Still, I agree it doesn't make much sense to only talk about who's writing the content instead of the content. Perhaps, generously, the other comments covered their opinions on fhat part already.
I think there is overall something here for current Claude which is there appears to be a hierarchy of conformance that breaks progressive disclosure and the utility of skills. It seems to honor the system prompt, user instructions, tool call results, and dead last skills. It applies a large amount of discretion as to whether to honor what skills say in the imperative and progressive disclosure seems to have at best a 20-30% recall. Other models like codex gpt 5.6 seem to be the exact opposite and slavishly adhere to the Agent/skills/plugins, to the point of being wasteful and dangerous. It feels clear there’s a tension being RL’ed around between conformance and skeptical behavior that neither has quite found the balance for yet, and is almost certainly an over constrained problem. I just find it funny Anthropic is the one you can’t trust with your wallet while OpenAI does precisely what you and your harness tell it to.
But this “science” and its editorializing are based on flawed techniques, don’t lead to the conclusion let alone the editorialized extrapolation, and are l
Someone betting—sorry, investing in a prediction market—means you can have even more confidence in their convictions.
But someone having a Patreon account for people to voluntarily donate money to them means they can't be trusted.
> That's fine. Do what you do.
I work on several projects on the side, and they all have agents files that give the context about what the project is, what the elements of it are, and what things we're typically working on. This allows my initial prompt to reference things that would otherwise not be in the context at all.
I mean, they're not magic, they're just some automatic context that's supplied. I feel like the study is trying to say that context with an LLM doesn't matter, which is obviously not a tractable position to hold.
The fact that they generated all the agents files instead of curating them with a human is probably part of the problem.