CC-Canary: Detect early signs of regressions in Claude Code
github.com
github.com
Don't lecture me on basins of attraction--we all know HK is a great programmer.
my claude usage has drastically dried up as i've personally realized the real bottom of this stuff is always gonna be genuinely learning and becoming excellent at a thing. i think claudes not bad for helping me get through early stages of that process, and for actual work i think claudes great for just ripping out something im too lazy to do, but something i know so well that i can catch him slippin. absolute coinflip on if it's worth the pain, many times now ive said "i shouldve just done this myself".
i've got my fingers crossed for like llm's to reach some sort of proverbial opus 4,5 territory here. even if thats gonna cost me a bit in hardware, that's kind of my personal bench mark for "good enough, i'm unplugging from all this craziness".
one thing is for sure, anthropic needs to stop adding _features_. claude code or vscode extension was their bread and butter reliability that garnered them a lot of goodwill from people who were willing to pay good money for a good service. seeing them launch their design thing just has me rolling my eyes. they're kind of microsoft'ing themselves here by trying to do too much, and they'll end up delivering a lot of subpar services that aren't best-in-class at any one thing. we're already seeing that i think.
In that situation, your model accuracy will look good on holdout sets but underperform in user's hands.
you can drift a tool via the harness in many ways
you can modify the system prompt
you can modify the underlying model powering the harness
you can use different “thinking” levels for different processes in the harness
you can change the entire way a system works via the harness, which could be better or worse, depending on many things
you can introduce anti-anti-slop within the harness to foil attempts from users using patch scripts
you can modify how your tool sends requests to your server depending on many variables
you can handle requests differently, depending on any variable of your choosing, at the server level
you can modify the compute allotment per user depending on many things, from the backend, without telling the user, it’s very easy. you can modify it dynamically depending on your own usage or the user’s cycle. Or their organization’s priority level as a customer. The weekly and daily usage management system is intricate, compute is very finite and must be managed
the user has literally no way to know and you have no legal obligation to tell them, you never made them any legally binding promises
the combination of so many factors that all affect each other means that you can, if you’d want to, create a new clusterfuck of an experience anytime any of these or unknown variables change, it may not even be deliberate, it grows exponentially complex, so you may not even be able to promise a specific standard to your users
drift is not imagined, sure, but admitting to it could expose you to unneeded liability
I learned to do that a while ago, after discovering not all of the block lists were enabled by default.
Anyone know of any other similar tools that allow you to track across harnesses, while coding?
Running evals as a solo dev is too cost restrictive I think.
Going to feed into my own.
Out of curiosity how have your agents evolved and metric changed.
This project is somewhat unconventional in its approach, but that might reveal issues that are masked in typical benchmark datasets
Claude Code changes all the time—it's the whole shitty trend of the day—but you can't tell which of those changes are better or worse from analyzing results on independent novel tasks.
And you're baking in certain conclusions: "HOLDING / SUSPECTED REGRESSION / CONFIRMED REGRESSION / INCONCLUSIVE". Where's an option for "better than previous baseline"? Seems certainly possible that a session could have better-than-average numbers on the measured things.
Overall, though, there's just so much here that's just uncontrolled. The most obvious thing that isn't controlled for is the work itself. What does the typical software project look like? A continued accumulation of more code performing more features? What's gonna make an LLM-based agent have to do more work? Having to deal with a larger, more complicated codebase. Nothing in this seems to attempt to deal with the possibility that a session that got labeled a regression might have actually been scored even lower against a month ago's Claude Code.
"It's harder to read code than to write code" and "codebases take more effort to modify over time as they grow" are ancient observations.
Drift detection would require static targets and frequent re-attempts.
I use it everyday and haven't seen worsening. (It's definitely not static but the general trend has been good.) But I use it on a codebase that was already very complex before we started using these tools, where overall every three months or so has brought significant improvements in usability and accuracy.
If you feel the need to do this, it’s time to move onto a tool you trust?