> There are dozens of ways to measure code maintainability.
There are no good ways. I'm averse to making absolute statements, but here I'll take that chance. I worked in dev producitivy for years with people who spent decades in that domain across multiple companies with very high volumes of code production. Everybody agreed: All metrics are flawed and even a combination of metrics is insufficient.
Just to give one fundamental reason (in addition to a lot of the sibling comments): for any given metric there are an infinite set of counter-examples that don't trigger any thresholds but are clearly bad code. So these metrics typically only help in trivial cases, don't catch a majority of the cases, and so often become more of an annoyance due to low SNR. A lot of dev productivity work ends up being wiring these metrics in and then providing escape hatches when they inevitably get too noisy!
And most relevant to this discussion: these tools do not say anything about higher-level concerns like architecture, over-engineering and design, which IME is where agents tend to mess up most. I've almost never had a complaint about the code itself; the logic, naming, functions, data structures, even a lot of the testing, are all on point. It's always been the higher-level structure and design: over-engineering, duplicate classes, suboptimal abstractions, redundant operations across layers that could be solved by adding a single variable in a class, etc. etc.
I think the problem, like with code written by humans, is lack of sufficient context while doing a task leading to tunnel-vision. This is why we need to oversee and ensure things are good holistically. I suspect models are now good enough to play the role of an architect as well, though, and I've read some indications of that online... I just haven't tried giving them that much control yet.