We only avoid doing it at scale because it's expensive. In particular if we want the measurement to generalise out of sample.
(In particular in this case, where once we're done, proponents will claim our data is too old to be a useful guide to tomorrow.)
The problem with this is that AI will create worse code that is going to cause more problems in the future, but the measurements won’t take that into account.
If we could even measure teams, against themselves, others and some kind of baseline, but we don't AFAIK.
Unironically, ai evaluating the impact of those lines might be getting close to a metric that would measure output better than having everyone print out their last 6 months of work for the new boss to look at.