There are deterministically computable metrics for cognitive complexity and readability[0].
They’re not perfect by any means — and I suspect they’re already included in the RL process for coding evals, and have been for some time. I do think we’ll see ongoing improvement in this area though.
[0] (pdf warning) https://www.sonarsource.com/docs/CognitiveComplexity.pdf