Reviewing a 50-line code change in 60% of the time it takes to review a 250-line code change means the shorter code review takes _four times_ as long to review per line.
Reviewing a 50-line code change in 60% of the time it takes to review a 250-line code change means the shorter code review takes _four times_ as long to review per line.
So not only is it more time per line, it's also more lines.
Basically the only benefit is PRs per day, which is... not useful in itself.
You could argue that shorter PRs are easier to review and therefore people are doing more review, and that should result in better changes (indeed the 15% fewer reverts supports this).
Or you could argue that shorter PRs are harder to review as they have less context, and therefore more time is taken but that a better output is not necessarily being reached.
I think both of these are entirely plausible, and the truth probably depends on other process factors.
If you split 200 lines of code into 4 x 50 you are actually going to have 0.85 x 4 reverts or 3.4x reverts with this strategy.
But overall the whole data is terrible and infuriating because likely the changes are very different in the dataset depending on the loc.
Unless the metric is actually 40% faster per LOC but that's not specified.