the paper finds that faults (and wmc (methods per class), cbo (coupling between object classes), response for a class (rfc), number of methods added (nma), etc.), correlate to class size (non-comment source lines of code), not program size
it only looked at one program so it could not measure any effects due to program size
moreover it wasn't measuring code quality in the sense of 'defects per kloc', it was measuring defects (whether a bug had been detected in a class in the field or not)
stripping away the acronyms, what they found was that classes that contained more code more often had at least one bug, and also had more methods, but that having more methods without having more code didn't make classes significantly more likely to have a bug
and similarly for the other complexity metrics like how many different methods a class calls and implements (rfc)
this is unsurprising, since things like the number of lines of code in a class and the number of different methods in a class are just alternative metrics of class size, and everyone knows that in general more code means more bugs
that's why we measure code quality in defects per kloc and not total defects. the paper didn't even try to measure code quality in that sense
that doesn't mean the paper is bad. if the paper's authors are correct that many other papers have failed to control for size in their defect metrics, they have identified a serious shortcoming in the existing research literature; haldar merely totally failed to understand them
so the paper haldar tried to summarize doesn't measure either program size or code quality
(some of their references did look at program size tho)
why are none of the other comments pointing this out
are they all just commenting without having read the paper
now i am sad