Sampling 3.3% of lines in a 72 million line codebase, especially across so many sub projects, is statistically significant enough to make claims like this with decent certainty.
That's a few million sample lines tested.
That's a few million sample lines tested.
Nevertheless I agree with you: there are no obvious way to do a better estimation, so words "fair" and "transparent" seem right in place there. Of course only if Karpov does not try to mislead his readers deliberately by choosing sub projects by some criterion that correlates with error rate.