> Just because more code has impact in the number bugs, doesn't mean more code by itself means more bugs. Correlation != causation.
So the only possible causations for the correlation we're discussing are:
1. more program code causes more bugs
2. more bugs causes more program code
3. some unknown third factor(s) simultaneously causes both more bugs and more program code
I think 1 and 3 are most often the case, where 3 could be something like developer inexperience, although some studies have shown that even experienced developers still introduce bugs at comparable rates to novices (just lower constant factors). I think 2 sometimes happens to address immediate needs, ie. hotfix for specific bug X may introduce more bugs, but I doubt it's the rule.
Regardless, my original claim still seems pretty undeniable, ie. more program code tends to yield more bugs.
> For a trivial example, 10.000 lines of "print 'hello world'" repeated won't have more bugs than a 1000 line complex C program.
But 10,000 lines of "print 'hello wrld'" would have more bugs than a 1,000 line complex C program. Probably on the order of 9,000 more bugs in fact.
The numbers we're talking about here are averages across all programs of comparable length, not to be applied literally to any specific program, because it turns out that those specific program qualities don't really matter, ie. LOC is still a more accurate predictor of bug count than cyclomatic complexity and other metrics.
Thus I can say that a 1,000 line program probably has about X bugs, and I probably won't be off by an order of magnitude unless the program was verified by a theorem prover or something along those lines. Something like verification is really the only confounder that I've come across.