As for why this metric and not others, it's the one I could actually get the data for. I don't have the time or space to download literally millions of repositories myself, so I used what I could get access to.
As for why this metric and not others, it's the one I could actually get the data for. I don't have the time or space to download literally millions of repositories myself, so I used what I could get access to.
That is only true in some programming communities.
Essentially what you are selecting for is languages whose users seem to have digested some of the same software development memes that you have. Those users are going to be generally drawn to expressive languages, and will have very focused commits in those languages. So there is a correlation between expressiveness and small commits.
But it is only a correlation. For instance I've programmed in both Perl and Ruby. Of the two, Ruby is more expressive. (Universal opinion of everyone that I know who has programmed both.) However Perl is very "unhip", so people doing open source software development in Perl these days tend to be people who have been programming for some time, which means that they've absorbed a lot of good programming ideas. (Seriously, once you get past the reputation for "unmaintainable line noise", a lot of surprisingly good code is written in Perl.) Thus Perl outranks Ruby in this list.
What would be much more informative is the ratio of lines of code/commit between languages for users that have programmed in both languages. It would take more work to do an analysis on the principles of that analysis, and you'd reveal similar trends, but the analysis would be far more informative.
1) You just assume that each commit adds "a single conceptual piece" -- where is your justification for that assumption?
2) Using the same descriptive phrase - "a single conceptual piece" - doesn't in-itself make the things described comparable.
Even if they were in some sense "a single conceptual piece" they could still be wildly different in-scale and in-complexity because of the problem-domain the project addressed not because of any intrinsic property of the language.
3) Ohloh tracks different types of repositories -- you seem to have ignored the possibility that locs/commit might have something to do with how much of a PITA it is to use the different repositories, and whether the history of language and repository popularity has meant some languages are represented more strongly on some repositories.
You've presented lots and lots of interpretation -- without exploring what's varying and what's constant in your dataset.