Programming Language Popularity on GitHub and Stack Overflow
langpop.corger.nl
langpop.corger.nl
Okay.
There are a lot of funny bits. For example the 'language' 'IO' is tagged quite a lot on stackoverflow. Or was it 'Input / Output'?
Common Lisp has fewer tags on stackoverflow - maybe because many questions are just tagged 'Lisp'.
'J' has very little lines changed. That could be a good sign, since it is based on APL, which is famous for its one line programs.
Perl6 is at the bottom. Okay.
CSS is now a programming language.
XML is now a programming language.
Self has some tags on Stackoverflow. But mostly in the context of other programming languages, not the programming language Self.
Etc.
This is flawed in so many details.
I got so fed up with "IO" tag being used for Io (http://iolanguage.org/) questions on Stackoverflow that I created the iolanguage tag and then re-tagged all the Io questions!
I did this a few years ago (2010 i think) so luckily there wasn't many questions to change :)
Io question count currently stands at 44 and not the 7,126 as shown on the chart.
With XSLT, yes.
http://fxsl.sourceforge.net/articles/FuncProg/Functional%20P...
Seconded. Could we please stop call CSS a language. If not, add PDF to the list.
"To simplify the processing of content streams, PDF does not include common programming language features such as procedures, variable, and control constructs."
> Could we please stop call CSS a language
If you're going to be pedantic, be pedantic right. CSS is indeed a language, even if it's arguably[1] not a programming one.[1] http://stackoverflow.com/questions/2497146/is-css-turing-com...
Github is a reliable source but I am not sure whether it is representative. At least it appears to be very popular with web folk. Other sources are missing such as bitbucket, especially since it allows for free of charge private repositories. Like this, the only thing this says is that github use correlates with the lines of JavaScript. Not sure whether this is a good measure for popularity. Much like a kid in kindergarten, who shows up every day and talks to everyone. Yes, his presence correlates with the presence of others, but what if he is just a jerk and nobody really likes him?
And of course basing the github measure on lines of code gives quite an unfair advantage to verbose languages like Java.
Perhaps they could look at compression ratios when (say) gzipping a decent sample of each language, and use these to weight the LOC metric?
I mean, yeah, most of the comments indicate that we agree that this kind of graph or "false statistic" is flawed.
But can we find any value in it? How do we interpret this graph, despite its basic problems? What good stuff can we do with it?
Perhaps the useful information is more in the correlations between different related metrics, than in drawing up ranked lists (surprise surprise, Java > Haskell!). This graph helps visualise the correlation between these two, which seems significant but far from perfect. Outliers like SQL can then be identified, which point to problems with the metrics (e.g. with SQL, presumably github fails to spot lines of SQL embedded in other langauges).
Also a lot of repos contain multiple languages (web languages more often than not do)
For #1, I would measure the frequency that a language appears in job ads.
For #2, I think the trending repositories in Github would be a better metric than commit lines.
This is where I think integrating these kinds of graphs with a filtration over distances on some network would be highly useful. It is not clear which network(s) would be the best, but there are some clear advantages to going with induced link through github collaborations. What I am suggesting might be more helpful for niche domains than it would be for the ``median'' programmer (if such a thing exists), but who knows, maybe there is some usefulness there too?
It is weird that Dart has 50M LOC when Go only has 100M LOC. My analysis of Dart's popularity suggests that it is actually a lot less than 2x less popular than Go. I though Go was actually 10x more popular that Dart at least.
I wonder how much double counting and so forth (as a result of forks) is present in this analysis. Maybe there is a bias as to what is released publicly to Github versus used in production.
- Programs written in this language tend to work
correctly the first time they are accepted by the
compiler
- The type system in this language helps me to be more
productive
- Code written in this language is easy to read and
understand at a glance
- I would choose this language to write a desktop GUI
application in
and asked people to choose which of a pair of languages each statement applied more to. The data could be ranked by language or question, and I found it really useful and discovered a few languages I wouldn't have heard of otherwise. But for the life of me, I can't find the site anymore. Does this ring a bell to anyone?Obviously Fortran is used a lot in HPC and numerical routines, but I always figured those were closed-source and rarely updated, with development focusing on their wrappers.
There are quite a few that are open-source, and often under very permissible licenses, even public domain, because they were developed under university research programs.
I remember finding at least half a dozen, if not more, open high-performance libraries that implemented a direct method for solving sparse linear systems. I'm fairly sure that applies to most everything else.
Perhaps even just the same commits replicated in lots of people's forks of (say) scipy -- do they control for this? what about rebases of the same commit?
There are known issues with distinguishing code between certain languages. For eg. There are issues between Perl & Prolog (and sometimes even Puppet). And up to very recently [1] most Rebol code was identified as R :)
[1] That is until this patch was merged in (although the repos language stats only get updated after a new push happens) - https://github.com/github/linguist/pull/1005
At the end of the day there's a clear correlation so I think this graph is a reliable source of information. Is there a way to remove logarithmic scale?.
Deleted comment
It has some strange coercion behavior and some floating point caveats, that programmers can learn and avoid in a week or so of learning the language. After that, it's fairly streamlined and easy to use, and it surely does not necessitate relentless rewriting to fix bugs much more than any other similar language.
Also, SLOC for ranking on Github? We don't use that to assess "productivity" for a reason.
Sure, a language can be 1/5 as terse as another. But not 1/100 as terse.