On the Impact of Programming Languages on Code Quality
arxiv.org
arxiv.org
From this paper:
The reanalysis failed to validate most of the claims of [18]. As Table 6(d-f) shows, the multiple steps of data cleaning and improved statistical modeling have invalidated the significance of 7 out of 11 languages. Even when the associations are statistically significant, their practical significance is small.
The reference [18] is https://dl.acm.org/citation.cfm?doid=2635868.2635922, which is the original paper that this paper is a replication study of.
I think the lesson here is that you can write good or bad code in any language, but choose one that's fit for purpose. Using proper abstractions (including those builtin to languages) probably does more to reduce the incidence of bugs than anything you can do other than writing automated tests.
For those who don't happen to know the difference, formal code review basically uses a check list of common best practices and possible defects to look for versus doing each code review ad-hoc.
It's difficult for me to see how, in the era of jsfunfuzz, Csmith, AFL, and other such tools, such a source should be taken seriously.
You would be doing yourself a disservice to throw out all of Newton's results because he didn't grasp relativity.
5.3 and 5.4 really nail it, IMO.
And you can't even use bug fixes as a metric, because there's a difference between semantic bugs which will be significantly language-influenced, architectural bugs which are partly language-influenced, and user-expectation/experience bugs, which are broad-brush logical errors or misunderstandings.
(There are probably other categories, but those are the obvious ones I can think of right now.)
The real takeaway is there is as yet no objective metric for software quality.
The best a working dev shop can do is implement a small but non-trivial sample project with known-good logic in a handful of candidate languages, and see what kind of trade-offs fall out.
I do think user experience/expectation bugs are very language dependent though. First you have the quality of the String and Datetime libraries. Ruby has all kinds of high quality string functions that are a source of pain and bugs in C.
Say your project wants to display the different between two datetimes as "5 minutes ago" or "last month" (in several languages). Saying "1.27 days ago" is correct to an extent but also not pretty. There exist high quality libraries for this, but not in every language, and a lot of devs would stumble on this feature.
The other is UI creation, I would say it is easier and faster to create a really nice UI with HTML/CSS/JS than with native Android. Maybe I'm wrong on this example but I bet there is a pair of UI frameworks where this is true.
I consider those to be bugs in the spec, not the code. I once worked at a place that simply swept those under the rug - if it worked as documented, it didn't matter how many users couldn't get it to work. The bug report was closed as "not a bug"; you can imagine how frustrating that was.
Reanalysis showed positive association between defects and C++, negative association for Clojure, Haskell and Ruby.
The original analysis found positive association for C++, Objective-C, C, PHP, Python, JavaScript. Negative association for TypeScript, Clojure, Scala, Haskell, Ruby.
* Python, JS, PHP, Java, C++ are all very common first languages; I'd imagine the average experience level for these is dragged down a bit compared to others like Golang, Haskell, or Scala. One of the recent Stack Overflow Developer surveys found Golang use to be strongly correlated with age, and since I'm assuming bug count would negatively correlate with age you'd see a confounding factor here.
* Ruby's community has an obsessive culture towards unit testing, that could impact bug levels.
* The trade-offs that make C and C++ attractive such as memory performance are double-edged swords and will naturally correlate with more bugs. For example, I work a fair bit in embedded systems which use MISRA C out of necessity; I've also worked in large scale data processing which chooses C++ for its memory mapping performance over Java or Go. If "number of defects" is your primary measure then you're ignoring some of the reasons languages are chosen. Sometimes a language that minimizes bugs is not a viable option.
* I'd guess there is a skill correlation here: an engineer who is willing to seek out a specific less-used language and learn to use it effectively is probably more likely (not guaranteed) to do a better job than one who is just going with the default choice.
It'd be very interesting to see a regression done that accounts for these factors.
Like Smalltalk, they're forced to do this to catch the errors that would've been caught by type annotation. So then, could the argument be made that less stringent type annotation hurts code quality by removing effort from unit testing?
It does not appear that the paper took automated testing into account at all. Whether greater use of automated testing in Ruby is responsible for the lower defect rate observed here relative to Python would definitely be an interesting topic for another paper.
Clojure is also dynamically typed by default and below the line. It does not have Ruby's strong unit testing culture, but it does have a design that encourages functional programming, which proponents argue reduces sources of error. It also has a culture of REPL-driven development that arguably fills some of the role of automated testing.
Figuratively speaking. None of the languages they test enforce any unit testing. What I'm suggesting is that since the Ruby community tends to be more obsessive about testing, that might lead to an improvement in code quality when comparing against other similar languages like PHP or JS which are less obsessive about testing. It's something you might have to control for to see the impact.
Type systems are a completely different factor; you'd have to control for that variable too to determine the impact of it. My experience with both systems makes me predict that static typing would improve code quality, but if you want to actually be scientific you'd want to add that factor to the statistical analysis to see if the data support that hypothesis. Anecdotes like personal experience don't count.
The trivial type-mistake "Unhandled Exceptions" in Smalltalk would leak out more often than was comfortable, if the project didn't have a good means of testing to catch those. I suspect that has something to do with the Ruby community being more obsessive about testing.
?
MessageNotUnderstood
Very few of the unit tests I've written for Ruby/Python catch errors that would be caught by typical type system use in C/Java/etc. (Haskell, maybe, but...) The main thing I miss in dynamic languages vs. static is the more powerful automatic testing possible when test specs can leverage type information (e.g., QuickCheck/ScalaCheck), not an absence of need for testing.
Types just make everything easier to do right the first time.
Plus, AFAIK, no one has written the libraries, making it even higher marginal friction in practice even more than than is true in the ideal case.
[1] https://twitter.com/unclebobmartin/status/113589437016571084...
Sure, I've seen global try/catch blocks swallowing all errors in C#/C++ code too, but definitely not as prevalent in my anecdata.
Despite these three languages giving far too much freedom to the programmers (which some might even like), high-quality code can be written in all three of them, as long as intelligent and strict coding guidelines are thoroughly followed.
And of course lots of out of band(non-language) processes and tools that influence this like unit testing/code reviews/strict coding guidelines. But it's always nicer when it's prevented at the language level as opposed to out of band.
Suppose I have to write 100,000 lines in language B to write something doable in 10,000 lines of A. Who cares if the defect rate is twice as high per line of A: I still have five times fewer defects due to ten times fewer LOC's.
Another thing to consider is the nature of defects. A defect that allows a remote execution exploit is not the same as an incorrect result being ignored and propagating to later calculations, is not the same as terminating with a diagnostic.
So even the comparison of C++ vs. Haskell, that has ~25% difference, is within that 1%. This is similar to a study comparing the effect of running shoes on running speed that shows that Nike positively affects your running speed 25% more than Adidas, to the extent running shoes affect your speed at all, which is less than 1%.
-----
[1] One should take care not to overestimate the impact of language on defects. While these relationships are statistically significant, the effects are quite small. In the analysis of deviance... we see that activity in a project accounts for the majority of explained deviance... The next closest predictor, which accounts for less than one percent of the total deviance, is language. (https://web.cs.ucdavis.edu/~filkov/papers/lang_github.pdf)
But if it was something like switching editor fonts? I'd do it for 1%.
I think this is probably one of those. Haskell and C++ are very different languages, that are designed for very different purposes, and are popular in very different problem domains. To start with, that would make it difficult to come up with a task that wouldn't handicap one language or the other. (For example, a video game would favor C++.) Even if you found that, you'd have a hard time coming up with a good test. Under random assignment, Haskell is presumably going to be handicapped relative to C++ by the fact that most languages are imperative and use Algol-style syntax. You could try to avoid that by using participants with zero programming experience, but then you're studying learning curve for absolute beginners rather than productivity among experts. And so on and so on.
Long story short, probably best to leave this one to the spittle-flecked Internet arguments.
For example, immutability provides similar benefits to GC. Instead of the developer having to manually manage references across the project by hand, the language takes care of that work. This frees the developer from doing additional work, and removes a source of errors. This also directly leads to the ability to do local reasoning about code resulting in developers needing less context to understand code. So, if GC plays a significant enough role then it's highly likely that immutability does as well because it addresses a similar set of problems.
A huge, huge, huge part of the replication crisis in social sciences is due to researchers getting overconfident about the capabilities of their experimental methods.
I'm not saying there aren't challenges associated with measuring these things, as there obviously are, but it's certainly possible to study them. And the more studies we have, the more confidence we get regarding the effects.
How do you measure that?
You could look at the project's bug tracker, but the number of bugs in it will reflect a project's popularity more than anything else.
You could look at commit messages, as the (original) paper does (see 2.2.1 / 2.2.2 in the reproduction paper), but the heuristic they use doesn't look very reliable to me; I've seen projects where every feature implementation had to have an issue in the bug tracker and every commit implementing that feature had to reference the issue, so almost every commit would be counted as a defect by their metric.
And of course, all of that is data about known defects, and fixed (presumably) defects in the case of commits - how do you compare the number of unknown defects?
There's an underlying assumption [edit: not implying it's yogthos' assumption, of course] that you can simply compare unrelated projects with vastly different policies and processes, and I believe that a lot more manual work is required to get reliable and comparable data.
It's a numbers game, if we look at a lot of projects written in a lot of different languages, we'll see a sample of projects with different kinds of popularity. If there is a trend associated with a particular language that's an outlier that would be an indication that the particular language may play a role.
>And of course, all of that is data about known defects, and fixed (presumably) defects in the case of commits - how do you compare the number of unknown defects?
Again, it's about trends. If you look at thousands or millions of projects these things average out across them. You can even group projects into different categories based on their domain, popularity, and so on. The underlying idea is that you should be able to see measurable trends when looking at large numbers of projects across languages. If such trends are established, then we can dig deeper and make a hypothesis as to what might be the cause.
For these kinds of research questions, human social behavior is not a distraction from the issue being studied. It's the very locus of the issue being studied.
At that point you can make a further hypothesis as to whether certain languages encourage a particular kind of social behavior, or whether particular kinds of people are drawn to particular languages. And of course it could just be the technical features of the language that end up playing a role because social aspect may be consistent across languages.
[the] association between eleven programming languages and software defects in projects hosted on GitHub.
Nothing more, nothing less.
If they wanted a "x is better than y for business", it would not be a serious study at all, just a bunch of journalistic edgy writing.
- This study tries to enhance the original work by better metrics (e.g. confusion with C/C++), using better statistics (e.g. uncertainty).
- Automatically evaluating the quality of a code is hard and time-consuming. They had to use peer review of labellisation.
- Most of the old claims are not confirmed.
- Only 2 languages in this study had both a significant impact (impact coeff) and enough fiability (p-value): Clojure and Haskell. But the impact is still small, and there may be other influencing factors.
- They propose several ways to further enhance their work (e.g. taking regression tests into account).
Selection bias? Consider the type of person likely to be attracted to languages like Haskell or Clojure.
(Note, didn't read the original study so I don't know the selection criteria. Were code samples in different languages from the same programmers compared, or were there different sets of programmers per language? My comment really only applies to the latter case.)
On one hand, Haskell or Clojure probably attract more academic types who have a reputation for (consistent with my experience) on-average worse-quality code.
On other hand, the barrier-to-entry for these languages is higher so they tend to be programmed by more experienced folks. I assume that people pick functional languages in part because they believe the functional paradigm leads to higher-quality code, so the sample probably cares more about code quality to begin with.
They went to school, they where not bashed in the head with hammers.
They also found a significant positive association between C++ and bug.
I remember Emery being quite confused at the methodology in the paper, and he was very confused with one aspect of it, which was how they came to conclude that C was more bug prone (something about void * being able to cast to everything? It was a while ago...)
Glad to see he wrote a rebuttal, a little awkward watching a professor debate with the paper's author :)
6.3 Grep considered harmful
Simple analysis techniques may be too blunt to provide useful answers. This problem was compounded by the fact that the search for keywords did not look for words and instead captured substrings wholly unrelated to software defects. When the accuracy of classification is as low as 36%, it becomes difficult to argue that results with small effect sizes are meaningful as they may be indistinguishable from noise. If such classification techniques are to be employed, then a careful post hoc validation by hand should be conducted by domain experts.
Smalltalk was wonderful in this regard. If you wanted to search for how something was used, you just right-click search for "senders." If you wanted to find implementors, the same.
This big difference occurred with chains of senders. In other languages, where there's 2 or 3 ways to get to something or call something, you have this O(2^n) or O(2.5^n) blowup in the work you have to do to properly follow chains of senders. With Smalltalk, it's just O(n).
On top of that, there were tools that could let you compose really sophisticated SQL like queries of the code base, dynamically, then pop that up in a browser.
Overall language effects do not appear to be a dominant factor in software quality. Manual memory management is error prone. Functional languages appear to produce better results than OO/imperaive ones. Static typing does not appear to play a measurable role. In fact, Clojure and Erlang were in a category of their own, beating out statically typed counterparts.
I wonder how they controlled for developer skill?
Is the code quality better? Well maybe. Probably. But dollar for dollar you’ll do better with a good language because you’ll get much better people for your dollar.
So choice of language does affect the number of potential bugs. But this ratio of effort a programmer gives to bugs vs features is almost completely independent. Actual bugs making it past the programmer depend on how much effort the programmer cares to expend
So a safer language doesn't reduce bugs. It's like farmers with back pain from driving tractors. Giving them comfier rides doesn't reduce their pain; it just means they drive faster upto the same pain limit
If this is true, then bugs can be reduced by making bugs more painful for the programmer
But also a safer language will let the programmer drive at a higher speed, releasing a greater percentage of safe code (even if the total number of bugs is the same)
However, the base methodology appears to be really weak.
The presence of fixes doesn't necessarily reflect the presence of underlying bugs: highly conscientious projects make many "fixes" that improve the software quality against an abstract ideal without directly fixing any defect, while highly unconscientious projects don't bother fixing anything.
There are also many confounders. For example, I've seen project institute restrictions on refactoring changes in response to the huge amount of brownian-code-motion popular projects can receive, only to then have some contributors start misleading describing their commits, "by splitting this function up, we avoid the risk that someone would later introduce an overflow!".
Considering how important software quality is to contemporary engineering, including life safety critical systems, it's disappointing that we don't see more randomize studies. E.g. take a pool of developers randomly assign them to teams using different tools each to accomplish the same task, then compare the results against each other and a hidden test suite.
In the early days of Python, some employers used it as a way to filter for candidates who were ahead of the curve. I believe Paul Graham mentioned Lisp being good for attracting better programmers, although I can't find the reference anymore [1]. If there's any value in this kind of filter, the effects will show up in the programmers' output.
1. I'm having a hard time finding references for these assertions; the best I've found is some Reddit posts and Joel Spolsky's post on finding great developers: https://www.joelonsoftware.com/2006/09/06/finding-great-deve...
I find that a surprising result! This might also call into question the theory that some languages attract better developers.
>Quite frankly, even if the choice of C were to do nothing but keep the C++ programmers out, that in itself would be a huge reason to use C
A shop that makes the decision to write their stack in Haskell do it with an intention to focus on correctness. I reckon that such companies would consequently invest in other practices like more tests, higher code coverage and more intensive code reviews resulting in code with fewer errors.
This is in stark contrast to a startup that may have hacked together a POC in Python, and once it worked, with a little bit more spit and polish put it into production.
Possibly "The Python Paradox" [1] and/or "Beating the Averages" [2]. (The first one, from 2004, is a bit of an odd read nowadays, since it no longer applies to Python.)
[1] http://www.paulgraham.com/pypar.html [2] http://www.paulgraham.com/avg.html
FWIW, the original authors (Ray et al) said as much.
I've seen terrible codebases written in vanilla PHP, JavaScript, Python (probably the worst ones). But I've seen very good looking, easy to understand, easy to maintain codebases written in Symfony, React and Django.
(React is an interesting case because it's not actually a framework but a view library, which doesn't force you to follow any structure to your project. But its declarative, component-based paradigm seem to help developers structuring their projects and writing simple, reusable and composable components in a good way).
If you have to use the "weird parts" of a language to get typical work done, you are doing something wrong. And if you are not using the weird parts, then programming languages look and do pretty much the same thing.
Braggings such as, "look! My language can do double recursive lambda backflips while blindfolded and chewing gum!" are mostly just 1950's car racing done in cubicles. We buy cars to get to work and shops as cheaply and conveniently as possible, not beat Fonzy.
See Paul Graham’s essay on the Blub Paradox [0], which discusses why some languages seem to have “weird parts” from the perspective of someone who knows a less-powerful language.
I can say I hated reading C because of maccros and different build systems. I hated C++ codebases because people abused generics. And I hated Java the most because of the verbosity and the directory structure and all the layers of abstractions and all the factories and all the singletons and all the...
It makes me wish the language creators wrote a single none-trivial 'this is how we intended it to look' project.
When learning a new language I find as many big projects written in it as I can and spend time exploring them.
It helps but there is often no guarantee those projects are doing it idiomatically either.
Fair, but a lot of those abstractions make life more difficult for new programmers (Why do I have to look through 12 classes to find out what this button does?), but make life much easier when you know the code base and have to maintain and extend it.
Or, legit question - is Java more prone to these types of issues due to its design?
Yes. Except on the actively powerless languages (mostly C and Go), people write code on other languages with the same flexibility that all that indirection creates on Java, but they use much simpler structures.
You can have that in PHP, too, when the project builds on Symfony, that is to say the style is not only dictated by the language but also on the framework used.
My favorite line from this paper is: "Correlation is not causality, but it is tempting to confuse them." I think most people believe "correlation is not causality, but most of the time it is", when in reality it is, "correlation is not causality, and almost never is." There was a great analysis of this in a book or paper I read years ago that I haven't been able to find, but the tldr is that if you have a set of events, if causality were truly common, the events would for a seized up network of causality.
I skimmed the paper. I misunderstood the results.
"Unfortunately, our work has identified numerous problems in the FSE study that invalidated its key result."
I guess it just comes down to needing a sufficiently large corpus of code and commits and having to eventually publish.
I mean, this type of study is somewhat open to interpretation. It's not like everyone should arrive to the exact same numbers..
The idea is to take a clear position, which can be attacked if wrong.
Unfortunately, our work has identified numerous problems in the FSE study that invalidated its key result. Our intent is not to blame, performing statistical analysis of programming languages based on large-scale code repositories is hard.
[...]
Acknowledgments. We thank Baishakhi Ray and Vladimir Filkov for sharing the data and code of their FSE paper. Had they not preserved the original files and part of their code, reproduction would have been more challenging
It’d be interesting to see whether programmers’ preferences uphold the division posited by the study, or whether they are widely variable—eg is it likely that if you prefer ruby you’ll also prefer typescript over JavaScript, if you prefer Haskell you’d choose clojure over c++ etc.
The winners in this paper are Clojure, Haskell, Scala.
I look forward to seeing Rust added. The authors go into substantial detail about what they consider errors, and in my personal experience Rust solves many of them. See e.g. "Some defect type like memory error, concurrency errors also depend on language primitives."
ETA: They actually did not assess Typescript in their reanalysis as most of the files were not actually Typescript, section 4.1.2
"4.1.2 Removal of TypeScript. In the original dataset, the first commit for TypeScript was recorded on 2003-03-21, several years before the language was created. Upon inspection, we found that the file extension .ts is used for XML files containing human language translations. Out of 41 projects labeled as TypeScript, only 16 contained TypeScript. This reduced the number of commits from 10,063 to an even smaller 3,782. Unfortunately, the three largest remaining projects (typescript-node-definitions, DefinitelyTyped, and the deprecated tsd) contained only declarations and no code. They accounted for 34.6% of the remaining TypeScript commits. Given the small size of the remaining corpus, we removed it from consideration as it is not clear that we have sufficient data to draw useful conclusions."
> ...only four languages are found to have a statistically significant association with defects, and even for those the effect size is exceedingly small.
Emphasis mine.