Why Don't Software Developers Use Static Analysis Tools to Find Bugs?
viva64.com
viva64.com
Don't tell me to hunt down random third-party tools that none of my coworkers has heard of, that I have to convince them to adopt, and which each solve a little (usually overlapping) bit of the problem so I have to run all of them in series. That way leads to madness, and horrible one-off UIs developed by people more concerned with their tool's special-snowflake analysis algorithm than its output.
Instead, ship an extensible, plugin-based linter with your language's toolchain (e.g. mix for Elixir, lein for Clojure, etc.) Make its output beautiful. Write one or two simple plugins, and make installing lint-plugins as simple as adding something to the project's dependency file. Then I (and everyone else) will use this stuff.
When the output is as long as your arm and you have pick through it with a fine toothcomb to find the things that matter (and even they aren't necessarily causes of bugs), the whole idea becomes substantially less appealing.
Sounds like you're in the majority.
Something s = function(....); return s;
This allows me to put a break point if I need one in the future. Unfortunately Sonar is configured to say this is a Major bug. It breaks the build until I have "return function(...)." Every damn time! Now I use (assert s != null);
If a company can come up with a default configuration that Jenkins or Bamboo can pull, great! I love it. Until then I literally have to push to SVN and wait for the Sonar nag email.
For example, you can get very fine grained control over pylint by asking it to generate a config file, and reading through it (it's remarkably well commented).
I think it boils down to being timid and not wanting to impose one standard on others which reduces the importance of coding standards to tabs vs spaces -_-'
Linters that can't rank rule violations sensibly, have bad defaults and opinionated stylistic rules quickly become more trouble than they're worth.
If somebody comes out with a better linter, though, that isn't verbose and just catches truly problematic code consistently without needing a ton of customization, I'll use that for sure.
But if you're in virmundi's situation, that's more a people problem than a technical one, and no tool will save you there.
I can often determine whether someone had used Eclipse or IntelliJ; there tend to be a lot fewer analysis warnings if someone has used IntelliJ (out of the box) => perhaps we need this analysis in real time?
Well, if a variable assigned but never appears in a path that might have side effects it's dead. It's only the halting problem if you pretend to catch 100% of cases under all inputs. You can find lots of dead code without it.
For more information about the specific technique of DCE under SSA form see http://grothoff.org/christian/teaching/2007/3353/papers/ssa....
But the PL community is not really good with incremental computations, especially ones that must consider arbitrary code changes (which, for example, can cause non-monotonic movements in your lattices).
I would be tremendously surprised if that proved to be true in the general case, but it might well be true enough in common cases to be practical. Very interesting notion...
But never reveal your secret tools!
Well, and if they get curious about your magical Bug-Finding-Skills, they are easy to convince to use these tools. Its a kind of marketing.
Aside from its code formatting rules, which are just silly. They're easy enough to modify or turn off, though, so there's that.
I used jslint on a JavaScript project a few months ago. It catched a few things I overlooked, I learned a couple of things (example: it's useless to define a var inside a loop) but it's also very opinionated and insisted that I write code in a way that pleases its author. I had to refactor working code to make most of the warnings go away. I ended up with something that worked as well as the original but was much more complicated to read (at least for me). I stopped using it and regretted making all those changes to the code. There are some configuration switches but nothing that could make it work for me. I should check jshint but unfortunately jslint primed me against that kind of tools.
I prefer to catch errors with tests and test coverage tools. They must be used anyway and they bend to me, not the other way around.
The sad thing is that the powerful analysis tools are truly amazing and do find things. They just are not what the majority of people have experience with. That does seem to be changing.
I say this as someone who took 10 minutes to generate and configure the pylintrc for a new project. After taking this time, pylint has been quite useful.
No, `for` loops have nothing to do with JavaScript scope. What you probably ran into was a situation like this:
for (var i = 0; i < length; i++) {
do blah;
}
for (var i = 0; i < length; i++) {
do more blah;
}
In this case, as a stylistic preference, JSLint will suggest you remove the second var, or you can also declare it at the top of the scope and then omit the var declarations in the `for` loops anyway.And yes, strictly speaking, the latter option is the most performance-oriented, although most JS interpreters optimize for this not anyway such that it makes no difference.
But this JSLint suggestion really has nothing to do with the `for` loop, this situation would have the same effect:
var i = 0;
i++;
use(i);
var i = 0;
i++;
use(i);
`for` loops have no impact on scope, which is mostly functional, with some prototypical complexities thrown in for good measure.ReSharper: you've enumerated this enumerable multiple times. This if statement is redundant. Basically stuff within the scope of a single method.
Code Contracts: you've passed an integer to this function. The function asks that you check that the integer falls into the length of the array that you are also passing in and that it's a multiple of two. Also, the function doesn't promise that it won't return a null value and you are using the result without checking it first. Basically looks at your entire project and can work out code paths across methods.
There's worlds of a difference. ReSharper is definitely useful, but doesn't come close to what proper static analysis can do. If you like what ReSharper is doing for your codebase I thoroughly recommend having a look at Code Contracts (it's a free Microsoft Research project).
I add the contracts anyway because they really do help find bugs at run time too and also act as executable comments.
If you add static analysis to a messy project, define success thresholds based on the current violation count, and/or disable violation categories that you can live with. The rule then should be "don't make it worse, try to make it a bit better when you can", rather than "don't commit anything to main until the checker says it's perfect".
After having surveyed most of the available static analysis tools I think that part of the answer is that often the tools simply do not provide enough value for developers, especially those analyzing dynamic languages like Python / Ruby / Javascript (for C++/Java the tooling is much better). We are currently trying to change that by developing a new, data-driven approach to code analysis, which (we hope) should improve the quality of the analyses quite a bit and provide better and more actionable feedback.
Part of the problem is also cultural of course: People have different ideas about what "good code" is and they usually do not like having their code critiqued. Using an automated tool rather than manual code review to check some aspects of code quality can be beneficial though, since it is normally easier to accept harsh feedback from a machine than a human.
The original article in PDF[1] was published at the NCSU COE
People site. It was translated and published at our blog
by the authors' permission. At the end of the article[2], we
added a short section about the PVS-Studio analyzer, where
we describe which of the recommendations suggested in the
article are or aren't implemented in our tool and why.
[1] http://people.engr.ncsu.edu/ermurph3/papers/icse13b.pdfhttp://2013.icse-conferences.org/index.html
These notes shouldn't have been tacked onto the abstract. Instead they should have been clearly separated from the paper itself.
In addition, the PVS Studio team "added a short section" which appears as if it were a modification to the paper, and and it also appears in the augmented table of contents. If one is not reading carefully, or if the section heading has scrolled of the top of the screen, one might be misled into thinking the "added section" was also authored by Johnson et. al. instead of by the PVS Studio team. The "added short section" should also have been published entirely separately.
The benefits of static typing (complie-time checks) are grossly exaggerated. If the claims were true, Java itself and Java projects would be much less buggy.)
Alice: "Doing X will prevent bugs like Y!"
Bob: "Oh, but it does nothing for bugs like Z. I just won't bother at all, then".
Why would you not want to try and remove an entire class of bugs if it were within your power to do so? Just look at all the effort companies like Facebook have poured into exactly this kind of problem with things like Hack and Flow (which make use of OCaml).
Haskell or ML style static typing is very useful, completely changes the way you do things and gives you many guarantees.
C or Java static typing is almost useless with regards to bugs (see null) and serves mostly to annoy you. It has many of extra costs of stronger static typing, but gives you very little of the benefits.
To repeat, you can harden a code base without rewriting it by using static analysis tools. This is not as true with static typing.
There is, of course, a tremendous pile of things it doesn't do well, and a bunch of ways you can make it less useful for yourself, but the last C project I worked on I found it a tremendous help in refactoring compared to the nightmare I would have had without it.
To me, simple static analysis can be over sold to the point that it is worthless. I swear, I see more effort put into detecting tabs versus spaces than I do things that actually reliably cause bugs. Seriously, unless you are writing make files, I just can't bring myself to care on tabs.
However, using some of the more advanced static analysis tools that don't just show where you forgot to do a null check, but also show where you pass in a null value... That is truly impressive and fixes bugs. Even better, these are things that can be used to harden a code base without having to rewrite it.
This is why you keep re-running it, by the way. You had your null check at the top of the function, and then in maintenance someone added something new at the top, not realizing that it needed to be after the null check...
I have no stats here (and neither do you :)), but based on my experience, Java code does tend to be much less buggy when compared with dynamic-typed code, keeping the features and quality of developers the same. Of course, logical bugs don't get caught by static typing. But it helps a lot when refactoring code, or collaborating on the same codebase with many people, or changing someone else's code. These things become really important once the code base hits a certain size.
I only see this happen when both the Java code and the dynamically typed code both have zero tests.
IME, once you actually start taking integration testing seriously and actually exercise your code even just a little, the benefits of static typing evaporate pretty quickly.
If a test becomes unnecessary if you have static typing then you should never have written it in the first place. It's a bad test.
That seems a very strong - and unsupported - claim. Could you elaborate?
In my opinion it is mostly poor integration into the coding workflow, and poor visualization of the analyzer warnings that lead to too little use of static code analysis, and another is a too wide-spread 'if it works at all, ship it' mentality.
There's not much a language can do to protect a programmer from doing nonsense. Terse, explicit languages that are easy to reason about have a slight edge in that they make it easier to grasp the big picture. A powerful type system like that of Haskell can certainly help too. Practices that involves many eyes looking at the code help the most.
I wouldnt call Java an "advanced type system". Take a look at Idris then try to say that with a straight face:
>Because neither an advanced type system nor static analysis could catch bugs in program logic
It certainly does if you indeed use an advanced type system.
EDIT: Additionally, given the changes in C++11 we should be able to push all the work onto the compiler for checking type problems, remove dangerous operations, dangling pointers, naked new/delete and avoid casting where possible. It would make for bug-free software. I know someone who writes their C++ like it is late 80s C and firstly, it's horrible to read. And secondly, it does really dangerous things.
Luckily the server side 'workflow' language, and the client side templates are all written in xml, so its trivial to parse.
The first version was painfully slow. The second version tracked the range of possible states as it did a depth first search of the AST, and then stored the requirements of each node of the tree before it moved on to the nodes next sibling or parent. That way if another part of the software later called into an equivalent subtree, I could compare the current range of possible states to the requirements of the first node in the subtree and record any mismatches, then move on. It went from taking over a minute to evaluate a fairly trivial app, to parsing the biggest apps my company has ever shipped in 1-2 seconds.
I'm sure the technique is quite common. I wish I could remember if I read about it, or applied it from a different context. In any case the exercise was more useful for learning than as a finished product.
Sophisticated solutions tend to fail from time to time when applied. And when they fail they are tossed aside for simpler more resilient ones.
Sometimes we forget companies do not want a perfect code or the best possible well designed software but a product that make them earn money.
My experience is that developers only use those kind of tools if they are forced to by their QA managers of bounded by contract. Programmers usually don't want to fix or track bugs.
If that it true, than I don't want to work with them.
A typical SCA tool can report hundreds or thousands of occurrences for a certain code base. How are developers going to deal with them?
I learned, that every error you can fix early on will cost you about 10x to fix in the next stage.
All the new principles like Agile have not changed that.
I think static analysis is a complex subject, because it can touch some sensitive subject like programming style and other more expert subjects like compiler back end and how the language defines such and such code behavior.
Static analysis should be made more mainstream, it would be such a great way to teach everybody how to write better code, including and especially students.
> Mike: "Clang is my favorite. Its built into the compiler. You don't have to invoke anything special."
I agree with Mike. I've also used Coverity Static Analysis and various other tools. Some of them are better than others.
Also, "-Wall -Wextra" or "-Weverything" are fantastic for a simple first-pass static analysis.
Everybody I know who actually uses static analysis does so because it's built into their IDE and very easy to use.
What I like to do is to force a deep analyze on every build. Takes a little longer to build but at least I catch some bugs when I introduce them.
1) PVS-Studio can instead be set to run in background immediately after the edited code has been successfully compiled.
2) PVS-Studio integrates with IncrediBuild. And soon we will publish an article about it. A little spoiler. :)
Do you not parallel build?
- syntax or undefined variables in exception handlers
- showing which modules are no longer used, so we can have a clearer import block
As for more complex bugs, most developers are aware these bugs exist but do not want to fix them immediately. The reason is on many occasions, these bugs represent some bigger problem in the code base that requires significant re-factoring. And applying the quick fix as suggested by the code analysis tool simply buries these problems instead of fixing them the right way.
But they are wrong. They do make quite a lot of them. Here's, for instance, a bug database we have collected and keep updating: http://www.viva64.com/en/examples/.
Moreover, some bugs can take quite a while to find, despite being simple. Here's a nice example:
The conclusion is: the bug we had wasted about 50 hours to track was detected at once with the first run of the analyzer and fixed in less than an hour! Source: http://www.viva64.com/en/b/0221/
True, false positives aren't good, but they are not that much trouble. Static analysis tools provide numbers of means to suppress them. At least, we in PVS-Studio do have a lot of false positive suppression mechanisms. But it's a long story, so you'd better refer to the documentation.
Ooops.
It's Alt+F11 in Visual Studio. Between that and Resharper, life is good.