Coverity Scan Update
community.synopsys.com
community.synopsys.com
In February of 2018 it was down for over a month with no word or ETA on when it would be fixed. I hadn't thought about it since then (we discontinued use), but researching it now they released a statement saying that it was hacked. There was not a single status update during the outage. https://www.theregister.co.uk/2018/03/19/coverity_scan_crypt...
Source: I worked on static code analysis product and we extensively black-box tested Coverity.
As best I can tell, most of the warnings are either things that can be figured out for a single translation unit and (some) compilers will eventually incorporate as warnings, or things that can only be figured out by analyzing / linking across many translation units - which a compiler can't do, but the actual "advance" is simple enough if you have all the function definitions to hand.
In other words, all the advances are in warnings you don't see. Yes, all Coverity warnings you see, are simple. I agree.
Quoting from https://cacm.acm.org/magazines/2010/2/69354-a-few-billion-li...
> Since the analysis that suppresses false positives is invisible (it removes error messages rather than generates them) its sophistication has scaled far beyond what our research system did. On the other hand, the commercial Coverity product, despite its improvements, lags behind the research system in some ways because it had to drop checkers or techniques that demand too much sophistication on the part of the user.
This is correct. Coverity does a bunch of specific analyses designed to eliminate False positives. This includes analysis to determine which data states are not possible for a given code path, and doesn't report those specific issues. This also works using data across function calls.
For C/C++/C#, Coverity has by far the lowest false positive rates, which make it probably the best in class for those languages.
Disclosure: I used to work at Coverity.
I now have quite different perspective wrt false positives; that false positive rate is not important. It's all about perspectives. Name the tool bug search engine instead of bug finding tool. You rarely look beyond the first page of search engine result. Develop a ranking algorithm such that all alarms in the first page is relevant. You can use any probabilistic voodoo to rank. Etc. I don't know whether this can work. But I think it's worth a try.
Of course, this is not obvious to people who haven't tried, so it is reported over and over and over. The most recently, from Google. Lessons from Building Static Analysis Tools at Google (2018). https://cacm.acm.org/magazines/2018/4/226371-lessons-from-bu... I mean, I could have told them.
Select quotes:
> Unlike compile-time checks, analysis results shown during code review are allowed to include up to 10% effective false positives.
IMPORTANT NOTE: 10% effective false positives means much less than 10% false positives! In above quote, Google defines true bugs marked as not-a-bug by a developer as effectively false positives. If what you say is true, but users misunderstand, it doesn't count.
We have a few high-profile projects using our automated code review (marketing people tell me I'm not allowed to call it 'PR integration' any more). One example is on the AMP Project, where we caught a regex injection vulnerability in a PR before a human looked at it: https://github.com/ampproject/amphtml/pull/13060
Our default analysis has found a few other vulnerabilities (remote buffer overflows due to misuse of snprintf in rsyslog and Icecast spring to mind) but, honestly, I think our strength lies in the fact that you can write custom queries that find bugs specific to a single codebase's foibles.
For example, my first ever CVE was for a vulnerability in ChakraCore. Google Project Zero found the original bug - type confusion caused by failure to check a flag indicating that the last element of a list should be cast to a different type - but we wrote a query to verify that code accessing that particular list always checked the flag. So when some new code got introduced with the same bug, we noticed as soon as we re-ran the query on the new commit.
Any technical comparisons with other static analysis tools?
It would be useful to have a public database of QL sample queries alongside matching OSS code snippets. In theory, this could be constructed from Github, but presumably LGTM already has that historical viewpoint.
Other QL references (such as the reference manuals for the language-specific libraries) are linked from here: https://help.semmle.com/QL/ql-home.html
Personally (having joined the company reasonably recently), I haven't found it particularly hard to figure out how to express my intentions in QL. Most complex queries start their lives as simple ones that return too many results; the complexity comes from adding extra conditions to refine the results, and is generally manageable. However, I do occasionally have difficulties where a query I'm developing returns no results, and my years of experience debugging imperative programs don't really help as much as I would like in figuring out what's wrong. It can also be quite easy for beginners to write queries that are interminably slow for non-obvious reasons; I believe we're working on documentation to help with that.
We do have a database of all the default queries here: https://help.semmle.com/wiki/display/QL/Built-in+queries
But I notice that not all of them have examples, and many of the examples we do have are fairly contrived. I like your idea of taking snippets from matching OSS projects.
Edit: our CEO gently reminded me that, if you look at the help page on lgtm.com for a particular query/rule/alert, you can get a list of all results across every OSS project we analyse, e.g. https://lgtm.com/rules/1506728586782/alerts/
> It can also be quite easy for beginners to write queries that are interminably slow for non-obvious reasons; I believe we're working on documentation to help with that
RDBMS have the same challenge. Data visibility for query parser internals can help, https://www.embarcadero.com/images/dm/technical-papers/sybas...
Where do these rank in chances of future support?
- Go
- Rust
- Lua
- Swift
- Haskelland no swift for now: https://discuss.lgtm.com/t/does-lgtm-support-swift/1575
As for the other languages there, there are no plans currently on the roadmap for 2019 to add any of them. However if this is something that particularly interests you, we'd encourage you to apply for a job and note that you'd like to add support for a particular language :)
https://github.com/tesseract-ocr/tesseract https://github.com/systemd/systemd
systemd have also written their own QL query: https://github.com/systemd/systemd/blob/master/.lgtm/cpp-que... https://lgtm.com/projects/g/systemd/systemd/alerts/?mode=tre...
(full disclosure, I also work at Semmle)
We are really happy with both the product and the team and I can recommend diving into it. Nevertheless, don't rely on a single static analysis tool, no matter how good it is. If you use C/C++, you should at least also have clang's static analyzer in use.
[1] https://rainer.gerhards.net/2018/06/how-we-found-and-fixed-c... [2] https://github.com/rsyslog/rsyslog/pull/3403
I used it for Nock, too, to make some quick fixes: https://github.com/nock/nock/pull/1301/files
In Shields we have the PR app turned on and it keeps us alert to problems during the review process. It's slower than CI, though usually problems manifest early enough in the review process that they can be fixed before merge.
All links are dead, and synopsis.com’s big Corp style website isn’t helping one bit.
There we go, I had no clue what this even was. Do a lot of people here use it?
I understand from speaking to C++ engineers who have extensive experience in embedded / industrial applications that Coverity is used extensively there.
Personally I've never seen the need to apply it to CSharp projects because the language is naturally safer than C++ and you get a lot of "bang for the buck" by using the tools built into Visual Studio and from JetBrains.
* He works at Facebook now :facepalm:
https://cacm.acm.org/magazines/2010/2/69354-a-few-billion-li...
https://github.com/znc/znc/search?p=1&q=Coverity&type=Commit...
https://github.com/bitchx/bitchx/search?p=1&q=Coverity&type=...
Unfortunately they lost their focus and got acquired; if they had taken a different course they could be an essential security tool.
Update is a change, this is an outage.