We have a few high-profile projects using our automated code review (marketing people tell me I'm not allowed to call it 'PR integration' any more). One example is on the AMP Project, where we caught a regex injection vulnerability in a PR before a human looked at it: https://github.com/ampproject/amphtml/pull/13060
Our default analysis has found a few other vulnerabilities (remote buffer overflows due to misuse of snprintf in rsyslog and Icecast spring to mind) but, honestly, I think our strength lies in the fact that you can write custom queries that find bugs specific to a single codebase's foibles.
For example, my first ever CVE was for a vulnerability in ChakraCore. Google Project Zero found the original bug - type confusion caused by failure to check a flag indicating that the last element of a list should be cast to a different type - but we wrote a query to verify that code accessing that particular list always checked the flag. So when some new code got introduced with the same bug, we noticed as soon as we re-ran the query on the new commit.
Any technical comparisons with other static analysis tools?
It would be useful to have a public database of QL sample queries alongside matching OSS code snippets. In theory, this could be constructed from Github, but presumably LGTM already has that historical viewpoint.
Other QL references (such as the reference manuals for the language-specific libraries) are linked from here: https://help.semmle.com/QL/ql-home.html
Personally (having joined the company reasonably recently), I haven't found it particularly hard to figure out how to express my intentions in QL. Most complex queries start their lives as simple ones that return too many results; the complexity comes from adding extra conditions to refine the results, and is generally manageable. However, I do occasionally have difficulties where a query I'm developing returns no results, and my years of experience debugging imperative programs don't really help as much as I would like in figuring out what's wrong. It can also be quite easy for beginners to write queries that are interminably slow for non-obvious reasons; I believe we're working on documentation to help with that.
We do have a database of all the default queries here: https://help.semmle.com/wiki/display/QL/Built-in+queries
But I notice that not all of them have examples, and many of the examples we do have are fairly contrived. I like your idea of taking snippets from matching OSS projects.
Edit: our CEO gently reminded me that, if you look at the help page on lgtm.com for a particular query/rule/alert, you can get a list of all results across every OSS project we analyse, e.g. https://lgtm.com/rules/1506728586782/alerts/
> It can also be quite easy for beginners to write queries that are interminably slow for non-obvious reasons; I believe we're working on documentation to help with that
RDBMS have the same challenge. Data visibility for query parser internals can help, https://www.embarcadero.com/images/dm/technical-papers/sybas...
Where do these rank in chances of future support?
- Go
- Rust
- Lua
- Swift
- Haskelland no swift for now: https://discuss.lgtm.com/t/does-lgtm-support-swift/1575
As for the other languages there, there are no plans currently on the roadmap for 2019 to add any of them. However if this is something that particularly interests you, we'd encourage you to apply for a job and note that you'd like to add support for a particular language :)
We are really happy with both the product and the team and I can recommend diving into it. Nevertheless, don't rely on a single static analysis tool, no matter how good it is. If you use C/C++, you should at least also have clang's static analyzer in use.
[1] https://rainer.gerhards.net/2018/06/how-we-found-and-fixed-c... [2] https://github.com/rsyslog/rsyslog/pull/3403
https://github.com/tesseract-ocr/tesseract https://github.com/systemd/systemd
systemd have also written their own QL query: https://github.com/systemd/systemd/blob/master/.lgtm/cpp-que... https://lgtm.com/projects/g/systemd/systemd/alerts/?mode=tre...
(full disclosure, I also work at Semmle)
I used it for Nock, too, to make some quick fixes: https://github.com/nock/nock/pull/1301/files
In Shields we have the PR app turned on and it keeps us alert to problems during the review process. It's slower than CI, though usually problems manifest early enough in the review process that they can be fixed before merge.