Security for Open Source Code: Dynamic Analysis Is the Only Way
blog.sourceclear.com
blog.sourceclear.com
http://cacm.acm.org/magazines/2010/2/69354-a-few-billion-lin...
(starting from "Laws of Bug Finding")
Their tool has a few major weaknesses:
1. Builds are not reproducible - Reproducible builds require pinned package versions, which they specifically avoid. This could result in security holes if a dependent package version was bumped after the parent project was tested.
2. Subject of analysis - Their main target of analysis is the build file. This severely limits the extent of their analysis and requires them to build new tools for each build tool used.
3. Underestimation - Since a code path must be exercised in order to analyze it, you are guaranteed that the set of vulnerabilities detected is a subset of (or equal to) the true set of vulnerabilities. This is the opposite of what I believe should be the default: Always prefer false positives over false negatives. Both static analysis and binary analysis allow the programmer to over-estimate their analysis, guaranteeing that all vulnerabilities will be found.This article is BS.
For #2, SourceClear doesn't build the software under test, it's hooked into the build via various methods. For Maven and Gradle, those are plugins. For NPM and Bundler, the existing build files contain the complete dependency graphs as determined by the build system. The analysis is quite accurate, I daresay more so than any other tool. Yes, it requires implementation work for each build stack, but that's the price you pay for accuracy.
In response to your #3, SourceClear doesn't report only vulnerabilities verified by call paths, it reports on all components will known vulnerabilities and denotes if a call path was found.
Disclaimer, I'm a Co-Founder and have spent a great deal of time writing scanning code.
Static analysis refers to analysis techniques that understand code without recourse to execution, usually during compilation. (Static binary analysis does exist, but the difficulties of disentangling code from data in native code makes it usually impracticable). The Java FindBugs tools or Clang's static analyzer are good examples of static analysis techniques.
Dynamic analysis refers to analysis techniques that rely on gathering information using actual executions of a program. Even pedestrian techniques like profiling and code coverage are considered dynamic analysis techniques.
Binary analysis is a completely orthogonal concept. It refers to the analysis of programs without recourse to the source code, and is usually understood to refer almost exclusively to native machine code as opposed to bytecode languages like JVM or .NET IL. Most binary analysis tools are dynamic (such as Pin or Valgrind) but there do exist static binary analysis tools such as GrammaTech's CodeSurfer.
This doesn't really capture the whole picture of techniques, though. Symbolic execution and its step-child concolic execution aren't best thought of as static or dynamic analysis.
For example, a profiler for a Makefile should record the time taken for each build step as the Makefile is running. It's not about profiling the program being built.
But, back to the premise: it would be really helpful if you [the author] could illustrate security defects which can be detected using "dynamic analysis" which cannot be detected with "static analysis." Legitimate, actually exploited/able vulnerabilities would be ideal.
You only need to perform whatever security analysis is necessary after all your dependencies are resolved. This does not require anything "dynamic". You just call the packager manager, wait for all the dependencies to be resolved, then verify there are no vulnerable versions, which is just a matter of looking up the relevant pieces in some database somewhere.
Which is fine if that's what they're doing but this article seems to be just smoke and mirrors mostly.
* Many tools just scan a text file where you declare your dependencies, which misses transitive dependencies and won't tell you exactly which versions are in use
* They use incomplete datasets for vulnerabilities, like CVE. Most OSS projects don't create CVEs for vulnerabilities, so it is a mostly useless datasource.
Besides having a complete view of the dependencies, we have a research team that is finding and disclosing new vulnerabilities all the time, which you can see here: https://www.sourceclear.com/registry/explore
Most dependencies are specified like "network-info >= 0.2" (taken from a random example: https://github.com/bitemyapp/blacktip/blob/master/blacktip.c...), which means you don't know which version will actually be used until you build the project. Which is our point in this blog post. If you just scanned that text file, you'd just be guessing at which versions of open source libraries are actually being used.
Basically, never try to guess what the package manager is going to resolve. Just let it do its job, just as it does for your production builds, and use that information to look up any associated vulnerabilities.
As an aside, and echoed by a few other comments in the thread, what you're calling "dynamic analysis" is what I know as "static analysis", like what I want Coverity to do (watch the build process, and monitor the source that actually goes into each of my build artifacts). "Dynamic analysis" brings to mind something more like Valgrind, or some other tool that monitors and profiles a program during execution.
(you could think of stack as analogous to ruby's Gemfile.lock plus rbenv, or python's virtualenv - it makes the whole build reproducible)