Semgrep: Semantic grep for code
semgrep.dev
semgrep.dev
For example, a webapp may have been designed such that authorisation needs to be explicitly added with a line or two to each controller. A semgrep rule can be written to match all the controllers which are missing this line. Then these controllers can be manually reviewed to assess whether unauthorised access should be allowed. Depending on what you are trying to match, this is something that may be very complex or even impossible to implement accurately in plain grep. Some languages like Ruby have powerful static analysis tools (Brakeman) that can also do this, but the benefit of Semgrep is the flexibility across multiple languages and how readable the rulesets are. [1]
[1] https://blog.includesecurity.com/2021/01/custom-static-analy...
I suspect there would be similar issues with your example of ensuring no use of eval() in PHP. So it seems okay to keep your own developers informed, but I wouldn't use it, alone, to vet outside code. PHP has eval-like functionality buried in preg_replace(), assert(), and probably other places. This tool also doesn't seem to dig into namespaced "aliases".
Semgrep does look at an AST; but that counterexample is not something you can "fix" solely by looking at an AST. You need actual Python-specific semantic analysis that knows that all "open" functions like 'print' come from the builtins module, and thus are bound to the same identifier. They're literally built into the implementation, it's not something you can "discover" from analyzing existing Python source. Even if you had a perfectly accurate python AST it couldn't "tell" you this fact, it's a priori knowledge, and all analysis engines need a base set of facts like this that they work from.
> but I wouldn't use it, alone, to vet outside code.
I mean, nobody seems to be suggesting this though, and the OP quite literally stated the major value of the tool is enforcing domain/codebase-specific rules among a team. Which is a really good use for it! There are tons of little useful patterns you can codify this way.
Perhaps I worded it poorly. Dumping the python AST for builtins.print() makes it pretty clear that it's "print" though. So I'm curious why that skirts the rule.
>I mean, nobody seems to be suggesting this though
Not specificially, but the context is using it for security purposes with phrases like "every use of a potentially injectable function such as exec,system,etc. in PHP". Felt like that was worth commenting on.
No, the AST only tells you it's a method call on something called "builtins." You need the separate semantic knowledge of what builtins is in order to figure it out. Parsing + AST just means it sees "method call of `print` on `builtins` object". Regular print calls would come through as "regular function call of `print`".
To echo the other replies, the AST for builtins.print() is the same as the ast for mymodule.print() and, in fact, if you stick a builtins.py in the right place, you'll be able to prevent the import of the standard library builtin module, while the ast's would be identical.
Writing the parser is nontrivial, but once you have it it should be straightforward to expose a programmatic API for doing this stuff instead of trying to hardcode every useful linter rule into a single program.
In the JVM world annotation processors or compiler plugins can have access to the AST during compilation. Both in Kotlin and Java.
All that's needed is a compile_commands.json file which can be easily generated via most build systems, or you can use Bear [3]/some other tool (or write a script that logs all syscalls and generate it yourself).
[0] https://releases.llvm.org/12.0.0/tools/clang/tools/extra/doc...
[1] https://releases.llvm.org/12.0.0/tools/clang/docs/LibTooling...
[2] https://releases.llvm.org/12.0.0/tools/clang/docs/LibASTMatc...
https://github.com/elanning/checkr
It is just simple regex at this time, but hopefully I can add something like CCGrep syntax in the future:
"GitLab SAST historically has been powered by over a dozen open-source static analysis security analyzers. These analyzers have proactively identified millions of vulnerabilities for developers using GitLab every month. Each of these analyzers is language-specific and has different technology approaches to scanning. These differences produce overhead for updating, managing, and maintaining additional features we build on top of these tools, and they create confusion for anyone attempting to debug.
The GitLab Static Analysis team is continuously evaluating new security analyzers. We have been impressed by a relatively new tool from the development team at r2c called Semgrep. It’s a fast, open-source, static analysis tool for finding bugs and enforcing code standards. Semgrep’s rules look like the code you are searching for; this means you can write your own rules without having to understand abstract syntax trees (ASTs) or wrestle with regexes.
Semgrep’s flexible rule syntax is ideal for streamlining GitLab’s Custom Rulesets feature for extending and modifying detection rules, a popular request from GitLab SAST customers. Semgrep also has a growing open-source registry of 1,000+ community rules.
We are in the process of transitioning many of our lint-based SAST analyzers to Semgrep. This transition will help increase stability, performance, rule coverage, and allow GitLab customers access to Semgrep’s community rules and additional custom ruleset capabilities that we will be adding in the future. We have enjoyed working with the r2c team and we cannot wait to transition more of our analyzers to Semgrep. You can read more in our transition epic, or try out our first experimental Semgrep analyzers for JavaScript, TypeScript, and Python.
We are excited about what this transition means for the future of GitLab SAST and the larger Semgrep community. GitLab will be contributing to the Semgrep open-source project including additional rules to ensure coverage matches or exceeds our existing analyzers."
The web page states: "Static analysis at ludicrous speed. Find bugs and enforce code standards"
"grep" is short for "global regular expression print". It finds matches for the given regular expression and prints them.
"Semantic Grep" is a static analyzer with configurable rules, style checks, etc. It does much more than search and print.
Perhaps a better name is needed?
Edit: How about "omnilint" or "omnicritic" since semgrep is more of a "lint" (https://en.wikipedia.org/wiki/Lint_(software)) or "critic" (https://en.wikipedia.org/wiki/Perl::Critic) type of tool that handles multiple languages?
Edit2: "Static analysis at ludicrous speed" ==> "turbolint"? ("ludicrous speed" reminds of the hilarious Space Balls scene :) "turbolint, GO!"
But at least to me, semgrep looks a lot more like "lint" than "grep".
Agreed, this project name is misleading about what it does. The name "grep" always indicated some kind of "find a text/pattern and print results to stdout" utility. Like pgrep, which searches running processes by name and then prints their IDs.
Hey, I'm a maintainer of Semgrep, and this sounds like a pretty good description of what the CLI can do, see this example for finding all function/class/method calls:
$ semgrep -e '$NAME(...)' -l python
flask_todomvc/extensions.py
4:db = SQLAlchemy()
------------------------------------------------------------
5:security = Security()
flask_todomvc/factory.py
15: app = Flask(__name__)
------------------------------------------------------------
17: app.config.from_object(settings)
------------------------------------------------------------
18: app.config.from_envvar('TODO_SETTINGS', silent=True)Looks like a pretty useful tool with a couple nice options. A bit strange that the `-e` option is only explained on the website, but to be fair it seems to be a lot to cover. Still, a kind of "cheat sheet" style summary in the help message would be fantastic, just as a little suggestion.
This is like saying that PubNub is a bad name because it does messaging, and has little to do with pub-sub. Or hell, even that Y Combinator should be called something like StartupFactory since it's not really a recursive tool.
In short, the metaphor's close enough.
https://semgrep.dev/docs/extensions/ describes how to do pre-commit.
Nvm, here's semgrep's own .pre-commit-config.yml for semgrep itself: https://github.com/returntocorp/semgrep/blob/develop/.pre-co...
`.git/hooks` directory in your repo for samples, e.g. `.git/hooks/pre-commit.sample`.
You can run any old shell script there, without having to install a python tool.
Pre-commit requires Python and pre-commit to be installed (and then it downloads every hook function).
This fetches the latest version of every hook defined in the .pre-commit-config.yml:
pre-commit autoupdate
https://pre-commit.com/#pre-commit-autoupdateA person could easily `ln -s repo/.hooks/hook*.sh repo/.git/hooks/` after every git clone.
[pip,] install pre-commit
pre-commit install
# git commit
# pre-commit run --all-files
# pre-commit autoupdate
https://pre-commit.com/Like, if you've never tasted lychee, it would never occur to you how to cook with it.
I'm going to need to see some useful, real-world examples to jumpstart my brain to think this way.
Alternatively, we curate 1000+ community rules that you can look through as well.[1]
[0]: https://github.com/hashicorp/terraform-provider-aws/blob/mai...
The starting point of semantic grep is very useful. When you have a big codebase, you often want to detect antipatterns, or not even antipatterns, but just uses of a thing, say you're renaming a method and want to track down the callers.
Being able to act on the AST, instead of hoping you searched up all of the variants of whitespace and line breaks and, depending on the specific example, different uses of argument passing, is really useful.
But often when you're semantically grepping, your goal is to replace something with something else (this is what refex was initially built for: to aide in large scale changes in python, as a sort of equivalent to the C++ tools that Google uses).
But then you want to shift left even further: once you have a pattern that you want to replace once, you can just enforce that a linter yell at you when anyone does it again. So it's very natural to develop a linter-style thing on top of one of these[2].
This is, as I understand it sort of the same thing that happens in C++: clang-tidy and clang-format are written on top of AST libraries that can be used for ad-hoc analysis and transformations, but you can also just plug them into a linter.
The thing is, for most organizations, enforcing code style and best practices is more valuable than apply a refactoring to 10M lines of code, because most organizations don't have 10M lines of code to refactor. That doesn't mean that these tools aren't also useful for ad-hoc transforms and exploratory analysis. They absolutely are!
Wait, is this a web app? I was expecting a command line tool to navigate my code locally.
You can write stuff like
# Check for Python == where the left and right hand sides are the same (often a bug)
$ semgrep -e '$X == $X' --lang=py path/to/src grep -E '(.+) = \1' *.py
I have trouble looking at the examples in the project website (many things inside iframes are adblocked). Do you have any example of a search that would be difficult or impossible with grep? import random
y = 0
def f(x):
print(x)
if random.randint(0,1) == 2:
y = 1
f(y)
This not only fails but crashes the program.given:
if 1 \
== \
1:
pass
then $ semgrep -e '$X == $X' -l python
1:if 1 \
2: == \
3: 1:
ran 1 rules on 1 files: 1 findings
to follow myself up, one can also check for expressions like what I used as an example and say "don't do that", but without regard to what is inside the if semgrep -e $'if ...:\n pass\n' -l pythonStill, it seems rather cool, I like the idea of being able to search code at a higher level than just raw source text.
Here’s an elaborate example: https://semgrep.dev/s/ievans:c-dataflow
Lots of workarounds it wouldn't find, like:
import builtins
builtins.print("whee")I don't think you can go in with the mindset that it will catch everything, but rather, it's about being able to iterate quickly with your rules.
I've found the flake8 API and documentation lacking, so perhaps just a cleaner interface?
Go down, see "brew install semgrep" and try to copy paste it. And it's an image :(
Go to https://semgrep.dev/p/jwt
Go to the page 2/5
Click "Run Locally", so you can copy the code
close the modal -> you're on page 1/5. Expectation would be to stay on page 2/5.
It would also be very useful to be able to filter by language and topic.
How does Semgrep compare to ESLint+a strict tsconfig?
If the commit hook rejects anything where rules are triggered, a way to force the comrit is needed for the cases when the rule finding is not reanny an issue.
Upd: I found that the --no-verify option can be used in many cases
E.g. const output = console.log; output(“hello”);
This wouldn’t be caught by ESLint (to my knowledge) but would be caught by Semgrep. I think you could do it with ESLint but given the interface for an ESLint plugin exposes an AST, you’d have to track this yourself. I’m assuming Semgrep could stretch to things like enforcing APIs are called with certain optional arguments present (even if the TS types don’t require it). Again, I think with ESLint you’d have to do more juggling with the raw AST.
That said, this is my understanding from a quick skim!
(I'm guessing from your comment that this is important to you (i.e. WSL/Docker is not a solution).)
Also kind of surprising it's written in Python given that they advertise its speed.
Add to that that the reason things fail on Windows is usually something Windows specific and “their fault.” So it’s unusual for me to fire up a Windows VM just to sanity check my code. And since Windows CI runners cost 2x or more, I don’t usually run cross platform CI.
Free projects only get 2,000 minutes per month so running on all three platforms means you only get 1/13th of the minutes of only using Linux.
I rarely use anything but Linux runners even on my paid projects. I like saving money, so unless I really need integration testing on Windows or Mac, I don’t do it.
[0] https://docs.github.com/en/github/setting-up-and-managing-bi...
On the speed front, the core of Semgrep is OCaml, with a Python wrapper to add niceties like pattern composition. I’m less close to that part of the codebase, but think more and more logic is being moved to OCaml for performance reasons.
It blows my mind how fast it is compared to many tools in js ecosystem. Tree-sitter was parsing millions of files in half a minute. JS, TS, Ruby, yaml, html, Css. It’s quite magical. Such great engineering.
try {
const parsedURL = new URL(url)
requestPath = parsedURL.pathname
} catch (error: unknown) {
// NOOP
}
It's complaining about : unknown bit which one of the newer typescript eslint rules enforces. import random
if random.randint(0,1) == 2:
print(“hello”)
Is also unparseable.Perhaps a more pointed set of questions: what is the error it emits, and have you considered submitting that case as an actual bug?
Anyone else know of a Service linting tool? OPA/conftest come close but lack syntax parsers for Ruby/Javascript.
Is there a way to search for functions in C (other than printf!) whose return value is ignored at the call site?