Introducing Semgrep and r2c
r2c.dev
r2c.dev
Having a (fast) single tool that can accurately parse most commonly used programming languages is incredibly useful, but it requires the maintenance of dozens of grammars, which is difficult without a large community effort. Hopefully increased adoption means more accurate parsers and support for even more languages.
Tree-sitter powers syntax highlighting on GitHub.com and (soon) neovim and OniVim 2. Hopefully regex-based syntax highlighting is a thing of the past soon. If you haven't seen the Strange Loop conference talk on tree-sitter [2] yet, it's worth a watch.
I think a Prettier-like code formatter using tree-sitter would be cool, both in terms of potentially broader language support and native performance.
If you can write code in a language, you can use semgrep. It also has a feature I have learned to love every time I find it in any kind of auditing tool: it’s ruthlessly effective as an exploratory and experimental tool, but it takes no effort at all to turn that into a persistent check. By comparison: ripgrep finds anything fast, but nobody uses it to write linters. Other off the shelf linters do a great job finding (simple) issues, but bandit doesn’t help me one bit to build a mental map of how a codebase works.
Do gcc/clang/any other preprocessor create "source maps" that could facilitate that? GCC looks like it has a `-fdebug-cpp` that "[...] dumps debugging information about location maps. Every token in the output is preceded by the dump of the map its location belongs to."
This thread would love to learn about Compiler Explorer https://gcc.godbolt.org/ which works for C++ and many other languages.
I don't write or even read a lot of C++ these days but I recall from when I did that a major pain point was deciphering compiler warnings/errors when there are a lot of templates, macros, or both. Seems like the problem has been around forever.
Some prior art for reference: https://github.com/bytedeco/javacpp/issues/51
Edit: There is a help entry: https://semgrep.dev/docs/faq/#how-is-semgrep-different-from-...
It doesn't appear to catch the following when searching for exec(...) in the following python code:
not_exec = exec
not_exec('rm -rf /')
Edited to include language $ semgrep -e "not_exec('somestr)"
will match foo = "somestr"
not_exec(foo)
Here's a more complete example: https://semgrep.dev/s/ievans:const-pythonIn your example, we don't propagate exec because it's not seen as a literal -- that's a TODO for sure. See https://github.com/returntocorp/semgrep/issues/1645 for a longer discussion!
Here are docs for what exists currently https://semgrep.dev/docs/experiments/#autofix
day = 'friday'
I want it to find
day="friday"
also!
I'd expect latency might be juuust in the range where it doesn't feel interactive yet? But honestly any search that isn't ripgrep or --omg-optimized-etags feels like that to me now, and people use symbol rename features in IDEs all the time that take multiple seconds, so maybe I'm just unreasonably picky.