Don’t underestimate grep-based code scanning
littlemaninmyhead.wordpress.com
littlemaninmyhead.wordpress.com
I work a lot with Go, where all code in our repository is gofmt'ed. You can get quite far with regular expressions for finding/analyzing Go code.
(And when regexps don’t cut it anymore, Go has excellent infrastructure for working with it programmatically. http://golang.org/s/types-tutorial is a great introduction!)
That said grep is pretty fast period, so probably doesn't make a huge difference in practice, especially if you're IO-bound, which is common.
But the skip-ahead stuff isn't the most important thing nowadays. The key is staying in the fast vectorized skip loop as long as possible.
I've seen function names misspelled, and then every invocation just doubling down on that misspelling.
The amount of cross team coordination is staggering.
I thought code reviews are the only way and then you need to have every Dev aligned and on the same page on this topic... Which never happens. :/
But I do love auto-formatters. (I'm doing web dev at the moment, so Prettier.) It is freeing to not worry about spacing, line breaks, parens, etc.
All I have to do is give the computer a valid AST and it does The Right Thing.
And if I'm doing whitespace repetition, there's no great advantage in an autoformatter.
Well that makes sense, the less reliable your input text, the more complex the regexp.
ag FooBar | grep -v Baz
It's in brew/apt/yum etc as `the_silver_searcher` (although brew install ag works fine too).Well, there's no true hierarchy/succession order but rg was written after ag was.
ag provides sane default and settings for developers. grep is ubiquitous and great, but to do what most developers want it requires some guidance, whereas ag focuses on being what you want most of the time.
What I mean by that is that I enjoy the smart-case sensitivity (as in, if there are not caps in my pattern, then it defaults to case-insensitive but if I have any caps in my pattern, it uses case-sensitive) or fast filename searching with -g or both with -G.
I've tried rg and while it was faster, it also didn't provide as good support for filename searching. Ag is still what I consider "so fast I almost don't believe it"
It's like the 'tldr' command (https://tldr.sh/). Of course I still use man pages, but having something that gets me what I really need very quickly is important.
Steps:
1. Use ag
2. If ag isn't present, try to install ag
3. If I can't install ag, then use grep or find, no big deal.
ag > ack > grep/find
Could you elaborate on this? Is it because you need to type more? If so, I'd suggest one of two things. 1) use `fd` for searching for files, which is dedicated to that purpose. 2) define `alias rgf="rg --files | rg"` (or similar) and use `rgf foo` just like you would use `ag -g foo`.
I do wonder if there is a multi-language aware one though.
int
foo_func(void) {
You can grep for `^foo_func\b` to get to a declaration or definition, or `^foo_func\b.* {$` to get to a definition or `^foo_func\b.* ;` to get to a declaration. This is instead of using something like `^\w.* \bfoo_func\(`, which is what you'd need for: int foo_func(void) {
By the way, anyone know of a way to insert a literal asterisk here without having to follow it up with a space? fooFunc(): Int {
"fooFunc returns an integer."Not
"An integer is returned by fooFunc."
This is were you could be wrong. We would need to give a reason for dismissing it and then the risk officer would need to approve it (or reject it). False positives can be a real pain in the ass.
Grepping leans-in to shell. Though if you have other environments available (python, javascript etc), it makes sense to lean-into them e.g I use JavaScript examine my package.json to ensure my dependency SemVers' are "exact".
That said, I rarely write static-analysis scripts: In JavaScript-world there is already a plethora of easily configurable linting & type-checking tools. If I wanted to focus in on static-analysis etc I'd probably reach for https://danger.systems/js/
SideNote: My CI generates a metrics.csv file, which serves as a "metric catch-all" for any script I might write e.g. grep to count "// TODO" and "test.skip" strings, plus my JavasScript tests generate performance metrics (via monkey-patching React).
I don't actually DO ANYTHING with these metrics, but I'm quite happy knowing the CI is chugging away at its little metric diary. One day I'll plug it into something.
While there's no install and initial results are quick to appear, the false positives that grep or any string search tool generates will make the cynics shoot down this simple attempt to find problems in the source code.
Problems that arose:
- what about use of those questionable APIs/constants in strings (perhaps for logging) or in comments?
- some of the APIs listed in the article were only questionable when certain values were used - sometimes you can get grep/search tool of choice to play along, but if the API call spans multiple lines or the constant has been assigned to a variable that is used instead, then a plain string search won't help.
- it's hard to ignore previously flagged but accepted uses of the API/constants.
- so there's a possible bug reported, but devs usually want to see the context of the problem (the code that contains the problem) quickly/easily. Some text editors can grok the grep output and place the cursor at the particular line/character with the problem, some can't.
If you go down that road to try and reduce false positives, you'll end up with a parser for your development language of choice.
My SAST generates tons of false positives and is unforgivably slow. If this is orders of magnitude faster, it might be worth the extra false positives.
As a side note, my dream is a SAST that comments directly in the PR like a human reviewer would. Maybe that exists?
If the SAST has to process C/C++ source code, then the SAST will parse all the #include'd header files. The SAST may track values to determine if illegal/uninitialized values are used.
A string search tool will skip doing all of that.
If the class of problems you're looking for contains only bad functions/constants, then a string search tool may be fine.
But as I mentioned before, the string search tool may get confused if these bad strings occur in strings/comments/irrelevant #if/#else/#elif sections.
There are another class of bugs dealing with data values which a string search tool can't deal with easily.
As an example, PC-Lint lists the type of problems the program may flag - https://www.gimpel.com/html/lintchks.htm. A string search tool won't know about classes and virtual destructors or other concepts relevant to the programming language in question.
For the string search tool, you'd either invoke the search string tool several times with different search strings for the same source code or slightly more efficient, have one long search string containing all your search strings as alternate search targets for the string search tool.
Either case, when the string search tool spits out a positive result, it won't explain why there is a problem. The dev will have to know or lookup the problem associated with that search result.
When I worked on this area, C/C++ compilers stopped at syntax errors. Most have gotten better at flagging popular problems like variable assignments within if statements, operator precedence bugs, and printf-format string bugs.
Some divisions at Microsoft required devs to run a lightweight SAST before committing changes to locate possible problems ASAP.
It's relatively easy to integrate an SAST into your build system to scan the modified source code just before you're ready to commit the changes.
> the resulting string in dest is always null-terminated.
If you want to copy out the first line from a buffer that happens to be a 10TB mapped file, that strlcpy call will take a long time to finish. If you are using strncpy/strlcpy because you don't trust the src buffer is properly null terminated but you still want to stop the copy at the first null or when the buffer is full, well, you're out of luck because strlcpy is going to blast past the end of the source buffer regardless.
I would have been much happier if it had just returned a flag indicating either successful copy (0), buffer was truncated (1), or an error occurred and errno was set (-1). Possible errors could be that the src or dest was NULL or the size was 0 (ERR_BAD_ARGUMENT).
So, you have to figure that out. It isn't hard (hell, it's trivial) but I think if you're either going to be aware of the pitfalls — and then these functions are mostly not going to help you — or you're not, in which case you're just as likely to pass the wrong value for the size (dest's size/src's size) and overflow the buffer anyways.
Honestly, if I had to do more than a trivial amount of string manipulation in C, I'd be wrapping that in a mini library to manage some sort of stronger string type or finding such a library (glib? ICU?) very quickly, depending on needs. std::string was one of the things in C++ that made me question why anyone was still using C, given how much less error-prone it is, comparatively. (std::string is not without problems / only as compared to char * in C.)
It's very very fast / almost instant even with hundreds of source code files and millions of lines of code.
I hit ctrl+T and then can search everything, this give me a drop down that filters out the more I type, select the item in the dropdown and it goes to that source file.
I can also type:
/t and search just types
/m members
/mm methods
/u unit tests
/f file
/fp project
/e event
/mp property
/mf field
/ff project folder
e.g.
/t Foo
will find all the Foos
/mm SavePhoto
will find any methods called SavePhoto
Same works in JetBrains Rider for C# stuff.
I couldn't dev without this now, and it's all built into my IDE.
It provides a tool called sgrep (syntactical grep) that lets you do some cool tricks, for instance:
`sgrep "some_func(X,X)"` returns all calls to some_func with the same argument repeated.
`sgrep "some_func(X,Y,...)"` returns all calls to some_func with 3 or more arguments.
It's come in very handy for refactoring some troublesome codebases.
Excerpt: "The result of this is that, in the limit, GNU grep averages fewer than 3 x86 instructions executed for each input byte it actually looks at (and it skips many bytes entirely)."
However note https://news.ycombinator.com/item?id=19522987
> It does not. ripgrep does not use Boyer-Moore in most searches.
> In particular, the advice in [the freebsd mailing list post] is generally out of date.
although the out of date bits are really the Boyer-Moore ones: https://lobste.rs/s/ycydmd
> much of Mike Haertel’s advice in this post is still good. The bits about literal scanning, avoiding searching line-by-line, and paying attention to your input handling are on the money.
> But the stuff about Boyer-Moore is outdated.
Try
/usr/bin/find . -depth -type f -print | /usr/bin/xargs -i /usr/bin/fgrep string '{}'
and run it several times so that the filesystem cache is primed.
That's BS. The fgrep on my system – GNU grep 3.1 – provides recursive search (-r). What now, are you claiming that's not "real fgrep"? [1]
> that would be implementing tools within tools, which is against the UNIX®️ philosophy
Even more BS. Or are you telling me that "rm -r" is also against the "UNIX philosophy"?
> /usr/bin/find . -depth -type f -print | /usr/bin/xargs -i /usr/bin/fgrep string '{}'
Terrible, terrible advice. Cumbersome, error-prone, and slow as molasses. A quick test: Searching for 'asdfadsgf' in the Linux kernel repository takes 0.25 s using rg, 12 s using GNU fgrep -f, and 228 s (!) using your command.
You know, when your ideology results in the worst results of all, you should really reconsider your ideology.
GNU stands for GNU is not UNIX®️.
Edit: this highlights the big weakness in the "UNIX philosophy", in which the only record delimiter that's conventionally recognized in pipelines is the newline but the shell recognizes characters as filename delimiters that are also allowed in filenames. Causing a cascade of delimiter bugs. Sometimes you really do need a bit more structure to your data.
(The UNIX philosophy is best understood in contrast to what went before - the COBOL or JCL style where files had fixed records, in turn based on fixed-column punchcard layouts.)
[0] https://mywiki.wooledge.org/BashPitfalls#for_f_in_.24.28ls_....
Yeah, spaces are nasty. find has -print0 and xargs has -0 to handle this gracefully, but one needs to know to use it.
foo | sort
with foo > file
sort < file
I like the idea of powershell, but every time I try to do something complicated with it I'm disappointed. $WhateverObject | Export-CliXml
$WhateverObject = Import-CliXml whatever.xml
Most of the time `>` also work.By the way, PowerShell serialization is by orders of magnitude better then anything *nix has to offer as you can use objects from other machine shell just like they exist on your local one.
> I like the idea of powershell, but every time I try to do something complicated with it I'm disappointed.
I did some very complicated things in PowerShell. For example, check out the script that maintains ~300 mainstream packages on Chocolatey up to date, all in few minutes with bunch of self maintenance features.
You need to learn it, simple as that.
https://gist.github.com/choco-bot/a14b1e5bfaf70839b338eb1ab7...
Which you'd know if you'd wondered, because that's something /u/burntsushi regularly explains: https://blog.burntsushi.net/ripgrep/#literal-optimizations
https://blog.burntsushi.net/ripgrep/
(It seems down at the moment you can try the cached version)