Grab – simple but very fast grep
github.com
github.com
ggreer@boron:~/code% du -sh .
8.3G .
ggreer@boron:~/code% time ag cpu_set_t
ag cpu_set_t 4.45s user 5.25s system 295% cpu 3.285 total
ggreer@boron:~/code% time grab -R cpu_set_t .
grab -R cpu_set_t . 13.31s user 21.67s system 35% cpu 1:38.28 total
30x faster, but these benchmarks aren't a fair fight. Ag ignores binary and hidden files by default. If I tell ag to do an unrestricted search, it's still 2x faster (43 seconds vs 98 seconds). Even with cold caches (echo 3 | sudo tee /proc/sys/vm/drop_caches between each run), ag beats grab handily: ag -u cpu_set_t 19.62s user 32.56s system 90% cpu 57.433 total
grab -R cpu_set_t . 15.48s user 37.89s system 37% cpu 2:22.67 total
I haven't profiled grab yet, but there's definitely some low-hanging fruit. For example, it looks like grab could get a big speedup by detecting literal patterns and using strstr() instead of a whole PCRE engine. Also, FileGrep::find is calling pthread_mutex_lock/unlock even if there are no matches to print. Adding a condition around that makes grab 1.5x faster.I'm glad I took a look at grab. Despite its shortcomings, I learned from it. I'll definitely try out a few tricks grab uses that ag doesn't, such as thread affinity.
One more thing: The author of grab is right about counting newlines. It does hurt performance. Still, I enable it by default in ag. I think the tradeoff is worthwhile.
1. https://github.com/ggreer/the_silver_searcher
Edit: The mutex locking change was so straightforward that I submitted a PR: https://github.com/stealth/grab/pull/2
the_silver_searcher is a great piece of software and I use it every day, but why would your main development platform be anything but Linux in this day and age? OS X has inferior performance, no official package manager and it's being produced by people hostile to open source developers. Why in the name of Jah would you subject yourself to that?
I really want* to use Linux, but I'm on OS X for a single reason: my living depends on my operating system reliably working well, after every update. OS X gives me this, and Linux, sadly, doesn't.
Also the UI is far better than any Linux distro can offer.
>Also the UI is far better than any Linux distro can offer.
I'm a programmer. I do most of my work in a terminal emulator. Openbox takes care of the rest.
Not any more: http://www.infinality.net/blog/infinality-freetype-patches/
Also, even though the ThinkPad X140e is certified by Ubuntu to be compatible[2], it took me 4 months of messing around before I could change the screen brightness. That issue ruined battery life and made my laptop unusable at night. I still can't get bluetooth to work. Others have been luckier than me, but hardware support on Linux can still be spotty. With a mac, I don't have to worry about that.
But that's just my experience. Others (such as yourself) love using Linux on their laptops. It's entirely possible that people simply have different preferences or workflows. Instead of feigning shock or getting upset, just use what you like.
1. http://abughrai.be/pics/DSC_8737.JPG
2. http://www.ubuntu.com/certification/hardware/201309-14195/
It's not shock, it's disappointment. You put up with this shit[1] because it's shiny?
[1]: http://openbenchmarking.org/prospect/1304096-FO-RARINGOSX81/...
Homebrew The missing package manager for OS X
If it's a crutch, it's a damn useful crutch that I can't live without on OSX, and that I'd use over any other package manager on OSX in a heartbeat.
This is probably faster on a lot of platforms with the common grep use-cases, yeah (i.e. very short literal strings). But as a common practice imo using the C library string functions in performance-sensitive code is tricky, because it can lead to some wildly variable performance between platforms, compared to directly using a portable-C implementation of a good algorithm.
I was bitten by strstr() in particular in the past. I had a large slowdown when taking something pretty simple (filter/extraction code) that ran fine on Linux, and trying it on FreeBSD. It turns out that strstr() on FreeBSD implements the naive O(nm) algorithm: a loop around strncmp() that tries a match at each possible starting location [1]. That's fine for short search strings but increasingly bad for long ones. I had assumed that a typical strstr() would use Boyer-Moore or something of that sort past a threshold, but on many platforms it doesn't (GNU libc, at least in recent versions, does switch to Two-Way [2]).
[1] http://svnweb.freebsd.org/base/head/lib/libc/string/strstr.c...
One thing I wish for is to cache the search result of a search term so that next search of the same term would be instant. This is good for large static file set.
Some of the optimizations are similar (mmap and pcre_study), while others are opposed (ag uses pthreads, grab claims disk I/O is the bottleneck and threads slow things down).
Ask, and ggreer (https://news.ycombinator.com/item?id=8781374) shall give (within 10 minutes). :-)
I suspect improving grep would have been a better idea than making tools that aren't as versatile as grep, but...