Thanks for noticing ripgrep's performance on Unicode. :-)
> If you're just searching for an ASCII string literal in a single file rg probably gets slowed down by its UTF-8 validation.
This is most definitely wrong. Firstly, ripgrep doesn't do any UTF-8 validation. That would make it much much slower. Secondly, searching a simple literal is one of the cases where ripgrep should actually do better than GNU grep in many cases. This is because of the slightly smarter way in which ripgrep keeps its vectorized loop active more often than GNU grep does by trying to search rarer bytes. I talk about this quite a bit here, which even includes a benchmark where I measured ripgrep at 2x the speed of GNU grep for a simple literal search: https://blog.burntsushi.net/ripgrep/#subtitles-literal
Anyone can try this for themselves:
$ cd /tmp
$ curl 'https://object.pouta.csc.fi/OPUS-OpenSubtitles/v2016/mono/en.txt.gz' | gzip -cd > subtitles.en.txt
$ pv < subtitles.en.txt > /dev/null
9.28GiB 0:00:01 [4.89GiB/s] [======================================================================================================================================================>] 100%
$ time grep ZQZQZQZQ < subtitles.en.txt
real 8.713
user 7.495
sys 1.212
maxmem 9 MB
faults 0
$ time rg ZQZQZQZQ < subtitles.en.txt
real 1.857
user 0.573
sys 1.282
maxmem 9 MB
faults 0
On my system, /tmp is a ramdisk, so this doesn't factor in I/O. If you look up this thread, you'll note that someone was trying to search a 30GB file. Unless that person could fit that file into memory, it's very likely that they were just measuring the I/O bandwidth of their machine.