Also, the limitations of Hyperscan might be quite noticeable. It doesn't support all pcre facilities (e.g. capturing and arbitrary lookaround). It has a considerable compile time - I wouldn't want to use Hyperscan to grep short files! Justin Viiret (another ex-Hyperscan team guy) has a blog post about this comparing our relatively heavyweight optimization strategy with RE2 (which gets down to the business of scanning a lot quicker than we do). You can find it here: https://01.org/hyperscan/blogs/jpviiret/2017/regex-set-scann...
(sorry about the giant 01.org "dickbar", we couldn't control that)
The upshot is that most people looking for big collections of regexen in huge amounts of data aren't really running 'grep' type tools. If someone wanted to do that, like I said, it would be a good project.
30 regex and 1 meg of data for example.
I'm somewhat curious as I have a couple of scraping things I do where I can compile the regex once and keep it hanging around (or save to/load from disk if that's feasible) for 3 to 5 minutes at a time.
However, the compile times of HS for 30 regexes might be not entirely trivial (maybe a second, probably less). A megabyte you'd probably see benefits but I think the benefits might not be drastic.
Probably the best way to find out would be to fire up hsbench (a tool that comes with Hyperscan now), figure out the slightly weird corpus format (sorry) and get some numbers for yourself on your own workload.
It would be an interesting project to add Hyperscan support to it. You would get all the extra goop you mention for free. I would be happy to give guidance on that if anyone's interested.