- ripgrep PATTERN | xargs jq
- find PATTERN -exec jq
In both cases you have a large amount of data in your file system that is in JSON and you want to extract a subset of it for further processing. Ripgrep is an extremely fast way to do a content based search for a pattern and find is a fast way to do a file path based search. Then jq lets you extract the data.
For other types of data I use the same technique of first using ripgrep to find the candidate files and then piping to a text processor like awk/perl/ruby to actually process the data.
If you need FTS rather than regex search then SQLite FTS5 is my go to.
So a full example is something like:
`find . -name “*.pod.json” -print0 | xargs -0 -P 12 -I {} sh -c ‘jq -r “select(.spec.containers != null) | .spec.containers | to_entries[]” sh {} \; | jq -s ‘sort_by(.image)’`
Something like that so it’s sort of like a map-reduce you first narrow the subset of inputs by first finding by file pattern, then you pull out the relevant data from each in a parallel xargs, then you reduce it with a jq -s. This technique is used because jq is very slow on large files and your later processing scripts might be slow so the first step of a good pipeline is to first throw away all the data you don’t need first.
The last time I tried, I think the reason I gave up on JQ for large inputs was that the throughput would max out at 7mb/s whereas the same thing with spark SQL on the same hardware (MacBook) would max out at 250mb/s. So I started looking into using other solutions for big data while I use jq in parallel for small data in multiple files.
I will test it out again cause this was 4-5 years ago when I last tested it, but I believe jaq is still preferred for large inputs. Still I prefer for big data to use Spark/Polars/clickhouse etc.
So the pattern is to reduce the candidate set of files with rg or find then extract the relevant parts using xargs jq then pipe to jq -s to produce the dataset.