Wikidata as a Giant Crosswalk File
dbreunig.com
dbreunig.com
> However, for reasons unknown to me, they wrap these neatly separated rows with brackets ([ and ]) and add a comma to each line
Well, the reason (misguided or not) is as you say, I imagine:
> so it’s a valid, JSON array containing 100+ million items.
> We are not going to attempt to load a this massive array. Instead, we’re running this command:
zcat ../latest-all.json.gz | sed 's/,$//' | split -l 100000 - wd_items_cw --filter='gzip > $FILE.gz'
That's one approach - I'm always a little wary of treating a rich format like JSON as <something> deliminated text - I'd be curious if using jq in streaming mode is much different in run-time. I believe this snippet, the core of which we lifted from stack overflow or somewhere does the same thing; split a valid JSON array into ndjson (with tweaks to hopefully generate similar splits: gunzip -c ../latest-all.json.gz \
| jq -cn --stream \
'fromstream(inputs|(.[0] |= .[1:]) | select(. != [[]]) )' \
| split -l 100000 - wd_items_cw --filter='gzip > $FILE.gz
Note on MacOS zcat might not be gunzip, hence the change.[1]: https://github.com/jimhigson/oboe.js
EDIT: looking around a bit, I found json-stream ( https://github.com/dgraham/json-stream ) for Ruby.
I love seeing these quick and dirty Ruby scripts used for data processing / filtering or whatever, this is what it is good at!
https://mix-n-match.toolforge.org/
It let's you review existing potential matches as well as upload CSVs or generate regex web scrapers to ingest database IDs to be linked against others.
Here some related to places, since the article is geographical:
https://mix-n-match.toolforge.org/#/group/ig_authority_contr...
But it's got all sorts of people places and things, anything that someone might have built a catalog or list of.