OpenRefine
openrefine.org
openrefine.org
I’m still excited to learn more about OpenRefine, but I guess maybe something like Google Colab might be better in terms of sharing and having direct access to our G Drives.
I had a look at how to use this. The video I watched is a couple years old, but probably mostly relevant still. https://youtu.be/nORS7STbLyk
The thing that really resonates with me here is the way they use faceting to find bad data.
When I write pipelines on the command line, I sometimes find it necessary to filter and select data in various ways. Because of this I end up rerunning cli pipelines multiple times sometimes. If instead I dump it to csv I could see myself using OpenRefine with its faceting to pick out the relevant data for processing
This is that lesson for getting started with OpenRefine. http://datacarpentry.org/OpenRefine-ecology-lesson/
Rewrite it in Rust+SQLite+Tauri+Typescript+Svelte?
The spread sheet ui is super useful and something that non technical people are much more comfortable dealing with. I've used Google sheets as an interface to business people over the years. Whether it is categorizations, place descriptions, addresses, etc. just put it in a spreadsheet.
Instead of building complicated UIs and tools, you just build a csv/tsv importer and let people do their thing in a spreadsheet, export, validate, import. Once you get it in people's heads that the column names are off limits for editing, they kind of get it. The nice thing about this stuff is that it is low tech, easy, and effective. And easy to explain to an intern, product owner, or other person that needs to sit down and do the monkey work.
Refine takes this to the next level. You can take any old data in tabular format and cluster it phonetically, minor spelling differences, or by other criteria, bulk edit some rows, and export it. It's also easy to enrich things via some rest API or run some simple scripts. But even just the bulk editing and grouping is super useful. We used it when it was still Google Refine more than 12 years ago to clean up tens of thousands of POIs. Typically we'd be grouping things on e.g. the city name and find that there would be a few spelling variations of things like München, Munchen, Muenchen, Munich, etc. Toss in a few utf-8 encoding issues where the ü got garbled and it's a perfect tool for cleaning that up.
Tens of thousands of records is potentially a lot of work but still tiny data. We had a machine learning team that used machine learning as the hammer for the proverbial nail. Google Refine achieved more in 1 afternoon than that team did trying to machine learn their way out of that mess in half a year.
It is SSR with Velocity templates.
What was dissapointing last time I used it is arguably not a problem of openrefine at all: the connectivity and response of wikidata queries was very slow. But that combination of local data harmomization with an open and globally available reference is super-important. I hope it somehow receives more attention and traction.
-There doesn't seem to be a canvas for visualizing the flow (graph) of transformations.
-It doesn't seem to support operations of 2 datasets, such as Join.
Isn't that rather limiting? Or have I missed something?
Does this allow you to save your steps and "re-play" them on a different dataset?
Im sure ot must be in the documentation somewhere but you can also read about it here https://guides.library.unlv.edu/open-refine/undo-redo