Odo: Shapeshifting for your data
odo.readthedocs.io
odo.readthedocs.io
I ended up learning how to do this from the various SQL shells. But, that was a bit of a cognitive load, especially when the CSV files were complicated. AFAIK, for example, SQLite won't properly ignore commas in quoted fields, so you have to throw in an extra utility (like csvkit) to change the delimiter before importing into SQLite.
sidenote: through Odo's homepage, I discovered this amazing library for generating network graph diagrams, NetworkX: http://networkx.github.io/
(Here's my Gephi map of NSFW subreddits according to their links to each other: http://electronsoup.net/nsfw_subreddits/)
Full disclosure: I still don't.
Odo is a shape-shifting being who is removed from his home planet. But we don't find that out til the 3rd season.
Min/max limits, truncation nulls, floating point precision, encodings, picking CHAR vs VARCHAR or string vs categorical, metadata like indices, etc are some of the hard problems behind bulk loading.
By the way, I have problem when opening that site. The site has problem with SSL certificate on Mozilla Firefox.
Part of the problem is the scale, but another part is that writing partitioned parquets seems poorly documented (I would love corrections, I spent a decent amount of time last week looking for good information)
It's literally one line of code. See http://labs.vistarmedia.com/2016/12/27/indexing-json-logs-wi... for an example (except you write to a local file system rather than HDFS).