Otherwise it seems more flexible to just fire up a python interpreter and do it in like 3 lines of pandas (or with sqlite and .import like another commenter mentioned)
Otherwise it seems more flexible to just fire up a python interpreter and do it in like 3 lines of pandas (or with sqlite and .import like another commenter mentioned)
This tool is weird because if I’m going to download some tool that’s not in the distro, I’d rather use SQLite or try to figure it out with awk.
I definitely used to use `sqlite` for every CSV before and I've shifted to using `xsv`. If I want to do any heavy lifting I'm going to pop it into PostgreSQL anyway.
Crucial feature is that it's easy to pop some `xsv` pipeline in the middle of your `for i in ...` or `find . -print0 | xargs -0`.
Sometimes I also use xsv to just do a step of the analysis and dive deeper on some subset using pandas.
In my experience both SQLite and Pandas aren't as fast as fast for large files. So they are not really good options.
Pandas is especially bad because it uses a column oriented data structure internally so reading from or writing to CSV is incredibly slow in Pandas. If you can use parquet that's not a problem but unfortunately parquet is not nearly is ubiquitous as csv :(
Too big for excel is not big data, and my laptop can load this 10G in RAM (not that it necessarily need all of it) so why not if the data is here and the laptop on your lap ?
Often this resulted in one-off data loads or reporting metrics. It was almost always easier just to knock something out with command line tools than it was to write, test, and deploy code. The deployment process itself could take longer than actually processing the file.