Probably the closest thing I'm aware of is handing around a sqlite file, but I'm a little uneasy using a format that's meant to be a database as a transfer format. Dolt looks promising here too. Are there other ways?
Probably the closest thing I'm aware of is handing around a sqlite file, but I'm a little uneasy using a format that's meant to be a database as a transfer format. Dolt looks promising here too. Are there other ways?
Note, we use mostly Python, some R, and a various range of ML or Optimisation tools, depending on the project.
We had to pick a single file format recommendation for sending 100GB+ tables on FTP servers or dropbox, scanning terabytes of useless stuff only to grap an key-value pair, and properly reading integer and UTF-8 columns. Turns out, Parquet is practical. Enough for users to start using it instead of CSV. It could be Avro, but it's just not as easy.
I actually think Parquet is pretty great in practice, I just have some issues with the sheer volume of abstractions necessary to implement it. I just wish it was anything other than Thrift.
I would probably choose Parquet over anything else, though.
https://medium.com/ssense-tech/csv-vs-parquet-vs-avro-choosi...
It's well-supported in Pandas and Kafka, has good schema support, and reasonably small compressed file sizes.
Behind the scenes, we store them as cstore_fdw [2] files which is a columnar storage format that helps with analytical queries.
[0] https://github.com/splitgraph/splitgraph/