I had no idea anything at AWS had that long of an attention span.
It's funny and telling that in the end, it's all backed by CSVs in s3. Long live CSV!
I had no idea anything at AWS had that long of an attention span.
It's funny and telling that in the end, it's all backed by CSVs in s3. Long live CSV!
Also, we mostly have Parquet data cataloged in S3 today, but delimited text is indeed ubiquitous and surprisingly sticky, so we continue to maintain some very large datasets natively in this format. However, while the table's data producer may prefer to write delimited text, they are almost always converted to Parquet during the compaction process to produce a read-optimized table variant downstream.
CSV is only good for append only.
But so is Parquet and if you can write Parquet from the get go, you save on storage as well has have a directly queryable column store from the start.
CSV still exists because of legacy data generating processes and dearth of Parquet familiarity among many software engineers. CSV is simple to generate and easy to troubleshoot without specialized tools (compared to Parquet which requires tools like Visidata). But you pay for it elsewhere.
1. Dynamically typed (with type affinity) [1]. This causes problems with there are multiple data generating processes. The new sqlite has a STRICT table type that enforces types but only for the few basic types that it has.
2. Doesn't have a date/time type [1]. This is problematic because you can store dates as TEXT, REAL or INTEGER (it's up to the developer) and if you have sqlite files from > 1 source, date fields could be any of those types, and you have to convert between them.
3. Isn't columnar, so complex analytics at scale is not performant.
I guess one can use sqlite as a data interchange format, but it's not ideal.
One area sqlite does excel in is as a application file format [2] and that's where it is mostly used [3].
[1] https://www.sqlite.org/datatype3.html