My takeaway is that csv has some undefined behaviours, and it takes up space.
I like that everyone knows about .csv files, and it's also completely human readable.
So for <100mb I would still use csv.
I like that everyone knows about .csv files, and it's also completely human readable.
So for <100mb I would still use csv.
I think in polars it's
df.filter(pl.col(pl.Utf8).str.len_bytes() == 0).shape[0] == 0
although there's probably a better way to write this.