Time to retire the CSV?
bitsondisk.com
bitsondisk.com
1) A truly open format is available and accessible. Csvs are textfiles. There is no system around that cannot open a textfile. If the format is binary or requires patents or whatever, then it's a non-starter.
2) Applications have a speed increase from using csvs. To wit, I loved csvs because often they finish preparing much faster than a "formatted" output in excel, etc. and sometimes I just want to see the numbers, not have everything colored or weird features like merged cells in excel, etc.
3) The new format should not be grossly larger than the one it is replacing. Extracts in Excel format are orders of a magnitude larger than csv in filesize. This affects run-time to prepare the extract, open it (memory constraints, etc.)
Is there truly a problem? The author is not forced to use csvs.
These other data formats all win on your #2 against CSV, as CSV is actually horrible at parse-time vs. other data formats — the fact that both of its separators (newlines and commas) can appear as-is inside column values, with a different meaning, if those column-values are quoted, means that there's no way to parallelize CSV processing, because there's no way to read-ahead and "chunk" a CSV purely lexically. You actually need to fully parse it (serially), and only then will you know where the row boundaries are. If you've ever dealt with trying to write ETL logic for datasets that exist as multi-GB CSV files, vs. as multi-GB any-other-data-format files, you'll have experienced the pain.
> The new format should not be grossly larger than the one it is replacing.
Self-describing formats like JSON Lines are big... but when you compress them, they go back to being small. General-purpose compressors like deflate/LZMA/etc. are very good at shearing away the duplication of self-describing rows.
As such, IMHO, the ideal format to replace ".csv" is ".jsonl.gz" (or, more conveniently, just ".jsonl" but with the expectation that backends will offer Transport-Encoding and your computer will use filesystem compression to store it — with this being almost the perfect use-case for both features.)
-----
There's also Avro, which fails your point #1 (it's a binary format) but that binary format is a lossless alternate encoding of what's canonically a JSON document, and there are both simple CLI tools / and small, free, high-quality libraries that can map back and forth between the "raw" JSON document and the Avro-encoded file. At any time, you can decode the Avro-encoded file to text, to examine/modify it in a text editor.
The data-warehouse ecosystem already standardized on Avro as its data interchange format. And spreadsheets are just tiny data warehouses. So why not? ;)
For CSV file replacements, I'd expect something like "one JSON array per line, all values must be JSON scalars". In that case, it's not much larger than a CSV, especially one using quotes already for string values.
But this demonstrates the problem with JSON for CSV, I suppose. Is each line an object? Is it wrapped in a top-level array or not? If it is objects, do the objects have to be one line? If it is an object, where do we put field order? The whole problem we're trying to solve with CSV is that it's not a format, it's a family of formats, but without some authority coming in and declaring a specialized JSON format we end up with a family of JSON formats to replace CSV as well. I'd still say it's a step up; at least the family of JSON formats is unambiguously parseable and the correct string values will pop out. But it's less of a full solution than I'd like.
(It wouldn't even have to be that much of an authority necessarily, but certainly more than "The HN user named jerf declares it to be thus." Though I suppose if I registered "csvjson.org" or some obvious variant and put up a suitably professional-looking page that might just do the trick. I know of a few other "standards" that don't seem to be much more than that. Technically, even JSON itself wasn't much more than that for a lot of its run, though it is an IETF standard now.)
However it is not particularly readable/diff-able if this is part of your use case.
Yes, this is a major pain. It can be avoided by using Tab separated value (TSV) files, which don't use escaping. But then you can't store Tabs or carriage returns in your data. Also there is no way to store metadata in TSV.
JSON is far from ideal for storing 2D data tables as it is a tree. This means it is much more verbose than it needs to be. The same is also true for XML.
Of course this is not always applicable, since you sometimes don't control the format you get your data.
In practice, I used TSVs a lot, as tabs do not usually occur in most data. Alternatively, you could use pipes (|) or control characters as field or row separators.
If you are going to displace a standard, it has to be significantly better than the old.
I realize that this is a specific use case here, but I was on Cognos for years, and then when they started shifting over to Tableau, it wasn't any better.. csv, MS formats, proprietary formats, etc.
Avro is not a lossless alternate encoding of what's canonically a JSON document. Yes Avro supports a JSON encoding, but it's not canonical.
In general though, you're point about Avro being able to be represented as text is valid, and applies to practically any binary formThere's also Avro, which fails your point #1 (it's a binary format) but that binary format is a lossless alternate encoding of what's canonically a JSON document, and there are both simple CLI tools / and small, free, high-quality libraries that can map back and forth between the "raw" JSON document and the Avro-encoded file. At any time, you can decode the Avro-encoded file to text, to examine/modify it in a text editor.at, which is why the whole "but it needs to be a text format" argument is garbage.
> The data-warehouse ecosystem already standardized on Avro as its data interchange format. And spreadsheets are just tiny data warehouses. So why not? ;)
I wish that the data-warehouse ecosystem standardized on anything. ;-)
That said, there are plenty of good reasons why a data-warehouse standard would not be advisable for spreadsheets.
I'm still looking for somwthing that can do it.
(Of course, Spark doesn't support timestampz which is probably why the formats don't.)
For what it's worth, I totally agree something like compressed json lines is a better data exchange format, but part of why csv remains as universal and supported as it is is that so much existing data storage applications export to either csv or excel and that's about it. So any ETL system that can't strictly control the source of its input data has no choice but to support csv.
There's a subset of CSV that forbids escapes that is super fast to parse. All fast CSV parsers I'm aware of take advantage of this subset. I try to never ever publish a CSV that has quotes, and always aim for a more restrictive grammar that is cleaner, better thought out data.
*shudder
I find this really hard to believe given it's a simple enough syntax. And parsing is usually not the limiting factor, usually fast enough to not be noticed alongside interpreting or loading the source data. Every (much more sophisticated) compiler I can think of uses a linear parser based on this assumption.
I'll admit though that "import JSON" and then being able to essentially convert the entire file into a dictionary is nice if the data has more structure to it.
If I import JSON data I have no idea what shape the result will be in, and it requires a separate standard to let me know about columns and rows and validation can get complicated.
CSV is fine .. usually
#include <json>
std::json myjson("{\"someArray\": [1,2,3,4,{\"a\": \"b\"}]}");
std::cout << (std::string)myjson["someArray"][4]["a"];
and the result is we have 50 different rogue JSON libraries instead of an STL solution. Until the STL folks wake up, boost::split can deal with the CSV....until there is a newline inside a field.
The moronic quoting mechanism of CSV is one half of the problem; people like you, who try to parse it by "just splitting strings" is the other half. The third half is that it's locale dependent and after 30+ years, people still don't use Unicode.
"1) A truly open format is available" : sqlite is open-source, MIT-licensed, and well specified (even though I am usually not so happy with its weak typing approach, yet in this case this precisely enables a 100% correspondance between CSV and sqlite since CSV has also no typing at all...)
"2) Applications have a speed increase from using csvs" : I think it should be obvious to everyone that this is the case...
"3) The new format should not be grossly larger than the one it is replacing" : this is also the case
For single tables a database is probably overkill, but it's nice to have around when you need something reasonably powerful without being overly complex or hard to get started with.
The author appears to be a consultant selling data prep/transformation services. As long as the market is using CSVs, he’s forced to use CSVs, at least as end-of-pipeline inputs and outputs.
Of course, “people optimize their workflows for something other than making my job easy” is a common, but also rarely persuasive in motivating action from others with different jobs, complaint.
Ascii #31 instead of commas, Ascii #30 instead of newlines. Now those characters can go into your values.
If that's no good, zstd-compressed proto.
The advantage of CSV is that it's as accurate as your plain text representation of your data can be. Since binary data can be represented by character data, that's 100% accurate. As soon as you introduce a storage format that has made assumptions about the type of data being stored, you've lost flexibility.
SQLite is not intended for data serialization. It's intended for data storage to be read back by essentially the same application in the same environment.
4) Changes to the replacement format should be human-readable in a diff
It's easy to write a diff filter for e.g. xlsx, it's not possible to make CSV any good.
4) can easily interact with Excel.
A lot of things 'around here' run off spreadsheets, because relational databases weren't invented here and no-one ever needs more than a few million rows apparently.
I seem to spend half my job on the current project exporting stuff to CSV, running it through my code, then opening the resulting CSV in excel and formatting it a bit and saving as .xlsx again.
Still, at least I don't have to use Visual Basic that way.
CSV should be better standardized, but ... whatever, what should be done to "fix" CSV is to advertise the proper use of the libraries and the nontrivial aspects of a superficially trivial format.
A format that is trivially useful in 99% of cases is far better than many other "worse is better" things in computing.
They really don't. In fact I'd go further and confidently state that they really can't, because tons of mis-parsed CSVs are heuristic judgement values, and those tools don't really have the ability to make those calls.
I've never seen a "mature CSV library for most major language" which'd guess encoding, separators, quoting/escaping, jaggedness, … to say nothing of being able to fix issues like mojibake.
2. Yes, mixing formatting with data slows down data processing, don't do it.
3. Excel is not the replacement for CSV, and CSV is not a compact format. I mean, maybe if you are used to XML it is, but otherwise, just no.
Yes, there is truly a problem.
That's only true if you're trying to send all your data in a single, monolithic CSV.
If you're sending multiple CSVs, you're capable of representing data as well as a relational data store. Which is to say, you're representing your data using a system of data normalization specifically designed to minimalize data duplication. A single CSV represents a single table, and in most cases with intelligent delimiter selection you can represent an entire data set with no more than one character spent between fields or records.
Yes, you do have situations where you're storing losing data density due to using plain text strings, but that's not a limitation particularly unique to CSV for data serialization formats. Additionally, it is a problem that can largely be mitigated by simple text compression. Furthermore, once you switch to a non-text representation, you're limiting yourself to whatever that data representation is. It's easy to represent an arbitrary precision decimal number in plain text. It's hard to find a binary representation that universally represents the same data regardless of the system on the other end. Again, that's not a problem unique to CSVs.
If you're working with an API, object by object, then JSON is certainly going to be better, yes, because you can use the application's object representation. If you're working with bulk data of many disparate, unrelated, complex objects, however, or where you're transferring and entire system, you're not going to do much better than CSV.
There are few reasons to continue using csv in this day and age.
"The only true successor of CSV should be forward/backward compatible with any existing CSV variant"
If we manage to write a spec that meet this criteria we'll have a powerful standard with easy adoption.
1. a fool's errand, CSV "variants" are not compatible with one another and regularly contradict one another (one needs not look any further than Excel's localised CSVs)
2. resulting in getting CSV anyway, which is a lot of efforts to do nothing
So, a binary format consisting of: (1) a text data segment (2) and end of file character (3) a second text data segment with structured metadata describing the layout of the first text data segment, which can be as simple (in terms of meaning; the structure should be more constrained for machine readability) as “It’s some kind of CSV, yo!” to a description of specific CSV variations (headers? column data types? escaping mechanisms? etc.) or even specify that the main body is JSON, YAML, XML, etc. (which would probably often be detectable by inspection, but this removes any ambiguity).
There is no reason to try to be "backwards-compatible" with existing CSV files - we don't have a single definition of correctness to use to check that the compatibility is correct. Every attempt to be parse existing data would result in unexpected results or even data loss for some CSVs in the wild, because there is no way to reconcile all the different expectations and specifications that people have for their own CSV data.
But some people are. There are entire industries built around the exchange of CSV files, and the producers and consumers don't necessarily talk to each other.
Sqlite?
> Applications have a speed increase from using csvs.
Sqlite?
> The new format should not be grossly larger than the one it is replacing
Sqlite it is.
--------
Oh, you mean something that Excel can open? Oh yeah, I guess CSV then. But lets not pretend #1 (openness), #2 (speed), and #3 (size) are the issues.
CSV is a format more for humans and less for machines, but that is the use case: a format that is good enough to be compiled by humans and read by machines. At the moment there aren't many alternatives.
Until someone gets excel to ingest and produce something in a better format, we're pretty much stuck.
After working with so many retailers and online sales channels, things that are considered "legacy" or "outdated" by the HN crowd doesn't seem like it will go away unless both sides make a change. There are numerous articles posted on HN about how "FTP is dead" or no one uses it anymore, when it's far from the case.
Even Amazon's marketplace and vendor files are still using SFTP and EDI files. They've recently made changes, but it's been slow and hasn't had widespread adoption.
There's also the universality and "simplicity" CSV provides to the non-computer literate, and convincing them to make a change to a new standard provides itself some non-technical challenges. CSV is a bad standard, but it's the best one given what it does and its flexibility.
Maybe a CSV killer would be a human readable columnar based file format.
Nevertheless, the article basically discusses the issues encountered with the manual "editability" of CSV files, not so much with its performance. It also mentions parquet or arrow and concedes that they require specialized format to read/write. If we are looking to that, then there are a lot of options such as sqlite format, BerkleDB (used by some cryptocurrency projects) among plenty of others.
Sometimes you need to look up a single code in a 50 MB file. And sometimes you need a quick check to see if one line or a million lines changes.
It's "exception not the rule" type stuff... but it sure comes in handy to be able to check this stuff quick with basic text tools than have to run it through some binary parser. Same as JSON. But unlike protobufs for example.
The best that can be said for its simplicity is that it's easy to write code that can dump data out in CSV format (and to a lesser extent, it's easy to parse it, though watch out for those variants). This is not a really strong argument, most everyone is going to use a library for serialization, there's no reason to write your own unless it's for learning.
It is a lowest common denominator format. That type of thing is incredibly hard to kill unless you can replace it with something that is simpler. Good luck with that.
Do you really think there were no other "portable" formats to exchange data and SQLite is the first one?
The reason CSV is so popular has nothing to do with technical superiority which it obviously lacks.
The reason CSV is so popular is because it is dead simple and extremely easy to integrate within practically any conceivable workflow.
And because it is text format which you can trust, even if you have no other tools, you can inspect and edit in a text editor which is exactly the reason why other formats like INI, XML, JSON or YAML are so popular.
How would you:
a) Run shell command on a locked (can't install anything) PROD server to find the files and write a report of the file location, date of change and file size so that it can be easily processed by something else? Return the data in stream on the same SSH connection?
b) Quickly add a line to an existing startup script to get a report of when the script is starting and stopping so you can import data to excel. You just need it for couple of days, then you delete it.
c) Produce a report you can send to your coworker knowing they don't know how to program? And you don't want spending time explaining to them how to make use of your superior file format?
d) You talk to another team. You need to propose a format and specify a report with as little effort as possible. You know they are not advanced technically so you keep things simple?
e) You are solving an emergency and have to cross-check some data with another team. You need to produce a file that they will be able to read and process. You are stressed for time so you don't want to start a committee right now to find out shared technology. What would be safe choice without asking the other team for options?
f) A manager asked you for some data. You want to earn quick kudos. How will you send the data to get your kudos rather than irritating him?
These are all real life situations where CSV is being used.
As much as I understand technical arguments for SQLite, not every person is a developer. And all that technical superiority is worth nothing when they struggle to make use of the file you sent them.
I would say, the news of CSV's death are greatly exaggerated.
> This column obviously contains dates, but which dates? Most of the world
It's time to retire local formats and always write YYYY-MM-DD (which is both the international and the Swedish standard, and the most convenient for parsing and sorting).
> A third major piece of metadata missing from CSVs is information about the file’s character encoding.
It's bloody the time to retire all the character encodings and always use UTF-8 (and update all the standards like ISO, RFC etc to require UTF-8). The last time I checked common e-mail clients like Thunderbird and Outlook created new e-mails in ANSI/ISO codepages by default (although they are perfectly capable of using UTF-8) - this infuriated me.
> If not CSV, then what? ... HDF5
Indeed! Since the moment I discovered HDF5 I wonder why is it not the default format for spreadsheet apps. It could just store the data, the metadata, the formulae, the formatting details and the file-level properties in different dimensions of its structure to make a perfect spreadsheet file. Nevertheless spreadsheet apps like LibreOffice Calc and MS Excel don't even let you import from HDF5.
> An enormous amount of structured information is stored in SQLite databases
Yet still very underused. It ought to be more popular. In fact every time I get CSV data I import it to SQLite to store and process but most of the people (non-developers) have never heard of it. IMHO it also begs to be supported (for easy import and export at least) by the spreadsheet apps. A caveat here is it still uses strings to store dates so the dates still can be in any imaginable format. Fortunately most of the developers use a variation of ISO 8601 conventionally.
And by the way, almost every application-specific file format could be replaced by SQLite or HDF5 for good. IMHO the only cases where custom format make good sense are streaming and extremely resource-limited embedded solutions.
Both also require colons to separate hours and minutes and this makes it impossible to use in file names if you want to support accessing them from Windows.
I personally use the actual ISO 8601 (with the "T") wherever I can, simple YYYY-MM-DD-HH-mm-SS-ffffff where I need to support saving to the file system (but this is slightly harder for a human to read) and mostly RFC 3339 (with a space instead of the "T") wherever I need to display or to interop with tools written by other people. As for SQLite - I usually store every field (years, months,... seconds etc) in a separate integer column and create a view which adds an automatically generated RFC 3339 date/time column for simpler querying.
And by the way, many (if not an overwhelming majority) of the non-programmers don't even understand what does "just text files" actually mean, how do text files differ in nature from DOC files and how are CSV files different from XLS files. They can only use CSV because Excel and LibreOffice support it OOTB and consider CSV just a weird XLS cousin needed for import/export purposes.
Because sorting. You can just sort a collection of dates stored as YYYY-MM-DD strings alphabetically and the result will always be in accordance with the actual time line.
Usually MMM refers to the 3-letter shorthand of the month, e.g. "APR" or "OCT", but I guess that's not what you meant because it couldn't be used internationally.
(In theory people could write yyyy-dd-mm, but I’ve never seen anyone actually do that).
I like elegance as much as anyone. And I think it's a good proxy for other important qualities. But don't prioritize it above building something that actually does the job. Be an engineer.
Not to mention the mess that is exchanging documents between different locales. It's all sunshine and roses until you get your CSVs from an office in a different country (which happens a lot in Europe).
CSV gets the job done until it doesn't.
Unexpected behavior is a potential problem with any tool or format, certainly no less so with the kinds of solutions the article is proposing.
Then you take it out of use where it doesn't get the job done.
> To usurp CSV as the path of least resistance, its successor must have equivalent or superior support for both editing and viewing
Since the author has already ruled (on spurious grounds) human readability and thus the usability of ubiquitous text editors as incompatible with requirements for a successor format, this is simply impossible.
> If Microsoft and Salesforce were somehow convinced to move away from CSV support in Excel and Tableau, a large portion of business users would move to a successor format as a matter of course.
Yeah, sorry, you’ve put the cart before the horse. Neither of those firms are going to do that until the vast majority of users have already migrated off of CSV-based workflows, it would be an insane, user hostile move that would be a bigger threat to their established dominance in their respective markets than anything any potential competitor is likely to do in the foreseeable future.
If you think you can do better than CSV, let's see your proposal. Hint: you probably can't, and if you could, you probably couldn't get Excel to export it, so you still probably can't.
"The status quo is bad, more recent popular formats aren't good enough either, but I don't actually have a specific proposal that's better than all of the above" is a lot faster to read than that article, and says the same thing.
I've got one! It's basically the same as regular CSV, but everything is UTF-8, the columns and lines are delineated by dedicated UTF-8 "delineator" codepoints (if they aren't defined in the spec, find reasonable surrogates and use them), and therefore nothing ever needs to be escaped.
More human readable than regular CSV, less prone to error and just as easy to write to in a for loop (easier, in fact, as there are no escapes).
Depending on how Excel handles delineators, it should be able to import it too.
TSV allows you to do stuff on a single machine and GNU parallel that people would normally create a Hadoop cluster or 128GB database for.
For those not aware, TSV and CSV differ by more than just the delimiting character. TSV has a dead-simple specification: https://www.iana.org/assignments/media-types/text/tab-separa.... CSV does not have a standard spec and implementations differ quite a bit, but often in subtle ways.
Here's the RFC for CSV - https://datatracker.ietf.org/doc/html/rfc4180
The author completely misses the point of what CSV files are useful for.
They are useful when both ends of the communication understand the context. They know what the data types are, they know what the character encoding is etcetera.
This is a very common situation. CSV files are easy to process, easy to generate, and can be read by a human without too much bother (they are not "for people" as the author so irritatingly asserts). I have written so many CSV (and other delimiters, not just ',') generators and processors I lost count decades ago. It is so easy.
And what does "retire CSV" really mean? Just stop generating them then! The consumers of your data will probably insist you go back to them as writing all that code just to satisfy your fetish with metadata is not useful
Now I am whinging.....
1. It is a very common use-case to load CSV data generated by an opaque system which did not specify the format exactly to you.
2. CSV files are _difficult_ to process I: Fields are of variable width, which is not specified elsewhere. So you can't start processing different records in parallel, since you have to process previous lines to figure out where they end. You can't even go by some clear line termination pattern, since the quoting and escaping rules, and the inconsistencies the author mentions, make it so that you may not be able to determine with certainty that what you're reading is an actual end-of-record or end-of-field. Maybe it's all just a dream^H^H^H^H^H quoted string?
3. CSV files are _difficult_ to process II: For the same reasons as the above, it is non-trivial and time consuming to recover from CSV generation mistakes, corruptions during transmission, or missing parts of the data.
However - if you get some guarantees and abot field width and about the non-use of field and record separators in quoted strings, then your life becomes much easier.
CSVs also are great because you can parse them one row at a time. This makes for a very scale-able and memory-efficient way of processing very large files containing millions of rows.
Let there be no mistake: Everyone reading this today will retire long before CSVs retire. And that's just fine by me.
Even RFC4180-compliant CSVs can be incredibly memory-inefficient to parse. If you encounter a quoted field, you must continue to the next unescaped quote to discover how large the field is, since all newlines you encounter are part of the field contents. Field sizes (and therefore row sizes) are unbounded, and much harder to determine than simply looking for newlines - if you were to naively treat CSV as a "memory-efficient" format to parse, you would create a parser that would be easy to blow up with a trivial large file.
2) I ask for TSV, whenever convenient. It's been more reliable, and I don't have a comprehensive why, but I think it's slightly more resilient to writer/reader inconsistencies, for me. It may be that there's just less need for escaping and quoting, so you might dodge a smart quotes debacle when asking for a one-off from an Excel user, for example.
3) Despite the issues raised, the notion we'd retire it makes me hug it tight, because for the majority of my requirements, it hits a sweet spot. I still reserve the right to raise my fist in frustration when someone does: ",".join(mylist).
CSV is a mish-mash of different, complicated, and under-specified standards and/or implementations.
Instead of finding a replacement for CSV it might be easier to standardize it and enhance it. Excel's version of CSV is the de-facto standard. If you want a written-down spec that is available too [1]. To this we need to add enhancements such as a way to specify metadata (i.e., data type of each field). No need to find an alternative to CSV!
If you're working someplace that uses Excel, then maybe. Otherwise, it isn't.
Anyway, if more tools, including Excel, properly implemented RFC4180, we'd at least have something.
Another option is some sort of meta-data file alongside a large CSV describing its format somehow - which, if missing, you have to fall back on RFC4180 or your parser's "try to figure out the quirks" mode.
"I Cannot Fathom A Use Case For CSV So I Would Like To Ban Everyone From Using It For Any Reason"
As soon as data has any form of structure to it (and most data does). CSV complicates everything. Even for unstructured data, the problem of escape characters often shows it's ugly head. The moment your data contains a comma, tab, or space, you run into a nasty mess that, in the best case makes your system fail, and in the worst case silently adds corrupt data into the system.
Neither JSON nor XML suffer from that problem and both can easily be used in any scenario you'd use CSV. The only argument against either format is they are a bit more bulky than CSV.
Just yesterday I was trying to add some metadata to a list of images that I handed off to a designer and I just piped a single "find . -fprintf" command into a file.
To "retire" something so simple and utilitarian makes no sense. There's plenty of formats that deal with all the issues described in this. It's like saying we need to retire plain text files because it's confusing to know how they should be displayed.
Writing code that can accept and parse arbitrary CSV files is a whole different thing. If I had to do that, I'd be yelling too.
Similar story can be said for JSON, markdown, etc, where standard is inadequate, non-existent, or in mutual competition.
It was weakest when it implied there is no real standard. There is, and it's robust for representing data, even data that includes any combination of commas and double-quotes.
The algorithm for creating well-formed CSV from data is straightforward and almost trivial: if the datum has no comma in it, leave it alone. It's good to go. If it has even one comma, wrap the datum in double quotes; and if it also contains double quotes, then double them.
Not complicated and covers every edge case. CSV is going nowhere. 1000 years from now, computers will still be using CSV.
The answer to his objections could be to extend the format to include metadata. Perhaps a second row that holds type data.
fruit,price,expiration
string,$0.00,MM/DD/YYYY
apple,$0.45,01/24/2022
durian,$1.34,08/20/2021
etcIt is very easy to overlook edge cases in CSV.
That was not a criticism from the original article, and isn't even true.
If you have newlines in your original data, and you "overlooked" this "edge case", then neither JSON nor YAML nor any other format will save you. The same fix applies to them all.
This is really a very poor criticism.
CSV is ubiqituous because it's not trying to solve difficult problems. The second you try to solve the difficult problems you necessarily fragment your audience.
So, question to the greybeards: when this format was coming about, why didn't we use one of the dedicated Data Structure separator control codes (e.g. File/Group/Record/Unit separator) that were part of the ASCII standard?
https://en.wikipedia.org/wiki/C0_and_C1_control_codes#Basic_...
It seems like it would've saved us all several decades of headache.
But commas can be seen, edited, and typed with ease in any text editor.
As per comment above: your CEO, SWE, or secretary can all use or contribute to a csv file. And using an easily recognizable and typeable separator has proven to be worth the downsides.
Consider also that IBM EBDCIC and other non-ASCII character sets were (and are) in common use, so C0/C1 may not have made sense 'back in the day'
Because you can't see them. CSV is, at its core, a text format. Using FS/GS/RS/US would effectively make it a binary format.
Also, what happens if one of those bytes appears in data?
The simple fact is you will never be able to come up with a text-based format that can handle all possible values without escape sequences, quoting, or length-encoding. And really, that's not that big a deal. It just means you have to write a simple state machine instead of using regex or your language's equivalent of split().
- Repairability by mere mortals. If you are competent enough to use Excel and are given bad data, you are competent enough to fix it. (Whether the dataset is too big for mere mortals to find the problem is a different issue.)
- Trivial serialization. It is super-easy to dump things to CSV. (This is probably also why there are so many annoying variants.)
- Usable by other tools. You don't need special libraries to read it, so pipelines involving the usual suspects is possible. Importantly, there are a ton of different ways to do this sort of thing, so non-experts can frequently find something that works for them, even if it looks wonky to programmers.
All that said, I hate dealing with them, too.
Anyone that has dealt with more than 1 CSV is aware of many of the aspects of their horrible nature.
The more interesting reason is why they're so damn successful.
1: network effect - not supporting CSV in a product is practically silly. everyone can do it why can't you?
2: ease of producing / consuming (not saying that you do it correctly in all cases :))
3: data is transparent (or feels transparent)
4: accessible - if you can write language X you can parse or generate reasonable CSV in a mater of minutes. No knowledge of any libraries or tech required
That said it's horrible - but it will always be with us.
Microsoft -> The name of the format is 'comma separated values' not 'semicolon separated values'!
* Parquet stores schema in the metadata, so schema inference isn't required (schema inference is expensive for big datasets)
* Parquet files are columnar so individual columns can be grabbed for analyses (Spark does this automatically). This is a huge performance improvement.
* Row groups contain min/max info for each column, which allows for predicate pushdown filtering data skipping. Highly recommend playing with PyArrow + Parquet to see the metadata that's available.
* Columnar file formats are easier to compress. Binary files are way smaller than text files like CSV even without compression.
I wrote a blog post that shows how a Parquet query that leverages column pruning and predicate pushdown filtering can be 85x faster than an equivalent CSV query: https://coiled.io/parquet-column-pruning-predicate-pushdown/
CSVs are great for small datasets when human readability is an important feature. They're also great when you need to mutate the file. Parquet files are immutable.
It's easy to convert CSVs => Parquet/Delta/Avro with Pandas/Dask/Spark.
The world is already shifting to different file formats. We just need to show folks how Parquet is easy to use and will greatly increase their analysis speeds & they'll be happy to start using it.
Small nit: the article implies Apache Arrow is a file format. It's a memory format.
>Most of us don’t use punch-cards anymore, but that ease of authorship remains one of CSV’s most attractive qualities. CSVs can be read and written by just about anything, even if that thing doesn’t know about the CSV format itself.
Yes, we keep CSVs. If you care about metadata and incredibly strict spec compliance, then yes: avro, parquet, json, whatever. But most CSV usage is small data, where the ease of usage, creation, and evaluation wins.
One of the problems with CSVs he cites is a great reason why I like CSVs:
>CSVs often begin life as exported spreadsheets or table dumps from legacy databases, and often end life as a pile of undifferentiated files in a data lake, awaiting the restoration of their precious metadata so they can be organized and mined for insights.
A benefit of a CSV is a skilled or unskilled operator can evaluate these piles of aging data. Parquet? Even SQLite? Not so much.
For small-to-medium sized datasets, CSV is great and accessible to a wider user base. If you're relying on CSV to preserve structure and meta, then meh.
And the important think to remember is that you can not and will not make them.
It's an exercise in hubris because while CSV has one glaring fault, it is far fewer than other competing methods of the simplest possible exchange of textual data.
No one can decide to retire CSV because it ubiquitous. Of the three issues OP compains about: line delimiter, field delimiter, and header, only one is really a problem for anyone that has used them for any period of time.
1. Every file format suffers from Windows/NonWindows CRLF issues.
2. Metadata has been an issue since forever, and there are plenty of painful formats that support it (looking at you XML), complete with ginormous parsers and even more issues.
3. Escaping, as OP points out, was pretty clearly defined in RFC 4180.
So yes, it has a wart: escaping.
Learn what your system expects, and modify accordingly. Because using a CSV will be much faster and simpler than any other format you can try, which is why it has been so pervasive for longer than most of HN has been alive.
>values stored in the files are typed. >most importantly, these formats trade human readability and writability for precision
Those two properties are advantageous only when CSV is used in cases meant to be just parsed and as a data transfer format. But in reality, CSV files are being used in many different contexts. For instance, data scientists love to leverage and chain Unix tools to create a subset of data to test models. Also, it is often used as an export format to validate outputs quickly.
I think the problem is that CSV is often the subject of abuse. In the same way that a spreadsheet is a subpar database, I don't believe we are nearly close to the time when we will retire Excel.
People want to read and interpret arbitrary CSV ... well you cannot read and interpret any format that is arbitrary.
*As a consultant, I’ve written more than one internal system that attempts to reconstruct the metadata of a CSV of unknown provenance using a combination of heuristics and brute force. In short, CSV is a scourge that has followed me throughout my career.*
Ideally this should not happen because you should talk with party that you agree on common format. Someone that would explain what each field means and what should it contain, or at least some documentation for the file, not that it just is a CSV. But of course it always is more complicated than that.
Garbage in - Garbage out, even in other formats you still can have the same problem.
JSON - and also CSV - are ubiquitous because they are actually human readable and usually quick to process. JSON is strictly defined it just doesn't over comments or structures, like XML. CSV is loosely defined but better than another competing standard ;)
A good data interchange format must be human readable, JSON and CSV are. Proprietary stuff not. And if you feel the need for speed? Seriously? Okay, then think about a binary format.
I prefer also the .conf format (so called INI) over other complex stuff for application settings.
The conclusion is that there are many binary formats which are more suitable, except that none are used widely enough, and the post even goes on to not make a recommendation. Finally it says that we have to accept this sad state:
"Ultimately, even if it could muster the collective will, the industry doesn’t have to unify around a single successor format. We only need to move to some collection of formats that are built with machine-readability and clarity as first-order design principles. But how do we get there?"
Whatever is used, I like that tab-delimited (or even csv) is a human-readable and always-machine-readable long-term data format in the same way as ASCII (or Markdown etc) is for text. A hundred years from now, assuming the storage medium is still usable, the content should be easily recoverable. That may not be the case with spreadsheet files or other structured data formats.
https://www.w3.org/TR/2015/REC-tabular-data-model-20151217/#...
... at the time RFC-4180 came out, it didn't even accurately describe how to read csv produced by Excel - which was already inconsistent in line endings and character set between office versions and platforms. The w3c spec at least tried to offer a model which would parse junk csv if you could guess the metadata (by eg scanning for the BOM, \" vs ""[^,], and so on)
When I worked on this stuff early 2010s, if you wanted to produce a non-ascii csv _download_ that could be opened by all office/openoffice variants you were out of luck. UTF-16LE-with-BOM, as I recall, would work in _most_ office variants but not consistently even across minor version changes in Office for OSX - so it was just a roll of the dice. We offered multiple download formats which _could_ handle this but csv was required by some customers.
Anyone saying csv is easy never worked with it in an international context.
1. You can read them without any software
2. Streamable row by row
3. Compress well
To be honest most of the points in this article could be addressed by standardizing a method of defining a schema for the CSVs. It could even be backward compatible by appending the definition as metadata on the header column or as a separate file.
One thing that would be good to have is a standardized method of indexing CSVs so that random access is possible too, though that would be more involved.
name::string,number of legs::int,height in meters::float,date of birth::date(MM/DD/YYYY),email adress::email,website::url
joe,2,1.76,12/12/1999,joe@joe.com,https://www.joe.com
bob,1,1.84,12/12/1944,bob@vietnam.com,null1. Encoding is UTF-8.
2. Header line with column names is mandatory. If there are no column names, the first line must be blank.
3. Each record is a line. A line is defined according to the operating platform's text file format.
4. In the light of (3) CSV does not dictate line endings and does not address the conversion issue of text files from one platform being transferred to another platform for processing. This consideration is a general text issue, off-topic to CSV.
5. CSV consists of items separated by quotes. An item may be:
5. a) a JSON number, surrounded by optional whitespace. Such an object may be specially recognized as a number by the CSV-processing implementation.
5. b) the symbol true, false or nil, optionally surrounded by whitespace. These symbols may have a distinct meaning from "true", "false" or "nil" strings in the CSV-processing implementation.
5. c) a JSON string literal
5. d) any sequence of printable and whitespace characters, other than comma or quote, including empty sequence.
6. In the case of (5) (d), the sequence is interpreted as a character string, after the removal of leading and trailing whitespace. (To preserve leading and trailing whitespace in a datum, a JSON literal must be used.)
7. In (5), whitespace refers to the ASCII space (32) and TAB (9) character. If there are any other control characters, the processing behavior is implementation-defined. Arbitrary character codes may be encoded using JSON literals.
The real problem with CSV is the lack of vision of data formats in general. Inside your programming language you need to say that you want to read a CSV? You need to import a different library and change both the parse code and your consumption code to instead do JSON? You need to change everything to use futures in order to stream the results? Are you out of your damn mind?
So now that disk space is so cheap it would make sense for any file format to just begin with a single few-kilobytes line that defines the parser for the coming file. Could be sandboxed, WASM-y or something, could be made printable of course... Sure, it might not be possible to get the full SQLite library in there but you could at least make it free to switch between JSON and BSON and CSV without having to recompile the software or force someone to design modular file input systems. Somehow the only flexible containers that hold different data structures are video container formats that do not care about what codec you used. They "get it"— can the rest of us?
2. I would feel safer storing longer-term data in CSV than in a binary format w/ a complicated spec. Having to make sense of compressed, columnar padded data sounds worse than parsing CSV.
3. Although it's easy to point at corner cases, I don't remember the last time I couldn't figure out how to parse a file because of inconsistent quoting or exotic char encoding – and I've spent a good amount of the past 13 years exchanging CSVs and TSVs w/ 3rd parties full of horrible legacy. Asking those 3rd parties to send me an Avro/Parquet/whatever file would've made the project fail or take 10x longer.
There's a reason why CSV stuck around so long, and the alternatives make different trade-offs but miss the pros.
The trick is that if you compile data into a SQLite file and then deploy the Datasette web application with a bundled copy of that database file, users who need CSV can still have it: every Datasette table and query offers a CSV export.
But... you can also get the data out as JSON. Or you can reshape it (rename columns etc) with a SQL query and export the new shape.
Or you can install plugins like https://datasette.io/plugins/datasette-yaml or https://datasette.io/plugins/datasette-ics or https://datasette.io/plugins/datasette-atom to enable other formats.
> then deploy the Datasette web application
Is a huge hurdle for non-technical folks holding on to their CSV workflows.
I think there is no barrier low enough that CSV cannot limbo beneath it.
U+FEFF"aaa","b CRLF
bb","cc"c" CRLF
zzz,yyy,xxx
Just an example of the wonderful world of CSV \(.)/ which I've used quite a lot, because it enables quick and easy (dirty) data dumps/ exchanges. But using it as a data exchange format between multiple parties often leads to problems, such as above.Suppose someone wants to do some data processing client side in the browser with an input file. How would you do this for SQLite, even a SQLite file with a known schema such as would be the case most of the time for CSV file uploads? There is sql.js which compiles SQLite to JS or WASM, but even the WASM version weighs in around 400kb.
If you use the right characters built specifically for this purpose, then the problems go away.
For a start, think about it from a game theory perspective. If I keep CSV import/export in my Easy Data Transform product, while all my competitors with data transformation software (Alteryx, Knime etc) remove it from theirs, it gives me a big sales advantage. What advantage do I get for removing already working CSV import/export code? Nothing.
It is self describing, rigorously defined for the machine, highly flexible, highly extensible, human readable, text-based and an open format, and it is only mentioned as a storage format in the post, not even nearly doing it justice!
XML is complex, but there's already very established libraries for handling it, so that shouldn't be an issue, and that's the only drawback XML has.
Hell, even alternatives like JSON/YAML + JSONSchema exist, still providing incredibly rigorous validation, but staying human readable and ubiquitous.
P.S. I think a bigger failure here is trying to generalize the concept of spreadsheets, rather than using a common encoding (XML) for domain-specific data formatting.
Parsing times are often horrible.
There’s no standard for tabular data. You invariably need some overly complicated XML map, because people can’t resist the temptation to over-engineer.
I can't remember a time that such a simple file format had such wild inconsistencies. Not to mention some CSV export functions just ignore them. Put a comma inside a CSV field? why not? Put a single double quote in a CSV field? sure! Insist that leading (or trailing) spaces in a CSV field are semantically important? OF COURSE!
If CSV was used consistently, it wouldn't be that bad. But it's apparently simplicity lulls developers into a false sense of security, which is part of what the original author seems to be saying.
And that's why they're not going away.
Trying to add typing to them is probably missing the point that >95% of humans that use them don't need or want to understand that.
When they give me literally anything else I get pissed off and push their task lower in my queue.
The format isn’t inherently flawed, though much like other less than perfect standards (I’m looking at you SMTP), the implementations frequently are.
And also, much like the author, I’ve been “professionally” dealing with data in all (most?) its forms for the length of my career (~25 years).
There's no such thing as schema-less - there's undefined schema
In fact, I will go even further and say I want something mergeable by source control.
- easy to parse
- easy to edit with a generic text editor
- easy to edit with a widely available GUI, like LibreOffice
- allow adding more data with only append operations
SQLite is wonderful, but it's definitely not a replacement for CSV
When I have lots of data in multiple related tables where I would benefit from defined data types, I reach for HDF5 or SQLite, but there are so many nice things to say about a simple CSV for lots of simply structured data.
Super simple to stream compress/decompress, and the ability to use great CLI tools like xsv, and being able to peak at the data by simply calling less or head... it's just hard to beat.
{ "$schema": "http://json-schema.org/draft-07/schema#", "type": "array", "items": {"type":"array","items": [ { "type": "number" }, { "type": "string" }, { "enum": ["Street", "Avenue", "Boulevard"] },{ "enum": ["NW", "NE", "SW", "SE"] } ]} }
[[3,"some", "Street", "NE"],
[4,"other", "Avenue", "SE"],
[5,"some", "Boulevard", "SW"]]How would any existing JSON parser handle 75GB of data in a single array of arrays?
Bad CSV is a PITA, but usually you can make reasonable sense of it. Merging bad/inconsistent/conflicting metadata tends to be an open-ended nightmare with no good resolution at the end.
The contract on what is transfered or how it's mapped in your csv is developer to developer and not in code.
This is a problem and also a benefit.
As long as both systems are designed to read and write the csv correctly(there in also the problem of verification), csv works flawlessly and retains its simplicity.
Sometimes -- frequently, even -- you have to go to some least-common-denominator format to move data around. CSV is often a candidate.
When I was younger I thought we'd eventually abandon "primitive" formats like this, but in my middle years I realize their extreme utility. CSV absolutely has a place, and absolutely does a job that no other format does as quickly or as easily.
As long as you are following RFC4180 it works.
I ended up exporting from Oracle DB using JSON and converting to CSV for one of our contractors to be able to import the data to MongoDB.
It is just CSV with a little bit information on the top. Great for define column types, so that you reader does not need to guess the column types.
See: https://csvy.org/
The only appropriate replacement for CSV would be a better defined CSV that gave hard requirements for things like field quoting
What we need is something akin to Strict Markdown. Something that qualifies every edge case to produce a strict CSV that can encompass human-readable metadata within the strict delimiters.
I'd instead rather see much, much more CSV as well as flattening data models to make data more suitable for tabular representation.
Handling newlines and quotes is really not that difficult.
Yet, everbody gets it wrong. Everybody. Thoroughly, fantastically, almost unimaginably wrong.
For example, in the Microsoft world: I come across CSVs in about 5 scenarios, all of them very common, most of them designed to interact: Excel, PowerShell's Export-CSV, SQL Server, Power BI, and the Azure Portal. None of these are obscure. None of these use CSV infrequently. Yet, they're all mutually incompatible!!!
That just blows my mind.
For example, SQL Server will output string ",NULL," to represent a null field instead of just a pair of commas (",,"), so every other tool will convert this to the string "NULL", which is not a null.
PowerShell helpfully outputs the "type" of the value it is outputting, like so:
PS C:\> dir | ConvertTo-Csv
#TYPE System.IO.DirectoryInfo
"PSPath","PSParentPath","PSChildName","PSDrive","PSProvider",...
Excel can't open such files! It is entirely unfathomable that the PowerShell team wrote this code and never once double-clicked the resulting CSV to see if it opens successfully in Excel or not. Absolutely mindblowing!It just goes on and on.
Different quoting rules. Random ability/inability to handle new lines. Different ways of handling quoted versus unquoted string values. Encoding is UTF-8 by default or not. Handling of quote characters within strings. Etc, etc...
Basically, within one vendor's ecosystem, flagship applications have at best a 50:50 chance of opening arbitrary CSV files.
Don't even get me started on the inherent ambiguities of the format, like interpreting dates, times, or high precision decimals, etc...
Oh, and before I forget: the SQL Server Integration Services team wrote lengthy articles on how their engine can process CSV files faster than the competition. Why is this a feature? Because CSV is a woefully inefficient format and processing it fast is an achievement.
Lastly: By default, SQL Server cannot export table data to a flat data file and round-trip it with full fidelity, in any format. Exchanging just a couple of tables between servers is... not fun.
So yes, a replacement format with an efficient, high-fidelity binary format is long overdue.
Languages and formats are not fundamentally about computers or efficiency, they're about people. Carry on.
It turns out people hate bloat and complexity more than they hate all the ills of CSV put together!
Yaml is my go to format when I need something human readable and somewhat editable ( very easy to ruin Yaml spacing)..
What's the real motivation behind an article like this. I don't imagine anyone has serious trouble with csvs, they tend to just work
They have open-source implementations in many languages, are much faster to load than csv, are natively compressed, are strongly typed and don't require parsing...
There are few reasons to continue using csv in this day and age.
Libra Office works in Windows.
Google Sheets does a great job too, no?
If it is stricter, it would have one type of field separator that is not commas since some locales use them as decimal places (I'm looking at you, France) but something like '|'. It would insist that dates were iso8601. It could define how fields can be escaped and quoted - although I would prefer if quoting was kept to a minimum. The format should also allow for comments i.e # so that people can comment their datasets inside the same file.
Alternatively or in addition, it could have some header lines:
1) A header that defines the encoding, separator, decimal separator, quote character, escape character, line ending character, date format ...
2) A header that defines each column's name
3) A header that defines each column's data type and formatting
4) A header that defines each column's unit like m/s or kg - ok, this is a bit of a stretch but it would be great to have.
or some variation of the above.
Fundamentally, this bsv format would still be csv and most programs would still be able to read it with the parsers that already exist or be quickly adapted to read it. It could still be easily edited by hand but the metadata would be present.
I suspect that this is just a pipe dream because people would find hundreds of ways to break it but toml took off and that didn't exist so long ago.
I don't think you can make a breakthrough through syntax alone. I think you've got to integrate some type of live semantic schema, something like Schema.org. If I used "bsv" and didn't just get a slightly better parsing experience but also got data augmentation for free, or suggested data transformations/visualizations, et cetera, then I could see a community building.
I think perhaps a GPT-N will be able to write it's own Schema.org thing, using all the world's content, and then a BSV format could come out of that.
we've resorted to csv files + readme + json files in git for version control.
It was outdated from the start because ASCII already has specific characters for file, group, record, and item separation. Using this would give broad compatibility and a wider feature set (eg. more than one table in a file) while retaining all the benefits of csv.
SQLite databases are in contrast a nightmare, and git merges often corrupt files.
Like
Id(int), Name, LastName, BornAt(date)
1, asd, dsa, 1992-07-12
I think the valuable insight here is that there needs to be a meme / movement that CSV is bad or deprecated. That's what's actually going to put the nails in its coffin, not private griping from developers when they get CSVs.
I'm all for it. Down with CSV :)
If I'm trying to make an exportable format for a data logger with an SD card running on an ARM microcontroller, it doesn't get much easier than CSV. Sure, I could save space by rolling my own binary format, but then I have to provide a PC application to read it (and realistically, the user is probably just going to want that application to dump to a CSV anyway!).
I agree that for many use cases there are much better alternatives, but one of the reasons CSV is so popular is because it's so simple. It shouldn't be used for multi-gigabyte datasets, but for many simple use cases it works great.
And the cycle continues.
Xkcd: https://m.xkcd.com/927/
There are plenty of formats available. CSV is useful. Not perfect.
CSV is the A-10 warthog of data formats
Unfortunately it is not widely adopted yet and the language support is yet to be improved.
no ♥
There are workarounds for sure; VB macros, unzip XSLX and parse the XML, write scripts to automate Excel-to-whatever using DCOM, or import into an intermediary service and re-export into the format de-jure, but that all takes time and costs money too, and often causes confusion when even spoken about. Asking for a CSV takes a couple of seconds and is easily understood by most people. Anyone experienced with importing / manipulating CSV data can deal with the variance in delimiters and escape sequences without major issues. It's a headache at times but easier than the alternative of alienating or confusing clients who are looking for a simple solution to whatever issue they have today.
On the other hand once the data is in the target system, if they're still asking for CSV exports I do probe to ask why, and try and figure out if there's a better way. Reporting is the usual reason, and plenty of the CRMs I work with have built-in reporting that can replicate and improve upon whatever spreadsheet they are using, and have APIs that can interface with cloud-based reporting services. But there's a lot of inertia against change in most small-to-medium organisations, not everyone is a data expert and you can't sell someone something they can't use or understand. Ultimately people win, and the solution ends up being a balance of hopefully incremental technological improvements that they can still integrate into their day-to-day. Not every organisation is able to undergo a full digital transformation with time, budgets, and skills at hand.
I agree with the sentiment but I'd hate to see a future where everyone's locked into proprietary ecosystems - not that that is what's advocated for in the post, but we have CSV because that's what the big platforms seem to allow, not because they don't know there are better options. There's no technical reason Excel couldn't export to WordPress or SuiteCRM. Take CSV away and it gets harder, not easier, to move between platforms.
The edge cases are a hassle but they don't become less of a hassle from a business perspective by switching to json or really any other format. We tried an experiment of using more json and eventually gave it up because it wasn't saving any time at a holistic level because the "data schema" conversations massively dominated the entirety of the development and testing time.
Obviously being able to jam out some json helped quite a bit initially, but then on the QA side we started to run in to problems with tooling not really being designed to handle massive json files. Basically, when something was invalid (such as the first time we encountered an invalid quote) it was not enjoyable to figure out where that was in a 15GB file.
That said, I fully concur with the general premise that CSV doesn't let you encode the solutions to these problems, which really really sucks. But, to solve that, we would output to a more columnar storage format like Parquet or something. This would let us fully encode and manage the data how we wanted while letting our clients continue working their processes.
What I would really like to see is a file format where the validity of the file could be established by only using the header. E.g. I could validate that all the values in a specific column were integers without having to read them all.
If/when I get a data source introduced into my workflow that differs from this variant I come up with a routine to normalize it, integrate that into my workflow and move on.
Every single one of their points are extremely subjective and very wrong. If you have an excel sheet full of equations and colorful cells and then to decide to export it as CSV and open it in Notepad, you really can't complain that that CSV is bad. I mean it's obvious that to each format a set of strengths and a set of weaknesses, and also obviously what could be a strength to someone is a weakness to someone else. The fact that CSV is so simple to parse (almost every modern language can very easily read/write a CSV) makes it a fantastic data transfer format for every single usecase I had (this doesn't mean that other formats are less important of course). Sure you'll lose your Google Sheet or Excel metadata, but this is NOT what CSV is for.
What I find fascinating though is that the author decided it's a good idea to make a blanket statement like "Time to retire?" just because they themselves have an issue with SOME use case. I mean the idea that having your personal needs not met to justify arguing that we should ALL stop using CSV is so very bizarre to me, like it's way beyond selfish.
Everyone knows it’s not capable of storing anything more than table data. It was never meant to be much more.
It doesn’t really store objects well. To write the values of a composite object would require following some format order. It would go past human readability and just be a bloated way of writing bytes with readable characters.