"The only true successor of CSV should be forward/backward compatible with any existing CSV variant"
If we manage to write a spec that meet this criteria we'll have a powerful standard with easy adoption.
"The only true successor of CSV should be forward/backward compatible with any existing CSV variant"
If we manage to write a spec that meet this criteria we'll have a powerful standard with easy adoption.
PS: I never said 100% forward/backward compatible with all variant at the same time and without any noticeable artifact. I meant compatible in a non blocking way.
Also UTF-8/ascii compatibility is unidirectional. A tool that understands ASCII is going to print nonsense when it encounters emoji or whatever in UTF-8. Even the idea that tools that only understand ASCII won't mangle UTF-8 is limited - sure dumb passthroughs are fine, but if it manipulates the text at all, then you're out of luck - what does it mean to uppercase the first byte of a flag emoji?
The assumption that most software uses is that the import file will be in the same variant of the format as what that tool exports. That seems to be more of a problem than anything else.
Sure, and they sometimes do that if they have to ingest CSVs whose origin they don't control (although not every system implementor cares enough to do it).
But that's still just a bunch of shitty faillible heuristics which would not be necessary if the format was not so horrible.
cat input1.csv input2.csv > output.csv
resulting in a single file containing multiple formats.
Also, what variant is this:
1,5,Here is a string "" that does stuff,2021-1-1
What is the value of the third column?Is this a CSV file without quoting? Then it's
Here is a string "" that does stuff
Or is it a CSV file with double quote escaping? Then it's Here is a string " that does stuff
This is fundamentally undecidable without knowledge of what the format it is.You can decide to just assume RFC compliant CSVs in the event of ambiguity, but then you absolutely will get bugs from users with non-RFC compliant CSV files.
So, yeah. Can't really be done without making too many assumptions that will break later.
Yes. And my software does that. But it is always going to be a guess which the user needs to be able to override.
1. a fool's errand, CSV "variants" are not compatible with one another and regularly contradict one another (one needs not look any further than Excel's localised CSVs)
2. resulting in getting CSV anyway, which is a lot of efforts to do nothing
So, a binary format consisting of: (1) a text data segment (2) and end of file character (3) a second text data segment with structured metadata describing the layout of the first text data segment, which can be as simple (in terms of meaning; the structure should be more constrained for machine readability) as “It’s some kind of CSV, yo!” to a description of specific CSV variations (headers? column data types? escaping mechanisms? etc.) or even specify that the main body is JSON, YAML, XML, etc. (which would probably often be detectable by inspection, but this removes any ambiguity).
Almost any current CSV parser, even the bad ones, tolerate a header line.
So it should be possible to define a compact and standardized syntax that is appended before the real header of the first cell (separator,encoding,decimal separator (often disregarded by most parsers but crucial outside USA),quote character,escape character,etc...). Following headers would just use special notation to inform on (data-type,length,comment).
Newest parsers would use theses clues, older ones would just append some manageable junk to headers.
Does this really count as compatible? You will get user bugs for this.
Yeah, that's why I chose the “thing that looks like a text file—including optionally CSV—but has additional metadata after the EOF mark” approach instead of stuffing additional metadata in the CSV; there's no way to guarantee that existing implementations will safely ignore any added metadata the main CSV body. (My mechanism has some risk in that there are probably CSV readers that treat the file as a binary byte stream and use the file size rather than a text stream that ends at EOF, but I expect its far fewer than will do the wrong thing with additional metadata before the first header.
There is no reason to try to be "backwards-compatible" with existing CSV files - we don't have a single definition of correctness to use to check that the compatibility is correct. Every attempt to be parse existing data would result in unexpected results or even data loss for some CSVs in the wild, because there is no way to reconcile all the different expectations and specifications that people have for their own CSV data.
Note the difference is that I am suggesting reducing the number of standards in-use by using only one already existing CSV format. :)