RFC 4180: Common Format and MIME Type for CSV Files (2005)
datatracker.ietf.org
datatracker.ietf.org
Not sure why this is trending now, but this RFC is being revised now: https://datatracker.ietf.org/doc/html/draft-shafranovich-rfc...
Suggestions and comments are welcome here: https://github.com/nightwatchcybersecurity/rfc4180-bis
Does any implementation support/generate comments that way? The most I've seen so far is an oversized multiline header.
The other part thinks this is quite cool and wishes your efforts well.
If you are trying to generate a file that plays nice with Excel, there is a way to force a specific delimiter with the "sep" pragma:
sep=|
a|b|c
1|2|3Simple is good.
If it's really simple even Microsoft will be able to implement it.
However, many terminal mode applications intercept those keys and expect them to be commands rather than data. Often there is some kind of quote key which you can press first. In readline and vim, the quote key is Ctrl+V. In emacs, it is Ctrl+Q instead.
GUI applications are more variable. But vim and emacs still support their respective quote keys in GUI mode.
And if you have to type this a lot, you can modify your vim/emacs/readline/etc configuration to remove the requirement for the quote key to be pressed first.
In this sense sticking to a relatively common separators is good, because they encourage you to do the right thing from the start.
(Even when some platforms have the facility to store the MIME type in an extended attribute, few applications will actually support retrieving and acting on that attribute.)
In the past I've used CSV to handle files that are several GB after being compressed and encrypted. Formats such as JSON would have added a lot more to the total file size.
If you store tables as list of objects, sure, but that's comparing the JSON way of doing something CSV can’t handle at all (lists of heterogenous objects) to CSV doing a much simpler thing.
A compact, dedicated JSON table representation (list of lists with the first list as header, or a single flat list with the header length — number of columns — as the first item, then the headers, then the data) is pretty closely comparable to CSV in size.
* https://en.wikipedia.org/wiki/Delimiter#ASCII_delimited_text
Escaping is too easy to ignore, until it comes back to bite you in the ass. Balanced delimiters break as soon as the content contains invalid data (or is delimited according to a different convention than the container).
I think the human-editability you mention may be a major reason why comma-separated has been such a practical winner?
Since ASCII 28 to 31 are invisible or indistinguishable in most editors, I suspect a format relying on them would be subject to routine corruption.
(I have some experience working with a "legacy" "binary" format which DOES use those ASCII separator values... https://en.wikipedia.org/wiki/MARC_standards)
Yeah, that's the same problem as UTF-8's almost-compatibility with ASCII. Things work, until they break spectacularly.
Ultimately, you either have to solve the issue properly, or keep it visible enough that people are forced to handle it.
Invisible characters also mean you can no longer call a data format "human readable".
The special ASCII characters you're referring to are mostly a historical oddity today.
Any CSV format description anywhere, will of course you need to use various internal escaping mechanisms for a literal comma, or forbid them entirely. Does that "categorically make it impossible to happen"? Obviously not, words in a standard can't make something categorically impossible. How likely a protocol/standard violation is to happen of course matters to actual practical pain.
Yes, of course the ascii separator values are a historical path not taken. This thread, not started by me, was musing about why, and what-if, by people who, believe it or not, have plenty of experience dealing with data, and bad data.
Unless "editing by hand" means using something like excel?
Note that length-prefixes aren't necessarily safer when nested. For example, `line(10 bytes, [cell(1000 bytes, "xxxxxxx...`. This can lead to vulnerabilities like buffer overflows (naively allocating 10 bytes for the line, then trying to put 1000 bytes into it)
This is why we have editors that do syntax checking as we type. Editing CSV by hand is just not important enough for such editors to be commonplace. CSV files are exported and imported.
And since we have JSON, whose spec is stable, known and at this point consistently implemented, CSV should really just be sun set rather than "fixed".
Also, the simplest escaping format is wrap in quotes and represent quotes by typing quotes twice. That's easy to remember and relatively hard to screw up even when editing by hand.
I generally use CSV as an import/export format, and would like for software to at least support the option (even if CSV/TSV remains the default).
Also, geneticists have been forced to rename certain gene names because Excel messed them up[0].
It would be nice if this standard would be generally accepted. This standard is from 2005 and it is pretty clear that it is not universally used.
[0] https://www.theregister.com/2020/08/06/excel_gene_names/
Which is an inevitability for anything that's been a dependency for long enough.
Either you admit that things have to change and sometimes shit that's been secretly broken for years will have to change too, or you give in to the shit and admit that you can never solve any problems that made it to production.
Both answers suck, but one is a lot more suck so I choose the "fix it anyways and if something else turns out to have depended on the brokeness then add that to the list of shit to fix" approach whenever possible.
Examples:
- double quotes in text fields
- using the wrong decimal separator, should it be "," or "."
- using the wrong column separator, should be "," or ";"
> This standard is from 2005 and it is pretty clear that it is not universally used.
Unfortunately a lot of non-conforming software can't change or they risk compatibility problems with old versions of themselves. 2005 is relatively recent in the history of CSV files.
Currently R data.table supports read and write csvy format. https://www.rdocumentation.org/packages/data.table/versions/...
http://thomasburette.com/blog/2014/05/25/so-you-want-to-writ...
It made an impact on my approach to CSV handling, and helped me understand, "just because I can, doesn't mean I should".
RFC 4180 also fails to define the default text encoding. It seems that text encoding can be given only when MIME type is given for the file.
JSON is so much easier.
Though be careful if you communicate with people using both locals. Things interpreted as dates will get corrupted in fun ways.