>
RFC4180 isn't an Internet standard. Says so right at the top in the first paragraph.I didn't say it was an Internet standard. I said standardised. Ok, I'll concede it is more of an informal or de facto standard but when IBM, W3C, IETF, OKF and others all publish the same parsing rules for CSV, it's hard to agree when people make statements like "there's no standard in CSV". The problem isn't that there isn't a standard to CSV, the problem is that people often don't follow those conventions. But you have that same problem with other file formats too.
> I think I have a few thousand files on my computer right now whose names end in ".csv", and I'll bet money not one of them agrees with RFC4180 to the letter except by accident of the data itself, and that's a Real Problem to me.
That's just conjecture. And even if that were proven true, it's still only anecdotal. That said, I do sympathise with your point. But you could make the same argument for
- JSON files that don't follow spec (support for comments, aren't UTF-8 encoded, have been manually written so don't follow the escaping rules correctly and thus only parse correctly by chance).
- XML files that have been manually cranked and so don't follow schema
- HTML documents that don't follow specification and thus browsers do a lot of non-specification interpretation work to render correctly
The IT industry is littered with example of people not following the docs. CSV isn't unique in that regard.
> I can concede that two parties could agree to interchange according to RFC4180, but as a general format I maintain that for archival and interchange purposes CSV cannot be divorced from the rather complex schema I alluded to without data loss.
I wasn't commenting on whether it's a better format than another file format. I was commenting on your points about standardisation saying there are an abundance of published documents on how to read and write a standard CSV file and C-style escaping isn't part of that specification.