Any such solution will always be broken, proof is left as an exercise for the reader. Hint: use Ogden's lemma.
Any such solution will always be broken, proof is left as an exercise for the reader. Hint: use Ogden's lemma.
So if you decide not to support "strings", just "text", then those separators Just Work.
But if I was writing a similar standard today and someone was pressuring me that I really had to make some kind of mechanism for nesting, I'd use something that transforms the field to remove the need for escaping, in a big and obvious way. Probably base64. That prevents most of the implementation issues you see with CSV.
CSV needs escaping, so don't do CSV.
> ASCII separators are equally as broken. How do you represent a string containing the record separator in a record-separated file?
ASCII separators may not be perfect, but they're certainly less broken than CSV. With CSV, you have the pervasive problem of how to represent a common, everyday character in the file. With ASCII separators, you only have the problem of representing an unusual, untypeable character, which is much less troublesome.
It's the temptation to just concatenate or split based on commas, or to hack a broken implementation of quotes on top, that causes so many issues.
There are two main options for that. Make it so that poorly parsing with self-written code is still very likely to do the right thing, or make the parsing hard enough that people go get a library.
I think a big reason they never took off is also that there's no visual representation of them, or input method for them.
I'd be happy to use them instead of CSV if I could edit records in Notepad or TextEdit the way I can with commas and quoted strings. See the separators, type them.
But of course once you do that, somebody now wants to insert a list of five values separated by unit or record separators into a field. They'd just be common ASCII characters along with tab, CR, LF, etc, and need to be escaped the same.
There's no escaping escaping...
> There's no escaping escaping...
This is the truth.
And in practice I do not expect CSV with embedded nulls to work properly, so there's already precedent to reject certain characters entirely in a CSV-like format.
This difference carries over to C, where NULL got the job done for string termination pretty darn well, even if there were strong critiques to be weighed against it.
From a purely linguistic standpoint, the choice is sound. Obviously it's objectively a disaster but for unrelated reasons.
They can, only some functions stop at the first zero:
char s[] = "gogo\x00gogo";
printf( "%d\n", (int)sizeof( s ) );Consider the following program:
#include <stdio.h>
#include <string.h>
int main()
{
const char * const s = "gogo\x00gogo";
const char t[] = "gogo\x00gogo";
printf("%d,%d,%d,%d\n",
(int)sizeof s,
(int)strlen(s), // DO NOT EVER USE THIS FUNCTION
// YOU WILL BE FIRED IF YOU DO
(int)sizeof t,
(int)strlen(t) // DO NOT EVER USE THIS FUNCTION
// YOU WILL BE FIRED IF YOU DO
);
}
> 8,4,10,4> They can, only some functions stop at the first zero:
> char s[] = "gogo\x00gogo";
And in what human language is that `\x00` a valid character? What rune/glyph/symbol represents it?
Nothing, you say? Well, looks like a good choice to me.
CSV is more broken than you imagine. Try opening an American CSV file with Excel in Germany or France. You’ll discover that it doesn’t parse, and the reason is that the ‘C’ in CSV apparently stands for “could be a semicolon too” because comma is the decimal separator in these countries and therefore must be available for use inside values.
- Make sure to write your own parser! The format is so simple!
- Is the first line a header or not? The RFC says you should look at the MIME type to tell. You can't make this up.
- Is the line terminator "\r\n", "\n" or "\r"?
- What happens if the number of records for a line is incorrect? Well, that SHOULD'nt happen...
- Make sure to involve Excel, nothing ever went wrong feeding data to Excel.
- Are end-of-line spaces relevant or not? Should we trim them?
There is no defending of CSV, it's a broken format and it's broken beyond repair. The point is that the separator is immaterial. Replacing commas with something else gains us exactly nothing, and results in an identically broken format.
For one, I think you'd almost certainly use Record Separator, and \r and \n would be normal text characters inside of a field. And no trimming spaces.
And the behavior upon encountering corruption isn't really up to the spec, is it?
Headers might be more trouble than they're worth, but if they're in I'd say they need to be mandatory. One code path.
You'll find yourself with a much less broken format.
Always the seemingly simple solutions have most issues in real world... At some level still trying to edit stuff by hand might be where we have gone wrong for long time. Why don't we have any sensible agreed structured format for which tools would be on all and every device...