Why isn’t there a decent file format for tabular data?
successfulsoftware.net
successfulsoftware.net
Won't it all be in one long line?
> Typing \u001F or \u001E in some editors might be a faff, but it is hardly a showstopper.
Doesn't seem negligible when the editability issues with CSV were also minor.
> No escaping. If you want to put \u001F or \u001E in your data – tough you can’t. Use a different format.
Wanting those characters specifically is rare, but wanting to safely store an arbitrary unicode string - maybe a filename - is very common. Parsers might invent their own escaping to handle this.
Depends on whether editors start a newline for \u001E. Notepad++ doesn't. Than is definately going affect readability. Hmmm.
>Doesn't seem negligible when the editability issues with CSV were also minor.
The issue with CSV, as far as I am concerned, is the escaping. And that isn't minor from either a readability or parsing point of view.
>Wanting those characters specifically is rare, but wanting to safely store an arbitrary unicode string - maybe a filename - is very common. Parsers might invent their own escaping to handle this.
Are \u001F or \u001E even legal filename characters on any OS? These codes could turn up by chance in a blob of binary data. But it isn't intended for binary data.
For parsing I'd agree, as you've eliminated the need to handle escaping by declaring some characters off-limits. But for manual editing, to start a .usv file with most editors people would be copying the characters from google rather than just being able to type commas and newlines.
> Are \u001F or \u001E even legal filename characters on any OS? These codes could turn up by chance in a blob of binary data. But it isn't intended for binary data.
Generally I believe NUL and / are the only ASCII characters considered safe to never occur in file/directory names.
That is an issue. However Notepad++ has an "ASCII codes insertion panel". Many other editors probably have something similar. So it isn't too hard.
[0] https://stackoverflow.com/questions/8695118/what-are-the-fil...
The main difference is the lack of escaping, which is what makes CSV slow to parse and error prone. With escaping you have to parse from the start of the file, which makes multi-threading hard, which limits how fast you can parse it.