Perfect deployment of David Wheeler's aphorism:
> All problems in computer science can be solved by adding another level of indirection.
https://en.wikipedia.org/wiki/David_Wheeler_(computer_scient...
Perfect deployment of David Wheeler's aphorism:
> All problems in computer science can be solved by adding another level of indirection.
https://en.wikipedia.org/wiki/David_Wheeler_(computer_scient...
Maybe if editors are fixed up we could adopt ASCII Separated Values (ASV) as the new standard.
Unicode works fine there too, so it makes no nevermind to me which flavor people use. I just think it's funny how "everything old is new again".
,<FS> for fields \n<RS> for records
This removes ambiguity in parsing and remains user readable. It's also relatively easy to auto-fix files edited by users in normal editors.
It also mostly removes need for escaping.
It's also smaller or same size as unicode multibyte characters (haven't checked).
How could you make the difference with a standard CSV file if it looks like a standard CSV file?
They explain why they don't use control characters. Editors are not consistent in how they show control/zero-length characters:
https://github.com/SixArm/usv/tree/main/doc/faq#why-use-cont...
https://github.com/pmarreck/elixir-snippets/blob/master/prin...
I'd also prefer if escapes were done in the "traditional" manner of, for example, "\t" for a tab because you can then read in stuff with something like input.split("\t").map(unescape); you know any actual tab character in the input is a field separator, and then you can go through the fields to put back the escaped ones.
What about input lines like 'asdf\\thjkl\tzxcvb'? That should be two fields, one the string ‘asdf\thjkl’ and the other the string ‘zxcvb.’
I think that your way is a bit like trying to match context-free grammars with a regular expression. The right way is to parse the input character by character.
The "\t" in "split" is not a "slash-tee" but an actual tab character and then escape sequences in fields are handled by the "unescape" function.
If you still need to implement escape mechanism, might as well do CSV/TSV.
https://github.com/SixArm/usv/tree/main/doc/faq#why-choose-u...
As for still needing escapes, using obscure symbols instead of ones that are extremely common in writing inherently means needing far far faaaaaaar fewer of them.
And yes, I read README and source code, so I know that newlines are optional, existing tools don't generate them, and multi-line examples are basically fake.
It doesn't have to be all squished in one line, it just doesn't hurt anything. Visually splitting squished lines for presentation or perusal is trivial because of the record separator.
> You are not going to be editing this in regular editor
I know (or at least I think) that you meant this in relation to squished lines getting very long, but maybe we can talk about it in a broader context, since record splitting is trivial...
One could easily say these same words about documents written in right-to-left languages. But people in Israel manage to create files too somehow, so that's clearly not an insurmountable barrier.
And yet, that's explicitly not the semantic purpose of those glyphs. The actual delimiters already exist at a lower code point. If we're asking editors to semantically support delimiters we should be asking them to support the semantic delimiters.
If it turns out that escaping is needed, it will still be far rarer than escaping commas and newlines.
Perhaps there is a software developer version of "Needs more cowbell" called "Needs more complexity"
Computer languages generally use the Latin alphabet. And even in a case like APL, which some HN commenters call "hieroglyphics", the number of symbols is limited and each is precisely defined (cf. potentially up to 1.1 million Unicode symbols and "emojis" that are open to interpretation).
> For every interesting HN post, there’s at least one smug commenter who thinks he knows better, but actually doesn’t
https://github.com/SixArm/usv/tree/main/doc/faq#why-use-cont...
If you don't understand why something is the way it is, it might be better to start with a question than with a statement implying the tech misses existing tech. Chesterton's fence still applies, and ignoring it means you're outsourcing your work to others. RTFM is a perfectly valid answer at that point.
My point above, though, is that everyone has opinions and you don’t have to be a dickhead about “correcting” them.