Show HN: USV = Unicode Separated Values
github.com
github.com
(Just imagine using USV to export a set of blog articles, except you can’t because you talk about USV in the blog articles using the picture characters)
If the data is HTML, then you can encode the special characters using ampersand notation. And that can be kept out of the USV spec entirely as an external convention. So that is to say, use the native escaping of whatever data format you are storing.
A similar principle is used in embedding JS or JSON in HTML. A sequence like "</script" could occur in a JS or JSON string literal. For that reason there is a way to escape the backslash \/. Or else the slash can be encoded numerically. This is all provided by the embedded language; the embedding language (HTML) knows nothing about it.
The user, or software, doing the embedding just has to know about this and put in the escapes if they occur in the to-be-embedded data.
USV is nice the way it is; don't make a complicated skullduggery out of it with escape mechanisms. The whole point of choosing those characters is to avoid it. If you're gonna have escapes, then might as well stick with commmas for field separation, newlines for records.
What's your opinion of USV being simple (as you describe above) and USVX being a different standard that is USV + extras?
White space trim means that if you don't expect trim, your significant white space is eaten, causing a bug.
Most of the time, USV will be machine-generated. If the machine generates a field with leading or trailing space, then that's what the data is.
What you want is null characters. Multiple null characters can be used for multiple separating roles/levels. Microsoft implemented this idea in the Registry; there are string tables where each string is null terminated, and then an extra null indicates the end of the table. e.g.
foo<NUL>bar<NUL><NUL>
This is an empty table containing no strings: <NUL>
This is a table with one empty string: <NUL><NUL>
This looks like it should extend to more dimensions with more NULs: record, table, group, file, ...The long-winded point here is that your fictitious blogger trying to write about the scheme in all likelihood will not have any luck embedding literal NULs into their article, so we are are safe.
Bad news: it looks like they got ␟ and ␞ mixed up, there are trailing ␞␟ characters after each record (should just be an ␞ between each record), and no characters at all after the last record:
$ cat test.csv
one,two,three,
four,five,six
seven,eight,nine
$ vd test.csv -b -o test.usv
opening test.csv as csv
saving 1 sheets to test.usv as usv
Think about what you're doing.
test.usv save finished
$ cat test.usv
one␞two␞three␞␟four␞five␞six␞␟seven␞eight␞nine␞␟$
$We do not have anyone who is deeply familiar with USV spec details on the team, and could use your help to understand what needs to be fixed.
The characters are pretty much just for this, when would you ever need to put them in the data? If you're working on tooling for unicode itself? Just use... any other format. Literally any.
For example, import/export is a breeze on the command line with typical Unix tools such as awk, sed, grep, mlr.
No-escape makes international standardization clearer for Backus–Naur form (BNF), the same way other standards use no-escape BNF.
For example, no-escape is the same choice as with the existing international standard for IANA tab separated values.
Someone will generate wrong usv e.g. escape or encapsulate values.
And if you get that wrong format no one cares if they gave you a wrong format. => Your parser grows and grows...
1. Editors don’t return carriage on RS
2. There is no US or RS key on my keyboard
I am keeping my tsv. Tab is _never_ a data character. And if return carriage is a data character, you probably shouldn’t use a dsv file in the first place.
Until it is. Just like a semicolon for CSV.
> And if return carriage is a data character, you probably shouldn’t use a dsv file in the first place.
Agree.
EDIT:
> I am keeping my tsv
As a solely import-export data format (ie no manual editing) it has some virtue. But at this point you can just use XML, CLIXML or even SQLite *shurg_emoji*
USV with 2 units by 2 records by 2 groups by 2 files:
a␟b␞c␟d␝e␟f␞g␟h␜i␟j␞k␟l␝m␟n␞o␟p
USV with typical shell commands and pretty output: $ echo "a␟b␞c␟d␝e␟f␞g␟h␜i␟j␞k␟l␝m␟n␞o␟p" |
sed 's/␟/,/g; s/␞/\n/g; s/␝/\n---\n/g; s/␜/\n===\n/g;'
a,b
c,d
---
e,f
g,h
===
i,j
k,l
---
m,n
o,p