Show HN: Comma Separated Values (CSV) to Unicode Separated Values (USV)
crates.io
crates.io
Perfect deployment of David Wheeler's aphorism:
> All problems in computer science can be solved by adding another level of indirection.
https://en.wikipedia.org/wiki/David_Wheeler_(computer_scient...
If you still need to implement escape mechanism, might as well do CSV/TSV.
https://github.com/SixArm/usv/tree/main/doc/faq#why-choose-u...
As for still needing escapes, using obscure symbols instead of ones that are extremely common in writing inherently means needing far far faaaaaaar fewer of them.
And yes, I read README and source code, so I know that newlines are optional, existing tools don't generate them, and multi-line examples are basically fake.
It doesn't have to be all squished in one line, it just doesn't hurt anything. Visually splitting squished lines for presentation or perusal is trivial because of the record separator.
> You are not going to be editing this in regular editor
I know (or at least I think) that you meant this in relation to squished lines getting very long, but maybe we can talk about it in a broader context, since record splitting is trivial...
One could easily say these same words about documents written in right-to-left languages. But people in Israel manage to create files too somehow, so that's clearly not an insurmountable barrier.
And yet, that's explicitly not the semantic purpose of those glyphs. The actual delimiters already exist at a lower code point. If we're asking editors to semantically support delimiters we should be asking them to support the semantic delimiters.
If it turns out that escaping is needed, it will still be far rarer than escaping commas and newlines.
I'd also prefer if escapes were done in the "traditional" manner of, for example, "\t" for a tab because you can then read in stuff with something like input.split("\t").map(unescape); you know any actual tab character in the input is a field separator, and then you can go through the fields to put back the escaped ones.
What about input lines like 'asdf\\thjkl\tzxcvb'? That should be two fields, one the string ‘asdf\thjkl’ and the other the string ‘zxcvb.’
I think that your way is a bit like trying to match context-free grammars with a regular expression. The right way is to parse the input character by character.
The "\t" in "split" is not a "slash-tee" but an actual tab character and then escape sequences in fields are handled by the "unescape" function.
Maybe if editors are fixed up we could adopt ASCII Separated Values (ASV) as the new standard.
Unicode works fine there too, so it makes no nevermind to me which flavor people use. I just think it's funny how "everything old is new again".
,<FS> for fields \n<RS> for records
This removes ambiguity in parsing and remains user readable. It's also relatively easy to auto-fix files edited by users in normal editors.
It also mostly removes need for escaping.
It's also smaller or same size as unicode multibyte characters (haven't checked).
How could you make the difference with a standard CSV file if it looks like a standard CSV file?
They explain why they don't use control characters. Editors are not consistent in how they show control/zero-length characters:
https://github.com/SixArm/usv/tree/main/doc/faq#why-use-cont...
> For every interesting HN post, there’s at least one smug commenter who thinks he knows better, but actually doesn’t
https://github.com/SixArm/usv/tree/main/doc/faq#why-use-cont...
If you don't understand why something is the way it is, it might be better to start with a question than with a statement implying the tech misses existing tech. Chesterton's fence still applies, and ignoring it means you're outsourcing your work to others. RTFM is a perfectly valid answer at that point.
My point above, though, is that everyone has opinions and you don’t have to be a dickhead about “correcting” them.
Perhaps there is a software developer version of "Needs more cowbell" called "Needs more complexity"
Computer languages generally use the Latin alphabet. And even in a case like APL, which some HN commenters call "hieroglyphics", the number of symbols is limited and each is precisely defined (cf. potentially up to 1.1 million Unicode symbols and "emojis" that are open to interpretation).
https://github.com/pmarreck/elixir-snippets/blob/master/prin...
Imagine the amount of pain that could have been spared if we had done it right from the start some 50 years ago.
Perhaps this would make an interesting personal project. Are you aware of any hurdles, missing key features, etc. that previous attempts at creating such a format have run into (other than adoption, obviously)?
(That said every text editor since ever should have had a "table mode" that uses the ASCII field/record seperators (or whatever you choose), I was always confused why this isn't common. Maybe vim and emacs do?)
I'm a Notepad++ person. When I needed to mock-up data typing the characters was easy-- just ALT and the ASCII code on the numeric pad. It took a bit to memorize the codes I needed to use. Their visual representation is just inverse text and initials.
I absolutely HATE this Parcquage.
(I don't think everyone has moved to it. I had never heard of it myself.)
You won't remember Parquet in 15 years, but you will have CSV files in 50 years.
You're probably right about CSV but probably not parquet. Parquet is already 11 years old, there are vast data warehouses that store parquet, it's first class in the spark ecosystem, and a key component of iceberg. Crucially, formats like parquet are "good enough" for a use case that doesn't appear to be going away. There is a high probability in my estimation that enough places are still using them in 15 years to be memorable even if it isn't as common or as visible.
Considering it also separates records with newlines, they really should have replaced newlines with "\n" and require escaping "\" with "\\".
Edit: this is what I have so far: https://github.com/tmccombs/ssv
Nice job. I hope you come back to finish the project eventually.
In time, editors and file browsers should come to render separators visually in a logical way.
Feel free to correct me, but I figure that as long as data can be from 0x00 to 0xFF per byte, no format that uses characters in that range will ever be safe. I’m not a big C developer but I figure the null terminated strings have the same limitation.
But if its something entered by keyboard you should be ok to use control codes.
Personally, I find tab and return to be fine for text driven stuff. Shows up in an editor just like intented.
The advantage of ASV is not that you can't have invalid or insecure data, it's that valid data will almost never contain ASCII control characters in the record fields themselves. Commas, quotation marks, and backslashes, meanwhile, are everywhere.
(Because CSV is a terrible data exchange format in terms of information per byte. But that makes sense, because it's an intentionally human readable data exchange format, not a machine format)
Hence https://github.com/SixArm/usv/tree/main/doc/faq#why-choose-u...
https://en.wikipedia.org/wiki/Optimus_Maximus_keyboard
(It's been almost 20 years and you still can't get one...)
Where's the key on my keyboard yo make one?
The point of text-based formats is that you can edit them in a text editor by hand trivially, if typing the character is nontrivial, then it entirely defeats the point (that's also why USV ads very little value IMHO).
There is no perfect solution, but I’d rather open a text file in a decent editor than having to deal with the escaping hell that is CSV.
They could have chosen the pipe character “|” at least, but the comma is the thousand separator in many languages (number formatting is kind of important for tabular data, if you ask me) and also, you know, general prose.
Alt gr+E? Like it's shown on the keyboard.
European ones tend to have one. US keyboards don't.
Not sure about British. Are they different from US?
There's one on French keyboards actually!
And it was there even before we got euro coins in our hands (I know this because I'm still using my first (mechanical) keyboard that I got with my first own PC in 2001: and there is a “€” symbol on it)
BELL Ctrl-G
RECORD SEPARATOR Ctrl-^
UNIT SEPARATOR Ctrl-_
ESCAPE Ctrl-[
You can think of the Ctrl key as clearing the two most-significant bits of the letter's ASCII code. Not all key combos are supported in all environments. Notepad++ doesn't support Ctrl-] (GROUP SEPARATOR) at all, but does support e.g. SHIFT OUT as Ctrl-Shift-N, for instance. The Windows CMD.EXE command line supports many combinations (but not UNIT SEPARATOR, unfortunately), displaying them as e.g. ^[ or ^G in the console.
[1] https://upload.wikimedia.org/wikipedia/commons/1/1b/ASCII-Ta...
If you need to train your employees on that, it's not "very easily".
"Very easily" is when I can take any family member who's seen a computer in their life, give them a keyboard and they can figure it out on their own without Google in 2 seconds (like csv).
Unfortunately that ship has sailed. We have standards for escaping commas, escaping quotes, it’s escaping all the way down
This can be fixed in the font
> Imagine the amount of pain that could have been spared if we had done it right from the start some 50 years ago.
I think it's Putt's Law: If you design something for idiots, someone will make a better idiot. In this case, it took less than 15 years.
I need to zoom to be able to tell these apart, so I'll need editor support for it to be convenient to work with these anyway. And then clicking through to the comparisons, it demonstrates the difference existing support for CSV "everywhere" makes - Github renders the CSV examples nicely as tables, while again I need to zoom in to see which separator is which for USV.
Maybe once there is widespread editor support. But if you need editor support for it to be comfortable anyway, then the main benefit vs. using the old-school actual separator characters goes out the window.
I don't really get this project at all.
I also think there are failed lessons here that reduces the incentive for switching.
E.g. If you're going to improve on CSV, a key improvement would be to aim to make the format trivially splittable, because the lesson from CSV is that when a format looks this trivial people will assume they can just split on a fixed string or trivial regex, and so the more you can reduce the harm of that the better.
As such, I'd avoid most of the escaping they show, especially for line endings, and just make RS '\n' the record separator, or possibly RS '\n'*. Optionally do the same for US. Require escaping LF immediately after RS/US, and only allow escaping RS, so unescaping can be done with a trivial fixed replace per field if you have a reason to assume your data might have leading linefeeds in fields - a lot of apps will get away with just ignoring that.
Then parsing is reduced to something like `data.split(RS).map{|row| row.split(US).map{|col| col.gsub(ESCAPE,"\n") } }` (assuming RS, US, and ESCAPE are regexps that include the optional trailing linefeeds and escapes leading linefeeds respectively). Being able to copy a correct one-liner from Stackoverflow ought to avoid most of the problems with broken CSV/TSV parsing.
I'm also not convinced adding GS, FS, ETB is a good idea, partly for that reason, partly because a lot of the tools people will want to load data into will not handle more than one set of records, and so you'll end up splitting files anyway, in which case I'd just use a proper archive format... Those characters feels like they're trying to do too much given they're "competing" primarily with CSV/TSV.
Their spec also needs to talk about encoding, because unless I've missed something, they only talk about codepoints, and they're likely to e.g. get people splitting on the UTF8 sequence etc. This to me is another reason for using the ASCII values - they encode the same in ASCII based characters sets and UTF8, and so it feels likely to be more robust against the horrors of people doing naive split-based parsing.
The point is that their stated "advantage" does not exist for me. I still need to make changes to my setup to handle them. In which case why should I pick this option? (as you can see elsewhere, especially as this isn't the only issue I have with their format choices).
> And the main benefit isn't anything to do with the editor, I have no idea what you meant by that.
The main benefit relative to using the actual control characters is only the tool support. Where this does not work for me without making changes anyway to how the symbols are displayed anyway. Hence that "advantage" does not actually buy me anything.
The thing about the actual separators is that an editor could and should probably display them as they were intended, as data separators. It should be a setting in an editor you control, sort of like how you control tab width and things like that.
Just because a glyph is "invisible" doesn't mean it has to actually be invisible.
The symbols for the separators are hard to read, like you're pointing out, which means someone would eventually replace them with some other graphical display, in which case you were just as well off with the actual separators themselves.
They would have been better off advocating for editor support for actual separator display.
> Unicode separated values (USV) is a data format that uses Unicode symbol characters between data parts. USV competes with comma separated values (CSV), tab separated values (TSV), ASCII separated values (ASV), and similar systems. USV offers more capabilities and standards-track syntax.
> Separators:
>
> ␟ U+241F Symbol for Unit Separator (US)
>
> ␞ U+241E Symbol for Record Separator (RS)
>
> ␝ U+241D Symbol for Group Separator (GS)
>
> ␜ U+241C Symbol for File Separator (FS)
>
> Modifiers:
>
> ␛ U+241B Symbol for Escape (ESC)
>
> ␗ U+2417 Symbol for End of Transmission Block (ETB)
>
> ␖ U+2416 Symbol For Synchronous Idle (SYN)
That would be neat :)
Edit: Apparently, kinda (e.g. https://www.compart.com/en/unicode/U+241E )
Not the most creative....
```USV works with many kinds of editors. Any editor that can render the USV characters will work. We use vi, emacs, Coda, Notepad++, TextMate, Sublime, VS Code, etc.```
I loaded an example in my fairly generic Emacs and it worked out of the box. The separators were pretty small so I had to increase my font size to distinguish US from RS. And of course I have no idea how to enter those characters. I'm sure there is, but cut & paste worked.
It feels like instead of fixing it properly, they went with an option that will still need tool improvements, will be controversial, and adds unnecessary details (e.g. the SYN they've added will be an active nuisance and I'd be willing to bet will get ignored by enough tools to become a hazard to data integrity).
I quite like an initiative to make use of proper record and unit separators, but this feels poorly thought through in several respects (e.g. their quirky escape characters that adds differently depending on the class of the following character will be a 'fun' source of bugs; that splitting records on LF requires three characters almost certainly will mean a number of tools will incorrectly treat those three characters as a unit, etc. -- these assumptions are based on how slapdash a lot of CSV parsing and generation is; if you want to compete with CSV you ought to learn those lessons)
You cannot edit it in regular editor, like csv/tsv/jsonlines.
There is no schema or efficient storage, like binary formats.
There is no wide library support.
Not all data is representable.
ASCII 1963 had 8 separators, 1965 reduced it to 4, and named them. See 6.3.12 of https://dl.acm.org/doi/pdf/10.1145/363831.363839
The only time you need to escape a character is if it's a control character that's rarely used, unlike the " and , characters
Literally untrue. (And were it true, it still wouldn't be a reason why one should use this over CSV—not sure what's so hard to grasp about the conversational/contextual premise here.)
If only there were shortcuts on modern operating systems to allow us to do things that aren't readily on our keyboards. Like upper case characters. Or copy and paste. Or close windows. Our lives would be so much better.
If ASV had caught on, there could be common shared shortcuts to type them, and fonts would regularly display them (just like the unicode characters proposed). But CSV was simple enough and readily type-able.
> There is no schema or efficient storage, like binary formats.
I'm not quite certain where you're trying to go with this. Binary formats aren't really meant to be human readable in an average text editor. It doesn't know to differentiate 1, 2, 4, or 8 bytes as an integer or a float. Even current hex editors to make it easier to navigate these formats don't really know unless you are able to tell it somehow.
> There is no wide library support.
It's a critical mass problem. Not enough people are using them, so no libraries are being made.
> Not all data is representable.
I'm not quite certain what data couldn't be represented. f you can represent your data in CSV, you can represent it in ASV. It's all plain text that gets interpreted based on what you need. They're nearly a 1:1 replacement. Commas get replaced by unit separators, new lines get replaced by group separators. Then you have record and file separators to do with for further levels of abstraction if you need.
Now, the readme actually has that optional newline separator thing, but the optionality of it makes it completely useless, it seems like an after-thought. Fr example the first "real" USV writer I found, the "csv-to-usv", does not put them [0] and thus makes uneditable files.
And if we are going to end up with uneditable files, might as well go with something schema-full, like parquet or avro. You are going to have the same "critical mass problem", but at least the tooling is much better and you have neat features like schemas.
[0] https://github.com/SixArm/csv-to-usv-rust-crate/blob/30a0324...
What do you do if you receive data already containing a unit separator, or a group separator, and you need to put it into a field? The whole value proposition of ASV over, say, TSV is that you should never need to escape anything, but that's only possible by rejecting some input data.
Everyone thinks they can do better, but nothing's more widely supported (for a sufficiently generous definition of 'supported')
I randomly generated some CSVs and fed them into Excel and Numbers and they were differently interpreted.
Generally my only interaction with CSV itself is to fling it through https://p3rl.org/Text::CSV since that seems to be able to get a pretty decent parse of every sort of CSV I've yet had to deal with in the wild.
Countries that use , as the thousands separator (e.g. 1,000) use ; as the CSV separator.
Why? Because that’s how Excel does it.
Tab separated files are much better imo in not getting confused with the delimiter for a sufficiently sane tsv file.
while IFS="$(printf \\a)" read -r field1 field2...
do ...
done
This works just as well as anything outside the range of printing characters.Heaven help you if you cat the source file in a shell, though!
OB vertical tab
<content>
1C file separator
0D carriage return
I wrote one of the most popular translators for MLLP, which converts it to HTTP [1].---
P.S. Ironically, HL7 messages have something literally called a "field separator" but don't use the field separator character, usually they use vertical bar.
That's a fair point. But you could argue that when the abuse is so widespread, it becomes a defacto part of the format (even if it isn't in the RFC).
Yes, rather than using U+1F (the ASCII and Unicode unit separator), he proposes using U+241F (the Unicode symbol for the unit separator). I almost feel like this must be an early April Fool’s joke?
Also, he writes ‘comprised of’ rather than ‘composed of’ or ‘comprises’ throughout his RFC.
My bet is that this will lead to implementations that wrongly treats "␞␛\n" (RS ESC \m) as the real record separator, the same way lots of "CSV" implementations just split on comma and LF.
Seems to me if you're going to add support for something like that you should just bite the bullet and declare an LF immediately following an RS as part of the record separator, or you're falling in the same trap as CSV of being "close enough" to naively splittable that people will do it because it works often enough.
[0] https://github.com/SixArm/csv-to-usv-rust-crate/blob/30a0324...
https://github.com/SixArm/csv-to-usv-rust-crate/blob/30a0324...
Speaking of which, last time I had a control code heavy file open in Sublime, it actually did show the control codes as special characters, and it was possible to copy/paste those. This proposal is so bad I suspect it will become a standard.
"We tried using the control characters, and also tried configuring various editors to show the control characters by rendering the control picture characters.
First, we encountered many difficulties with editor configurations, attempting to make each editor treat the invisible zero-width characters by rendering with the visible letter-width characters.
Second, we encountered problems with copy/paste functionality, where it often didn't work because the editor implementations and terminal implementations copied visible letter-width characters, not the underlying invisible zero-width characters.
Third, users were unable to distinguish between the rendered control picture characters (e.g. the editor saw ASCII 31 and rendered Unicode Unit Separator) versus the control picture characters being in the data content (e.g. someone actually typed Unicode Unit Separator into the data content)."
- https://github.com/SixArm/usv/tree/main/doc/faq#why-use-cont...
Also I am not convinced about the need for an escape character. If you really need to use ASCII unit or record separators as data - tough use a different format.
If only editors would display the ASCII unit separator (Notepad++ does) and treat the ASCII record as a carriage return (Notepad++ doesn't) then .asv format would be a huge improvement on CSV.
Maybe I'll consider it when it does not belong to a company, has two more zeros in the number of stars, and has RFC/ISO attached to it. Because right now it is not much more of a "standard" than a hobby project I create on a whim.
Man, I preferred it when people could just write up and propose things. The insufferable "is that professional?", "What about consensus?", "Wow the ego to propose something".
Time to return to monke.
Additionally I'm sure that those ("they have such a big ego") types of comments and thoughts existed in the early internet as well since its fairly human reaction whenever anyone tries to build or propose something that disrupts the status quo.
By the way, USV doesn't belong to a company. It's just me. RFC/ISO is work in progress, and I submitted IETF ID 00 last week.
https://github.com/SixArm/usv/tree/main/doc/comparisons#asci...
Main point: ”USV provides typically-visible letter-width characters (such as Unicode 241F), whereas ASV provides typically-invisible zero-width characters (such as ASCII 31).”
This way people can initially use the visible glyphs while editors don't support the format, and this will always be supported. But, as editors add support and start to generate the files via tools or manually in tabular interfaces where the codes themselves disappear, usage will automatically transition over to the control codes.
Why would this go in-band inside a document format? Just why? If you want keep-alives, use a kind of connection that supports out-of-band keepalives.
If you download the same document twice, and the second time the server is heavily loaded (or it's waiting on some dependency, or whatever), presumably the server will helpfully generate some SYNs in the middle of the document to keep the connection alive (?), but now you've got the same document "spelled" two different ways, that won't checksum alike.
SYN along with the weirdness of
> Escape + [non-USV-special] character: the character is ignored
means that you have arbitrarily many ways of writing semantically-same documents.
Why does a file format need a transport protocol?
---
Existing transport protocols (TCP, QUIC) already provide this.
As for including commas in your data, it could just have been managed with a simple escape character like a \, for when there's actually a comma in your data. That's it.
Not quite. What if there is a \ in your data? Then you have to escape that.
No problem, any character following a `\` is a literal character. `\\` => literal `\`. `\,` => literal comma. `\a` => literal `a`, etc.
Parsing this is easy, generating it is easy, and there is only one rule to remember for humans reading or generating it.
Each rule added for parsing is one more added complexity and point of failure.
SomeCommas,MoreCommas,OnlyOneComma,ALotOfCommas
,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,
It will be difficult to find pathological cases for this grammar, because they don't exist.
As a result spreadsheets almost always fail to automatically parse a CSV.
I do like the idea of having a dedicated separator character, that would work right worldwide. And then just standardize the use of a dot as decimal separator in these files.
*edit Apparently emojis don't fly here, but it was an index finger pointing right.
Only a matter of time before something breaks catastrophically but it hasn't happened yet.
e.g. if it's a frowny face you know it's an invoice.
is used for tuples, for sets, for frozen sets, for pickles, for bytes and for bytearrays.
I thought it was pretty ingenious but clearly I’m not the only one to think of it.
* Editors will play nicely with the graphical representation. If you need better graphics, it's done with font customization, which everyone already supports.
* It announces that the data is source text, vs transmitted bytes. The type/token distinction is not easy to overcome.
* It sits way out in Unicode's space where a collision is unlikely. The whole reason why CSV-type formats create frustration is because the tooling is ad-hoc, never does the right thing and uses the lower byte spaces where people stuff all kinds of random junk. This is the "fuck it, you get the same treatment as a Youtube video id" kind of solution.
That said, if used, someone will attack it by printing those characters as input.
https://www.ietf.org/archive/id/draft-unicode-separated-valu...
https://datatracker.ietf.org/doc/draft-unicode-separated-val...
But we already keep stumbling over missing support for the on-demand quote character even with separators like comma and tab, using more exotic characters as the separator will only make it worse. The value of less escaping is negative.
We write 3.000,00 for exactly three thousand, instead of 3,000.00
Now imagine how often parsing breaks.
I think rarely used ' to group thousands is actually most sensible solution.
If you want to just double click and get to work, no.
Here's a snippet that runs it in your browser:
// Simple example to run this in your browser! But will work in Go, PHP, Ruby, Java, Python, etc...
const extism = await import("https://esm.sh/@extism/extism");
const plugin = await extism.createPlugin("https://cdn.modsurfer.dylibso.com/api/v1/module/a28e7322a6fde92cc27344584b5e86c211dbd5a345fe6ec95f1389733c325541.wasm",
{ useWasi: false }
);
let out = await plugin.call("csv_to_usv", "a,b,c");
console.log(out.text());> cdn.modsurfer.dylibso.com
Do people routinely do this - just run random code from arbitrary endpoints.
Yikes
This is a nice idea, and all, but seems unlikely to become a meaningful standard without some major backing behind that "we".
USV terminology is units, records, groups, files.
Spreadsheet equivalents are cells, lines, sheets, folios.
Database equivalents are fields, rows, tables, schemas.
The USV Rust crate provides iterators str.units(), str.records(), str.groups(), str.files(), so it's easy to get the parts you want.
It doesn't need any escaping or quoting: a field just has to be valid UTF-8.
The trick is that the delimiters are bytes that are invalid UTF-8.
The spec fits on a napkin, parsing is trivial, you can jump to the middle of a doc and find the nearest row, etc.
Main downside is you need an editor/viewer that can handle it.
What they say about display and input really depends on the specific editors and viewers that you are using (and perhaps on the fonts as well). When I use vi, I have no difficulty entering ASCII control characters in the text. However, there is also the problem with line breaking, with ASV and with USV, anyways; and they do mention this in the issues anyways.
Fortunately, I can write a program to convert these formats without too much difficulty, even without implementing Unicode (since it is a fixed sequence of bytes that will need to be replaced; however, it does mean that it will need to read multiple bytes to figure out whether or not it is a record separator, which is not as simple as ASV).
I'd previously given up using ASV because of the printability and copy/paste problems described. Replacing the control characters by their printable glyphs solves all my previous problems and is as genius as it is naughty.
I sympathize with the arguments people here present against and agree the SYN character and Group Separator are weird -- but cause no harm. I'm not bothered by the same data having multiple representations since I'm insisting on human readability rather than byte-by-byte perfection in the first place.
It took 20 minutes to convert my project and I'm very happy.
Only tooling change I had to make was adding digraphs to vim
digraph rs 9246 us 9247
etc. Easy to type directly in my .usv file. Easy to type and read in some Python consuming it.
Regardless of it becoming a standard and my lingering grouchiness about multi-byte characters, needing to use non-xterm, etc. this works very well for me.
Thanks for the digraph vi info. Thanks to you, I added a vi-specific page in the repo for your tip.
Have you maybe lost track of what post you're commenting under?
(I believe the answer is no BTW, the tool only supports , as delimiter in its input.)
If I work with CSV files they are most often not comma-separated but semicolon-separated because of the numbers. An Excel installation localized for decimal comma would not read 'real' CSV files correct.
If csv-to-usv cannot cater for this type of CSV files, it would not be usable in a large part of the world.
From the top of my head, I can highly recommend SML
https://dev.stenway.com/SML/SimpleML.html
Recommend watching the, 'stop using CSV video' too
a <comma> b <comma> c <enter> d <comma> e <comma> f
why not using header-character:
<row><cell> a <cell> b <cell> c <row><cell> d <cell> e <cell> f
The primary USV library implementation uses Rust, which is notably good at UTF-8 and conversions between operating system string encodings (such as ASCII) and UTF-8,
ESV: eggplant-separated values. Because who is ever going to put AUBERGINE (U+1F346) into a dataset? It's the perfect record separator!
• Whatever is best for the data model and/or languages you use. JSON is a common modern choice, suitable for most things.
• If you want something more tabular, closer to CSV (which is a valid choice for bulk data), use strict RFC 4180 compliant data.
• If you want to specify your own binary super-compact data, use ASN.1. I am also given to understand that Protobuf is a popular modern choice.
If you aren’t in a position to choose your standards, just do whatever you need to do to parse whatever junk you are given, and emit as standards-compliant data as possible as output; again, RFC 4180 is a great way to standardize your own CSV output, as long as you stick to a subset which the receiving party can parse.
Nobody needs “USV”, and nobody should use it.
Some tools will randomly convert " to 'LEFT DOUBLE QUOTATION MARK' and 'RIGHT DOUBLE QUOTATION MARK' if they see UTF-8 flagging. Thus, the file is converted without your voluntary participation.
Different kind of character sets and character encodings will be good for different purposes. Unicode is "equally bad" for many uses.
I'll tell you why, it's pretty simple. The characters this... thing is stealing, exist to represent invisible control sequences. That is their use. The fact that they can be mentioned by direct input is inevitable, but not to be encouraged.
I will be greatly disappointed if this is accepted as a standard. The fact that a USV file looks like a rendered ASV file is a show stopping bug, an anti-feature, an insult to life itself. Kill it with fire.
I know we're not supposed to do low-brow dismissal here, but everything about this idea is wrong.
Rationale, solution, representation, everything.
If you don't have a human-readable file, might as well be compressible, queriable, and metadata-enabled I think.
Someone should write a family of filters of the form CSV2ASV, CSV2USV, CSV2JSON ,USV2XML , TOML2USV, USV2Cuneiform.......