Perhaps this would make an interesting personal project. Are you aware of any hurdles, missing key features, etc. that previous attempts at creating such a format have run into (other than adoption, obviously)?
I'm a Notepad++ person. When I needed to mock-up data typing the characters was easy-- just ALT and the ASCII code on the numeric pad. It took a bit to memorize the codes I needed to use. Their visual representation is just inverse text and initials.
Edit: this is what I have so far: https://github.com/tmccombs/ssv
In time, editors and file browsers should come to render separators visually in a logical way.
Nice job. I hope you come back to finish the project eventually.
(That said every text editor since ever should have had a "table mode" that uses the ASCII field/record seperators (or whatever you choose), I was always confused why this isn't common. Maybe vim and emacs do?)
I absolutely HATE this Parcquage.
(I don't think everyone has moved to it. I had never heard of it myself.)
You won't remember Parquet in 15 years, but you will have CSV files in 50 years.
You're probably right about CSV but probably not parquet. Parquet is already 11 years old, there are vast data warehouses that store parquet, it's first class in the spark ecosystem, and a key component of iceberg. Crucially, formats like parquet are "good enough" for a use case that doesn't appear to be going away. There is a high probability in my estimation that enough places are still using them in 15 years to be memorable even if it isn't as common or as visible.
Considering it also separates records with newlines, they really should have replaced newlines with "\n" and require escaping "\" with "\\".
The advantage of ASV is not that you can't have invalid or insecure data, it's that valid data will almost never contain ASCII control characters in the record fields themselves. Commas, quotation marks, and backslashes, meanwhile, are everywhere.
Feel free to correct me, but I figure that as long as data can be from 0x00 to 0xFF per byte, no format that uses characters in that range will ever be safe. I’m not a big C developer but I figure the null terminated strings have the same limitation.
But if its something entered by keyboard you should be ok to use control codes.
Personally, I find tab and return to be fine for text driven stuff. Shows up in an editor just like intented.
(Because CSV is a terrible data exchange format in terms of information per byte. But that makes sense, because it's an intentionally human readable data exchange format, not a machine format)
Hence https://github.com/SixArm/usv/tree/main/doc/faq#why-choose-u...