CSV 1.1 – CSV Evolved (for Humans)
csv11.github.io
csv11.github.io
* From the documentation "No quotes needed for values ... Use dual quotes ... Use triple quotes"
* You haven't solved the problem of describing what the columns are. In fact, its worse because you are encouraging people to put units in the field
* Spaces matter! What if the data is literally " Word " vs "Word". This format makes them both the same.
* For some reason, the header row is removed. The only thing in a CSV that provides structure is gone. In the first "improved" example, I have no way of knowing what the columns meaning without them.
The only good improvement in this is the "#" convention as a comment. If we could get every CSV parser in future to agree (we already can't get them to agree now) to ignore lines that start with "#" then CSV's would be vastly improved.
If leading and/or trailing whitespace are significant, you can quote the value. Problem solved.
Having said that, almost always when you see data persisted with leading or trailing whitespace it should have been trimmed away before storage.
It's not clear to me that it's possible to convert this to a standard CSV; it can't guess the header and guessing whether a column contains units (consistent units?) is asking for trouble.
More pragmatically I can't share these with other people because they're almost like a csv, but incompatible.
It's great to try to make CSV more readable (I like the leading whitespace, but it'd be finicky to maintain with simple text editors), but not worth the technical risk of format confusion.
Original: I'm not even sure I agree that 'comments' are good, particularly if they're going to be intentionally ignored by GUI s/s apps.
If you have meaningful information, put it in the furthest right-column, so it appears in spreadsheet editors on the right of the data.
If it's not meaningful to the person who will open the file, why include it?
There are no problems :-). You're making up things / problems.
Nowhere says it that the header row is removed. You CANNOT auto-detect a header, you have to tell your parser (see the csvreader library as an example - https://github.com/csv11/csvreader - from my humble self.) if you have a header or not (it's optional).
If Spaces matter! but them in quotes. By default you don't need quotes (and, thus, discouraged).
> You haven't solved the problem of describing > what the columns are
That is solved / done by a schema with a (tabular) datapackage, see https://github.com/csv11/csvpack as a real-world example how that works in practice or use csvrecord, see https://github.com/csv11/csvrecord
So, just like a regular CSV?
1. Add a version identifier / content-type on the first line!
2. Create a formal grammar for this CSV format
3. Specify preferred character-encoding
4. Provide some tooling (validation, CSV 1.1 => HTML, CSV => Excel)
5. Add the option to specify column type (string, int, date)
6. Specify ISO-8601 as the preferred date format
7. Allow 'reheading' the columns in the file itself. This is useful in streaming data.
8. Specify the format of the newlines.
"CSV on the Web: A Primer" http://www.w3.org/TR/tabular-data-primer/
"Model for Tabular Data and Metadata on the Web" http://www.w3.org/TR/tabular-data-model/
"Generating JSON from Tabular Data on the Web" (csv2json) http://www.w3.org/TR/csv2json/
"Generating RDF from Tabular Data on the Web" (csv2rdf) http://www.w3.org/TR/csv2rdf/
...
N. Allow authors to (1) specify how many header rows are metadata and (2) what each row is. For example: 7 metadata header rows: {column label, property URI [path], datatype URI, unit URI, accuracy, precision, significant figures}
With URIs, we can merge, join, and concatenate data (when e.g. study control URIs for e.g. single/double/triple blinding/masking indicate that the https://schema.org/Dataset meets meta-analysis inclusion criteria).
"#LinkedReproducibility"; "#LinkedMetaAnalyses"
CSV's problems are the nature of a very flexible convention. It's so simple that everyone writes their own generators and parsers that are slightly different. That's what happens when you use a convention like csv.
Revving the spec won't help anything... Because csv is the kind of convention where no one reads the spec anyway!
ID,Name,Capital,Area,Tags
bc,British Columbia,Victoria,922509,en|western canada
vs ###########################
# Oh, Canada! 10 provinces and 3 territories
#
# see en.wikipedia.org/wiki/Provinces_and_territories_of_Canada
#
# note: key is two-letter canadian postal code
#
# for regions tags see
# en.wikipedia.org/wiki/List_of_regions_of_Canada
bc, British Columbia, Victoria, 922 509 km², en|western canada
In 1.1 I don't see what is what - what is `bc`? What is `en|western canada`? In vanilla CSV you can clearly see what is what because it shows you on top.Why not just write a CSV formatter?
Also, seems like specification-by-example, unless I’m missing something.
Removing column headings would be a killer. Although efforts in the past to add type information are misguided, by that point you may as well just use XML.
If you are worried about people in the year 3000 understanding your data, add a .txt file alongside the CSV explaining the fields.
I disagree. CSV is horribly underspecified and many parsers have conflicting ideas on how things should work. I've hit areas where, for example, an API was generating a CSV that another API could read, but NOT if it was opened and then re-saved by Excel first (just re-saving, not editing), because Excel was making some trivial change in the format that the consuming parser couldn't handle. Then there's issues with encoding, localisation, and escaping (see, eg, https://chriswarrick.com/blog/2017/04/07/csv-is-not-a-standa... although that's just scratching the surface).
That being said, this proposal seems worse in every way that existing CSVs, but I don't think the issue is "there's a well specified CSV standard with tons of parsers that won't change to support this new one", it's more that there is no standard, all the parsers are incompatible, and no meaningful improvements are possible. :(
14"
exports as , '"14"""',
It feels downright silly to me. 14"
would be ,"14""",
And I verified that all four of Excel's CSV varieties write this, but this is also what RFC CSV writers would emit. Not saying you can't pick up oddities elsewhere, but as stated, it's not that bad. (Not that I condone use of CSV… it's an awful format.)1. Look at how everybody is extending CSV in the real world
2. Make a new format which does those things, as closely to all the existing implementations as is reasonable
3. Write a fully-specified new syntax for it, including a complete parser (state machine including error cases)
4. Make a complete set of tests, which make it very obvious when you don't pass, for shaming everyone into compliance
5. Get buy-in from a major standards organization, and all of the major corporate players
I don't see any improvements/replacements succeeding if they don't hit all of these points. "CSV 1.1" here hits #1 and #2 only.
And while CSV is primarily used for data transfer, Excel could hypothetically be a great debug-viewer if it didn't do brain dead stuff like unrecoverably corrupt UPCs and other long numbers by default.
There's an Excel Uservoice about this very issue from 2015:
https://excel.uservoice.com/forums/304921-excel-for-windows-...
They're "looking into it." Just as they've been looking into it for the last twenty years. Any. Day. Now.
It's so bad I check any csv file I get if it's been opened in Excel, and if I find any evidence that it has I reject it and make whoever I got it from give me the original source, or transform it with literally any other tool.
Good point. That's what (tabular) data packages are for, see https://github.com/csv11/csvpack for some real-world examples.
If you want it to be easily parsable by a human then there are hundreds of applications designed for this. The most common being spreadsheets like Excel.
And what happens if you have a list of domains, but one contains a really long url (300 characters wide?). It will mess up the columns for every single row above and below it as well.
Also adding comments will completely break any existing software (I think most programs can handle some more whitespace, but you couldn't import the final example into Excel, I don't think (untested))
Another hack to improve CSV workflows is OCHA's HXL[2] that is used by humanitarian organizations. Basically adding a row of hashtags in addition to column names, which is surprisingly useful considering the ease of adding these to a file.
[1] https://frictionlessdata.io/docs/tabular-data-package/ [2] http://hxlstandard.org/
[1] https://frictionlessdata.io/specs/table-schema/ [2] https://github.com/frictionlessdata/goodtables-py
https://github.com/frictionlessdata/specs/issues/537
Feel free to add your use cases to that issue.
See https://github.com/csv11/csvpack for real-world data package examples.
By the way, the tabular data csv dialect specification - is a great start/initiative (mostly a 1:1 copy from the python parser :-), really - would need an update, for more options, to reflect the reality of the csv formats out there. The big insight and breakthrough - csv is NOT one spec or format - but various flavors / formats / dialect - let the computer (that is, csvreader library) handle it.
PS: csv-next is NOT an informal specification - these are notes (collection of ideas).
PPS: How do I know? I'm the author of the website - I should know ;-).
I guess you might not have understood my complaint (my bad), so let me give some examples how the specification should look like.
An informal specification looks like this [1]. You should give examples and rules enough to use your format and sufficient to implement most of the things. You should give definitions for keywords that are specific to your specification (but you can omit common definitions). You should give some (but probably not all) ideas where the specification can go wrong: Unicode whitespaces, Byte Order Mark, platform-specific newlines, escape sequences, duplicate keys in the front matter (well, actually there is no provision how to put the front matter at all), numeric overflows, and so on.
A formal specification looks like this [2]. In addition to what an informal specification provides, you should give lots of examples and formalized rules (most frequently ABNF [3]) to implement all the things. There should be clear and reasonable error handling policies. The wording of the specification should be clear, unambiguous and preferably standardized (there are specific meanings to "MUST", "SHOULD" etc. [4]). The specification should be honest about its pros and cons. You should be explicit about the flexibility of the format: you should give a list of what can be extended or modified later and what can't.
I strongly suggest you to put an informal specification at the least, and to prepare for the eventual development of a formal specification by pondering about missing pieces (it is a daunting task, I know).
[1] https://github.com/toml-lang/toml/blob/bb47759841ac368d86eb7...
[2] https://tools.ietf.org/html/rfc7049
[3] https://en.wikipedia.org/wiki/Augmented_Backus%E2%80%93Naur_...
That's right, but your "specification" was not even enough for writing a CSV 1.1 file. For example I have mentioned a front matter issue---there is no example using it, and I'm deeply confused how the front matter looks like (or even where it is). I'm seeing numeric units there, but I don't know how they are handled (or, say, what happens if different units like m² are there) from your page.
I'm not saying I'm a good specification writer (and I'm not even a native speaker of English), but I at least try. I have once written an informal specification [1] which should be almost enough to write valid files and start implementing parsers, without being too complicated. You can avoid a complicated specification without being too vague.
> For inspiration, here's an example from your humble self - https://github.com/csv11/csvreader
That's frankly much better (in terms of explicitness) than the front page, why not linking to it? :-)
But I believe that it is still not sufficient. Rather, I'm now unsure about the goal of your project: is it a set of Ruby libraries with fancy, modern but incompatible (by itself) CSV extensions? Or is it hopefully going to be a universally used format? If your intention was the former, then your front page should have been clear about it---it is not about a format but about a library. And if you are going to have a format, then my suggestions hold.
[1] https://github.com/lifthrasiir/cson (with a bit of formal materials, just for the clarity)
It's NOT about one universal csv format and the ultimate specification to settle the matter until the end of history etc.
Thanks for the detailed suggestions. I see your points. I appreciate your helpfulness.
On the whole, some ideas to make CSV files human-friendly, but neglecting backward compatibility (comments, use of spaces...) and introducing major syntactic and semantic cans of worms (multiline values, named fields, defaults...). I think CSV files should evolve towards tighter constraints instead.
Is it slightly more verbose? Sure, but why does it matter? Having to quote all strings is really not a big deal.
Now… most diff tools also have options to ignore whitespace changes, for cases like this.
To get comma separated CSVs to show properly in Excel we have to mess around with OS language settings. CSV as a format should have died years ago, it's a shame so many apps/services only export CSV files. Many developers (mainly US/UK based) are probably not aware of how much of a headache they inflict on people in other countries by using CSV files.
Even if a `spec` is created which clearly defines these cases and standardizes how they should be handled, does it qualify to be called CSV 1.1 (or 2.0? really, as it probably won't play nice with existing CSV 1.0 implementations). It is almost as-good-as creating a new format altogether. And there are many to complete with.
I also wonder if it REALLY solves the problems it aims to solve (even if the tech specs were in place). CSV in any form is not human-friendly. This is especially true when you have wide columns (say >30). Even if the records were spaced-out in a human-friendly way in STATE-1, editing records where values are of highly varying width (within a single column) will soon mess the justification when you get to STATE-2. If you skip the requirement for `fixed-number-of-fields` and `key:value` style named values; that messes readability (by humans) even more!
I've had good success with `reading` CSVs using the `csvtk` (https://github.com/shenwei356/csvtk). It provides excellent support for pretty-printing CSVs, filtering select fields etc. `writing` is still a pain, but I still feel spreadsheets are the way to go. They have been around for decades. Its sad if the formatting by specific tools is screwing CSVs, but, solving the composition/modification requirement purely by way of formatted-plain-text is a really tall ask.
https://en.wikipedia.org/wiki/C0_and_C1_control_codes#Field_...
[1] https://github.com/openfootball [2] https://github.com/openmundi [3] https://github.com/openbeer
I work in the public sector, we use a lot of CVS, even to non tech savvy people.
I’ve never heard the use case for this project.
Now that might not matter, but you’re breaking CVS to fulfill a nonexistent use case.
Just for the record to quote from the memo:
This memo provides information for the internet community. IT DOES NOT SPECIFY AN INTERNET STANDARD OF ANY KIND. IT DOES NOT SPECIFY AN INTERNET STANDARD OF ANY KIND. IT DOES NOT SPECIFY AN INTERNET STANDARD OF ANY KIND.
Just repeating it three times in case you missed it, see https://www.ietf.org/rfc/rfc4180.txt
How do you define headers? There's no header for the bottom csv version so I have to manually type headers for my data frame?
I got a dataset with 87 obs and ~8700 columns here. I'm not going to manually name those columns. What's the solution?
See the CSVJSON format - http://csvjson.org - Love it. I will add a new variant called CSV <3 JSON shortly to the CSV 1.1 repo and csvreader etc. too.
Everything old is new again.
Yes, ideally the tab is perfect (no escape rules needed, etc). In practice you cannot tell if you're human if a tab is a space or a space is a tab and than you will get into trouble reading / writing your data etc. Anyways, both are great and needed and use tab2csv or csv2tab to convert or pipe (when using command line tools) :-)