More importantly, CSVs are handled natively by statistical programs such as Excel and R (you can change the expected delimiter while importing, but in my experience that leads to unexpected results)
in excel or R? Because I've never encountered problems with tsv files with excel
That was embarrassing.
The CSV format is so messy that even big corporate implementations will have subtle bugs like that. Do not trust it.
except it is. (g)awk is more often part of the pipe(d)line than other programming langueage(s) interpreter(s).
comma vs. tab: It doesn't matter, in the end, you have to use a character to separate fields, and if you are going to use a "normal" character, there are chances that it might appear in your text, so you have to either escape that character, or quote your text, which csv does in a fairly standard way, so comma vs tab doesn't really matter.
I always found it best practice to quote text. Makes the debate a moot point.
It's a shame that the ASCII file, group, record, and unit separators that have been around since forever did not catch on.
(BTW, if inputtability mattered at all, Unicode wouldn't exist. ;-)
And unicode is of course easy to use with "less exotic" languages like Norwegian that can't be represented in basic 127-bit ascii -- via a simple keyboard map.
Tabs might be somewhat less frequent than commas in data, but you're still gonna have just as bad a time when they do turn up. And don't forget quoting and escaping. Without proper parsing you can neither read CSV nor TSV safely. My go to solution is pandas.read_table()
And with developers (especially Python devs) setting their text editors to translate tabs-into-spaces, copy-pasting TSVs results in data corruption. Or even opening a TSV in your text editor with default settings.
Use ASCII's `RECORD`, `GROUP` and `UNIT` separators, they were literally designed for it.
Record separator => C-v C-^
Unit separator => C-v C-_
in emacs I think you can use C-q in place of C-vTabs have a similar problem: They look like spaces in a simple text editor like notepad.
But, if you can't input the character...you can't input the character.
There's been one time where a client provided a couple hundred gigabytes of data where all the reasonable characters one might use were used in the fields themselves, including their own separator character(the comma). So I made up some rules for guessing which commas were delimiters and which were part of a field value, and stored it using the ASCII control character as a delimiter. The only other use of this character at my company are in jokes, like "We should ask the client to provide or ingest files separated using the ASCII control codes" which always gets a laugh.
It's trivial to change all those commas to something else before making the CSV file and importing into the database.
They can then be converted back to commas by processing the query result.
This type of conversion before storage is routinely done with other characters, such as newlines, quotes, etc. For example, look at the JSON for the 10mHNComments data dump.
Personally, from a readability standpoint, tabs drive me nuts. I am glad TSV is not the default.
I'm also a huge fan of Ruby's `csv` package in the standard library, particularly http://ruby-doc.org/stdlib-1.9.2/libdoc/csv/rdoc/CSV.html#me....
All that only works for properly quoted CSV though... for companies who can't generate quoted csv files pipes or tabs are the unfortunate way to go.
#!/usr/bin/env python
import csv
import sys
writer = csv.writer(sys.stdout, delimiter='\t')
writer.writerows(csv.reader(sys.stdin))Abbreviated Unix paths? ~/.bashrc
Approximated values? ~30º
Problem solved.
"Richard, Martin", "23 NS, North Coast, NY"
Comma frequently occurs inside texts, and then awk fails.That's apostrophe, quote, comma, space, end quote, apostrophe.
Will not work were numeric fields are not quoted.
But nice solution, nevertheless.
Name, Age, Address
"James aka ""Jim"", ""licensed"" attorney", 42, "New York"
That's three values: James aka "Jim", "licensed" attorney
42
New York
And there are other possible irregularities: zero or N spaces after the comma separators; unquoted values when they're not needed; backslash-escaped special characters; escaping newlines.