How to Analyze Every Reddit Submission and Comment, in Seconds, for Free
minimaxir.com
minimaxir.com
More importantly, CSVs are handled natively by statistical programs such as Excel and R (you can change the expected delimiter while importing, but in my experience that leads to unexpected results)
That was embarrassing.
The CSV format is so messy that even big corporate implementations will have subtle bugs like that. Do not trust it.
except it is. (g)awk is more often part of the pipe(d)line than other programming langueage(s) interpreter(s).
in excel or R? Because I've never encountered problems with tsv files with excel
comma vs. tab: It doesn't matter, in the end, you have to use a character to separate fields, and if you are going to use a "normal" character, there are chances that it might appear in your text, so you have to either escape that character, or quote your text, which csv does in a fairly standard way, so comma vs tab doesn't really matter.
It's a shame that the ASCII file, group, record, and unit separators that have been around since forever did not catch on.
(BTW, if inputtability mattered at all, Unicode wouldn't exist. ;-)
And unicode is of course easy to use with "less exotic" languages like Norwegian that can't be represented in basic 127-bit ascii -- via a simple keyboard map.
I always found it best practice to quote text. Makes the debate a moot point.
Tabs might be somewhat less frequent than commas in data, but you're still gonna have just as bad a time when they do turn up. And don't forget quoting and escaping. Without proper parsing you can neither read CSV nor TSV safely. My go to solution is pandas.read_table()
And with developers (especially Python devs) setting their text editors to translate tabs-into-spaces, copy-pasting TSVs results in data corruption. Or even opening a TSV in your text editor with default settings.
Use ASCII's `RECORD`, `GROUP` and `UNIT` separators, they were literally designed for it.
Record separator => C-v C-^
Unit separator => C-v C-_
in emacs I think you can use C-q in place of C-vTabs have a similar problem: They look like spaces in a simple text editor like notepad.
But, if you can't input the character...you can't input the character.
There's been one time where a client provided a couple hundred gigabytes of data where all the reasonable characters one might use were used in the fields themselves, including their own separator character(the comma). So I made up some rules for guessing which commas were delimiters and which were part of a field value, and stored it using the ASCII control character as a delimiter. The only other use of this character at my company are in jokes, like "We should ask the client to provide or ingest files separated using the ASCII control codes" which always gets a laugh.
It's trivial to change all those commas to something else before making the CSV file and importing into the database.
They can then be converted back to commas by processing the query result.
This type of conversion before storage is routinely done with other characters, such as newlines, quotes, etc. For example, look at the JSON for the 10mHNComments data dump.
Personally, from a readability standpoint, tabs drive me nuts. I am glad TSV is not the default.
I'm also a huge fan of Ruby's `csv` package in the standard library, particularly http://ruby-doc.org/stdlib-1.9.2/libdoc/csv/rdoc/CSV.html#me....
All that only works for properly quoted CSV though... for companies who can't generate quoted csv files pipes or tabs are the unfortunate way to go.
#!/usr/bin/env python
import csv
import sys
writer = csv.writer(sys.stdout, delimiter='\t')
writer.writerows(csv.reader(sys.stdin))Abbreviated Unix paths? ~/.bashrc
Approximated values? ~30º
Problem solved.
"Richard, Martin", "23 NS, North Coast, NY"
Comma frequently occurs inside texts, and then awk fails.That's apostrophe, quote, comma, space, end quote, apostrophe.
Name, Age, Address
"James aka ""Jim"", ""licensed"" attorney", 42, "New York"
That's three values: James aka "Jim", "licensed" attorney
42
New York
And there are other possible irregularities: zero or N spaces after the comma separators; unquoted values when they're not needed; backslash-escaped special characters; escaping newlines.Will not work were numeric fields are not quoted.
But nice solution, nevertheless.
Platform is OSX, title font is Source Sans Pro, axis numeric fonts are Open Sans Condensed Bold (the same fonts used on the web page itself; synergy!)
http://stackoverflow.com/questions/1898101/qplot-and-anti-al...
it should now be possible to get AA on both Windows and Linux via the Cairo-library.
I'll be playing a bit with R notebooks for my stats intro class, if I find the time I'll try to see if a) I get AA out of the box, or b) I can get AA easily by setting some parameters.
[ed: And if I do, I suppose pr to https://github.com/minimaxir/ggplot-tutorial will be in order]
In a nutshell, the idea was to compare lists of emails and test to see how many accounts in the given list were also in the database. We'd typically get 10k item lists to test against the database.
Initially, I was creating a temporary table and doing a unique intersection on it. That took about 45 minutes on a modest machine. Next, I hacked up a bloom filter search (since all we cared about was whether or not the items in the sample were in the database); that ran in a little under a second for a set of 10k, but got a little unwieldy with sets approaching 1MM+.
We talked about it, and decided to just use Elasticsearch, and to just split the database into a bunch of buckets. Search time went up a small bit for small samples, but went waaay down for larger samples.
Also of note, going from spinning platters to SSD did make a pretty huge improvement in things when we were using a relational database. Now that it's all basically in memory, it doesn't matter so much.
DDR4 also theoretically supports 512 GB sticks [3] Can't wait.
1) http://www.crucial.com/usa/en/support-memory-speeds-compatab...
Case in point, many comments in this page point at the "best time to post" chart. But what if that time varies by subreddit?
The answer at https://www.reddit.com/r/bigquery/comments/3neghj/qotd_best_...
"median": "2389", "seventy_fifth": "3543", "ninetieth": "5191"
How did the Internet page renders get slower since 2013?
2.2s median, 3.3s for 75th percentile, and 4.7s for 90th percentile in httparchive:runs.2013_06_01_pages]
I quite like that visualization and would like to use it for something else.
"2.5 seconds, 2.39 GB processed, and a few R tricks results in this:" http://minimaxir.com/img/reddit-bigquery/reddit-bigquery-2.p...