N-gram Analysis of the New York Times Weddings Section
news.rapgenius.com
news.rapgenius.com
You can clearly see the recent tech boom by searching "Google," "Facebook," "Twitter," and "Apple" http://www.weddingcrunchers.com/?q=facebook%2C%20google%2C%2....
The key takeaway here is Google:
Google has raced ahead of establishment NY law firms: http://www.weddingcrunchers.com/?q=wachtell%2C%20cravath%2C%....
Google has also recently overtaken top investment banks: http://www.weddingcrunchers.com/?q=goldman%20sachs%2C%20morg...
Ditto for consulting: http://www.weddingcrunchers.com/?q=mckinsey%2C%20boston%20co....
When do you think Google will start hosting a debutante ball in Chelsea?
Either way -- stuff like this is a delight to read.
Fred Wilson initially told them "I think lyrics is a very crowded space and almost entirely reliant on Google for traffic" and they admitted "our pitch back then was a bit too lyrics-focused.."
http://news.rapgenius.com/Lemon-how-rap-genius-raised-s18m-i...
You can find those quotes in an annotation in the above link (which makes me realize the problem with annotations is you can't ctrl-f them).
You can get that by clicking "share" in the annotation footer
That said, it might be nice to have an optional plugin to easily discover annotations and add your own to pages as you browse.
Given that that's the defense people seem to proclaim every time someone mentions that disabling JavaScript is now buried as an arcane config flag in Firefox.
They say: "just use a plug-in", "just use an add-on"...
What they mean to say is: "just don't disable javascript at all... ever."
They are the same people.
let me fix that for you
"This makes it possible to put numbers on our preconceived notions and play around with them."
It may be entertaining, but rigorous? I don't think so.
One simple one that comes to mind here is that you need to analyze to what extent changes over the period of the data set are caused by underlying societal changes, versus changes in the NYT itself; the end result will be a mixture of those two changes, some of which may be magnifying and others offsetting. The 1980 NYT and the 2013 NYT are not the same newspaper, not edited by the same people, not sold to the same readership demographics, and not soliciting the same advertisers, so it's somewhat questionable to treat it as a stable proxy for a social group.
Another common pitfall is language change screwing up all kinds of measures (since n-gram models just work on word counts). For example, if two words are used roughly interchangeably in 1980, but by 1990 one of them has fallen out of usage, and been replaced wholly by the other one, searches for just the one word will look like the word's on an upwards trend, but it would be misleading to infer an increase in the underlying concept over the period. Of course, you can account for this by merging words into equivalence classes (most analyses will do basic stemming and merging of alternate spellings), but you have to be very careful to get all the equivalence classes (which is not a well-defined notion). Just a list of the top words in a year will tend to be some mixture of 1) top concepts; and 2) concepts expressed using only a small number of wording variations, so their count doesn't get diluted.
The post just does a good job of hiding it by smoothing the plots. Compare an unsmoothed plot: http://www.weddingcrunchers.com/?q=Democrat%20%2B%20Democrat... with the smoothed plot in the article: http://s3.amazonaws.com/rapgenius/HhvuocYI3raAnYpWPE4HaeCh9a...
While the % of republicans does appear to fall, the % of democrats in the last year is lower than in the first year, the opposite of the conclusion they want you to draw!
It wouldn't be a bad idea to factor in the number of Democrats vs Republicans holding offices in the area around NYC during that time, either. I know NY state leans Democratic, and Democrats do well in city-level elections. Holding an actual office would probably make you more likely to mention your party.
>What does the y-axis mean exactly? The y-axis represents the frequency of each phrase, as a percentage of all phrases that contain the same number of words. For example, if you search for from New York, the graph shows the number of times those words appear in exact order, divided by the total number of 3 word phrases in all of the articles
I think doing it at a per-article level makes more sense for an analysis like this, but 0.02% is actually pretty significant when n is on the order of millions.
Thanks for the clarification.
That's actually a slightly dubious analysis. The question you need to ask is 0.02% of what? In this case, I would take a guess it means 0.02% of all the words analyzed. As a very simple example, imagine analyzing all the letter in a book. If English were perfectly balanced, we expect to see all 26 letters at 1/26 or 0.038%, so seeing the letter 'e' appear at say, 1.0% (or even 0.1%) would be a notable statistical result.
I understand that, like college admissions, you can hire a wedding planner or consultant who can considerably raise the chances of your wedding being listed.
The NYT obits are another interesting read.
The factors that enter into getting in (from my observation strictly) are a combination of things like:
- parents who live in ny metro
- the parties getting married living in ny metro
- having gone to school in ny metro
- parents or parties getting married working in ny metro
- what the parents do for a living
- any lineage "grandparent governor of NY"
- what the parties getting married do for a living
- school attended as far as perceived impressiveness
- whether an impressive job or title of any of the parties mentioned.
..and so on. That's off the top.
For example, "physician" and "went to school in NY" is probably almost assured to get the announcement printed.
"father a mechanic, mother a homemaker, inlaws are nobodies, parties are cashiers who work at walmart, no college, live in jersey city"[1] and so on either don't get in, don't care to get in, or don't have the drive to even submit a form to get in.
[1] Unless of course one of the parties is related to a famous former politician or some other mitigating factor.
I've read the wedding announcements on and off since about 1991, but much less these days, because I only get the online edition now.
Getting a write-up in the times is one affirmation that you're a power couple in a certain northeastern old-school way, or a human interest angle.
A day or two before the wedding, he told us that he wrote both a long form piece and a shorter, more typical piece. He wasn't sure which would get published, but he was obviously pushing for the longer piece to get in. It did.
Oddly enough, our write-up isn't included in the Rap Genius dataset. Maybe it's too recent or the longer write-ups aren't included.
As far as being in the dataset, certainly they aren't analyzing every flavour-style thing article, just "announcements"? Just like a 2-page life-in-review article on someone famous when they die doesn't really go in the obit section, does it?
Together my wife and I check quite a few boxes that the NY Time typically looks for. We aren't famous or all that noteworthy, but we do have an interesting story of how we met. That was the main focus of the article.
http://ldc.upenn.edu/Catalog/CatalogEntry.jsp?catalogId=LDC2...
This is how we do it (examples below are not weddings, but random topics):
http://blogdotitrendcorporationdotcom.files.wordpress.com/20...
http://blog.itrendcorporation.com/2013/04/10/social-media-on...
http://www.grantland.com/story/_/id/6769919/matrimonial-mone...