VOLUNTARY RESPONSE POLLS
As I commented previously when we had a poll on the ages of HNers, the data can't be relied on to make such an inference. That's because the data are not from a random sample of the relevant population. One professor of statistics, who is a co-author of a highly regarded AP statistics textbook, has tried to popularize the phrase that "voluntary response data are worthless" to go along with the phrase "correlation does not imply causation." Other statistics teachers are gradually picking up this phrase.
-----Original Message----- From: Paul Velleman [SMTPfv2@cornell.edu] Sent: Wednesday, January 14, 1998 5:10 PM To: apstat-l@etc.bc.ca; Kim Robinson Cc: mmbalach@mtu.edu Subject: Re: qualtiative study
Sorry Kim, but it just aint so. Voluntary response data are worthless. One excellent example is the books by Shere Hite. She collected many responses from biased lists with voluntary response and drew conclusions that are roundly contradicted by all responsible studies. She claimed to be doing only qualitative work, but what she got was just plain garbage. Another famous example is the Literary Digest "poll". All you learn from voluntary response is what is said by those who choose to respond. Unless the respondents are a substantially large fraction of the population, they are very likely to be a biased -- possibly a very biased -- subset. Anecdotes tell you nothing at all about the state of the world. They can't be "used only as a description" because they describe nothing but themselves.
http://mathforum.org/kb/thread.jspa?threadID=194473&tsta...
For more on the distinction between statistics and mathematics, see
http://statland.org/MAAFIXED.PDF
and
http://escholarship.org/uc/item/6hb3k0nz
I think Professor Velleman promotes "Voluntary response data are worthless" as a slogan for the same reason an earlier generation of statisticians taught their students the slogan "correlation does not imply causation." That's because common human cognitive errors run strongly in one direction on each issue, so the slogan has take the cognitive error head-on. Of course, a distinct pattern in voluntary responses tells us SOMETHING (maybe about what kind of people come forward to respond), just as a correlation tells us SOMETHING (maybe about a lurking variable correlated with both things we observe), but it doesn't tell us enough to warrant a firm conclusion about facts of the world. The Literary Digest poll
http://historymatters.gmu.edu/d/5168/
http://www.math.uah.edu/stat/data/LiteraryDigest.pdf
is a spectacular historical example of a voluntary response poll with a HUGE sample size and high response rate that didn't give a correct picture of reality at all.
When I have brought up this issue before, some other HNers have replied that there are some statistical tools for correcting for response-bias effects, IF one can obtain a simple random sample of the population of interest and evaluate what kinds of people respond. But we can't do that here on HN, nor can we for the online vocabulary estimation.
Another reply I frequently see when I bring up this issue is that the public relies on voluntary response data all the time to make conclusions about reality. To that I refer careful readers to what Professor Velleman is quoted as saying above (the general public often believes statements that are baloney) and to what Google's director of research, Peter Norvig, says about research conducted with better data,
http://norvig.com/experiment-design.html
that even good data (and Norvig would not generally characterize voluntary response data as good data) can lead to wrong conclusions if there isn't careful thinking behind a study design. Again, human beings have strong predilections to believe certain kinds of wrong data and wrong conclusions. We are not neutral evaluators of data and conclusions, but have predispositions (cognitive illusions) that lead to making mistakes without careful training and thought. Here, the conclusion "those other guys are cheating and that dragged down my vocabulary percentile score" is an example of a conclusion resulting from human predispositions.
Another frequently seen reply is that sometimes a "convenience sample" (this is a common term among statisticians for a sample that can't be counted on to be a random sample) of a population offers just that, convenience, and should not be rejected on that basis alone. But the most thoughtful version of that frequent reply I recently saw did correctly point out that if we know from the get-go that the sample was not done statistically correctly, then even if we are confident (enough) that HN participants are young or that their vocabularies are large, we wouldn't want to extrapolate from that to conclude that the users of any technology site are young, or that respondents to online surveys as a whole have large vocabularies.
On my part, I wildly guess that most HNers are younger than I am in part because this kind of poll recurs often on HN. I similarly guess that participants in online surveys of vocabulary size are likely to have larger vocabularies than average people in the general public because most people I meet find discussions of word meanings boring. But neither guess gives me a good quantitative basis for estimating how much users here differ from the general population.
I chose 'Canada' as my region since I'm from Montreal. My first language is English and I'm fluent French. I did the first half of elementary school in French. Firstly non-Quebec anglophones tend to have better grammar and larger vocabularies than anglophone Quebecers. Secondly it doesn't take into account that English can be a 3rd language . Most immigrants to Quebec are required by law to attend French language elementary and high schools (there are exceptions). Immigrant children who's first language isn't English or French (the majority) take on two new languages, English being their 3rd after French. English tends to be the social language for many.
Montreal has a strong tech industry employing bilingual/multilingual people many of which read HN and possibly took part in the survey. My gut feeling is that English speaking Quebecers are skewing the stats. More granular control over region will be useful; show some insight to this reality.
Note: I traveled through China and south east Asia last year and found the quality of English to be much better than I expected. Considering Indochina ruled by the French I didn't find a person who could speak it. To possibly classify any country as "non-English-speaking" is kind of silly. Every country is "other-English-speaking" but then again it's a subjective classification isn't it. Doesn't China have the largest English speaking population now...
Doesn't matter. Look at the density instead.
It certainly seems unlikely that someone who can produce an excellent 500,000 word work of fiction (aside: thanks very much by the way,) in addition to reams of technical writing, has a vocabulary not in the 95th percentile of the population. OTOH, HP:MOR has fewer words in it than I expected, and even the upper bound of 14795 seems low. Maybe the working is wrong; it's shown below.
$ cat Harry\ Potter\ and\ the\ Methods\ of\ Rationality\ 1-72.txt |
tr -cs 'a-zA-Z' '[\n*]' | tr '[:upper:]' '[:lower:]' |
sort --uniq | cat - /usr/share/dict/words | sort |
awk '{count[$1]++; if (count[$1]==2) print}' | wc -l
12685
$ cat Harry\ Potter\ and\ the\ Methods\ of\ Rationality\ 1-72.txt |
tr -cs 'a-zA-Z' '[\n*]' | tr '[:upper:]' '[:lower:]' |
sort --uniq | wc -l
14795
[1] http://www.fanfiction.net/s/5782108/1/Harry_Potter_and_the_M... — Eliezer's amazing Harry Potter fanfiction in case you're missing out.Harry also presumably does some practical limiting of his vocabulary in conversation, because only shared vocabulary is useful if you're trying to actually communicate and don't want to stop to give definitions all the time.
It might be more interesting to compare the unique word count of HP:MoR against some of the "real" Harry Potter books, if you can get your hands on the text.
Total word count is 1,122,131 which is longer than HP:MoR by a factor of three. Plotting mean unique word count for the whole, halves and quarters of MoR gives a fit of uniques=168*length^0.3357, which makes sense given Zipf's law. That formula predicts about 18,050 words for a work of the same length as the original HP.
(Edit to add obvious test in the other direction.) The first 386,829 words of the original HP contain 12,255 unique words. The last 386,829 words contain 13,635 uniques. So, its comparable but perhaps slightly more varied (MoR had 12,685).
In light of those figures, is it possible Eliezer's vocabulary is less good than he thinks (Dunning-Kruger)? Especially as the Harry Potter book were written for children and presumably edited as such.
On the other hand, the fact that Eliezer seems to have used fewer words in his writing than you'd expect if his vocab was excellent doesn't mean that his known vocab is poor — he might just not use all the words he knows in writing.
Additionally, given the success of J K Rowling as an author, you might expect her vocabulary to be excellent, so it is conceivable that he's good and she's better.
† I have all the Harry Potter books on a shelf at home. Is torrenting the pdfs at work so I can word count them infringing copyright? I could have done it manually, it just would have taken longer.
Excellent work, by the way; thanks for the analysis.
And when they see some egghead friend on Facebook has posted their vocabulary score and is challenging them to respond... they'll roll their eyes, and move on to their Farmville updates.
Don't get me wrong -- I love these things, and it came back with 37K for me -- but there's no way I'm posting that score, or even the link, to Facebook. I know how to maintain friendships, and saying "look how smart I am; I'm probably smarter than you" does not figure into it.
I think it would be a good idea to weight the survey with some test questions that ask if you know a definition to some of the less common words and then ask you to pick a correct definition from a list of 5 with 4 incorrect answers. At least this way they can approximate how much someone may exaggerate their knowledge.
However as someone who answered as honestly as I could (without spending the time to verify my definition of each word) it is cool to know what my personal vocabulary is.
Also this article does claim 24k-30k is the average for native english speakers: http://www.independent.co.uk/news/world/americas/english-lan...
I scored far lower than I would have expected, which although it hurts my ego a bit, I can easily dismiss because of the nature of this test.
In all the exams I had in my life, from 1st grade to BS in CS, I only had one multiple choice test.
E.g.: One of the words I encountered in the test was 'terpsichorean'. While I knew that Terpsichore is one of the Muses, I did not know which one and left the box unticked. Had there been multiple choices, I might have guessed the correct solution.
BTW, the word is 'deduce' ;)
To my defense, they are both derived from 'deducere', to lead away, and 'were not distinguished in sense until the mid 17th cent' according to my system’s dictionary
Google tells me it means dancing so I was way way off (maybe conflating turquoise and cerulean?).
But I have no way of knowing how many words that I feel comfortable defining are actually nowhere near correct, so to be any kind of accurate, they need to do some verification of correctness. All 'honesty' means is 'don't deliberately cheat' not 'don't be dumb'.
However, I think a multiple choice test could also inflate scores unless the definitions were very cunningly constructed.
I scored 75-80th percentile (32,800) which surprised me. It seems quite a lot of words, for one. For another I consider my vocab' to be very good and I don't think I'm being bigheaded in that. Ergo I expected to be ranked higher.
On the second page there was an entire column of words of which I recognised only three sufficiently to provide a guaranteed accurate definition. One of that column was terpischorean, another tatterdemalion.
Whilst looking up tatterdemalion I found little use of it after the 1930s except as a proper noun (a Marvel Comics character for example). What I did find however is that Google Books is useless for finding dates. One citation from an author Sir Edward Bulwer Lytton is given a date of 1999. That's a reprint date, the author died in the 19th century.
And yet I look at the vocabulary used in the comment posts of those claiming high scores, and wonder how they ever scored so highly (accepting that comments are not neccessarily reflective of ones general writing or vocabulary). I believe that as this test is so open to cheating, that using responses to it as the corpus for determining median scores renders the entire exercise completely meaningless.