Estimate your English vocabulary size
testyourvocab.com
testyourvocab.com
This also shows how extremely time-consuming it is to learn a second language. I started in school, 10 years old, am moderately well educated (some college drop-out), and use English on a daily basis. I also watch most movies in English (very seldom for people in a German speaking country to do), read some English novels, and also most non-fiction books I have are in English. Internet use is nearly English only.
Still, I probably have the vocabulary of an average 12 year old native speaker. After 17 years of learning and using the language, and at least 10 years of that using it _daily_.
Ouch.
I very clearly remember my first time in Italy, when I saw a family at a restaurant and realized their 5/6 year old spoke much better Italian than I did. Kind of humbling.
I'll admit though that that point was rather tangential to the post I was replying to.
Sorry for the pun but since this is a thread about vocabulary and it's sunday when most of the regulars are away instead of procrastinating at work anyway I hope I may be forgiven.
Especially seeing the second and third page surprised me as I can't recall ever hearing most of these words. My girlfriend noted that many of them are older terms, but still I expected to have picked up some through movies or books.
On the other hand: studying and working in the US with just 11k words works quite well.
It's easy to notice the difference in vocabulary in practice. Native speaker's speech is much more varied.
Also, it's maybe good for native speakers to realize that someone who may seem a bit rude or dumb on the Internet may be a non-native speaker who has difficulty expressing him/herself.
Not really surprising since most of English scientific vocabulary comes from greek or latin roots and is more or less the same in other European languages.
(I scored 35,300; of the words I recognized only about six were ones I don't either read or use with some regularity. I read a fair amount of literature though.)
The words that are measure here are the words that you'd use in "general" conversation I guess.
15,500 words
Add another 2000 to cover technical/domain specific/slang words. And you should have no reason not to be able to communicate freely on day to day basis.
I got 26K (English is my 3rd language) and most of the 'strange' words are from reading too much science fiction as a kid.
+ on showing percentiles of non-native speakers tho
For non-native speakers, I would guess that there are quickly diminishing returns past about 10,000 words. Much more important is how well you're able to use those first 10,000 (or maybe even 5,000). It's one thing to recognize a word when reading it, and another to be able to use it in a natural way, choosing words with the right shade of meaning, and avoiding awkward constructions or unintended connotations.
There are plenty of words I know that I generally avoid in conversation for practical reasons. Entertainment reasons, though... I caught my wife teaching our not-yet-two-years-old daughter to use the word "etiolated" when referring to a green bean that was wilted and yellowish.
Not a modern corpus, but reasonable.
Strong's Exhaustive Concordance of the Bible (that would be the King James version) lists 8674 Hebrew root words for the pentateuch, and 5624 Greek root words for the New Testament.
Wikipedia suggests that full literacy in Chinese requires knowledge of about 4000 characters.
Most modern Mandarin words are bisyllabic, and represented by the combination of two characters, and their meaning does not necessarily derive in a straightforward way from the meanings of their constituent characters (for example, the word for "thing" is composed of the characters for "East" and "West" in conjunction)
Making up new words is an easy way to boost vocabulary.
I knew general meaning of many words I didn't check, but I couldn’t exactly define them. I know that I will understand the meaning from the context when I see them in some sentence.
I used the former because I thought it was what people would do (as it is easier to do; people will not what to spend 30 or more seconds on each word), and got 26600. That is below, but about average. I am not a native speaker, but would have expected to score higher in the general population.
I am fairly sure that this test does not reach the general population, though (by the way, it is nice to see that the test adapts to one's level. Try answering almost none of the words on page 1. I got 'my' score down to 28 by confessing to know only one of the words)
I'm also not a native English speaker, and one area where I know there's a huge difference between my Swedish vocabulary and my English is when it comes to names of plants and birds and spices and animals and trees and fish and rocks and flowers and vegetables and fruits. I know maybe thousands of names of such things in Swedish, but in English I know much fewer names. That's a few thousand words I lack and probably will never learn because it's so specialized.
(Btw, bananas are berries, but cows are not deers. :) )
Not sure if you're joking or not, but to be clear to non-native speakers: bananas are definitely not berries. Berries are smaller and rounder. For example: blueberries, raspberries, blackberries. I guess strawberries, too, but they're outliers.
Also the plural of deer is deer.
Berries have seeds on the inside; a strawberry's seeds are on the outside, and raspberries are little clusters they are called something else. It's a famous quiz question in the UK :)
Same with tomato - it's a fruit botanically, but in culinary terms it's a vegetable and definitely not a fruit. As long as you get your context right, you won't have problems communicating :)
"Based on previous research, Nation and Waring (1997) estimate that the receptive vocabulary size of a university-educated native English speaker is around 20,000 base words, while Goulden, Nation, and Read's (1990) intervention indicates that the receptive vocabulary size range of college-educated native English speakers is 13,200 - 20,700 base words (Goulden, Nation, & Read, 1990), with an average of 17,200 base words."
I'm a non-native speaker and got 18 700 words, which I found rather disappointing when I read that that was, according to the site, roughly equivalent to (or even below) the average 15 year old native speaker. Thinking a bit about it however I'm quite sure that that is nonsense - I am regularly asked by native speakers to review their English texts and scholarly articles and am relatively often commented on my broad vocabulary. When people review mine and I push them on why they suggested certain grammatical or stylistic changes, almost invariably it turns out that they are influenced by personal style preferences or local customs (as in, local to the area they grew up in). Now I'm not God's gift to the world and I'm sure that there is much to be improved on my English, but still I'm quite sure I can outperform the 5th percentile of the general population on English vocabulary knowledge. (I mean - that's people with an iq of 80 or less ffs, again not to say that I'm a genius but I find the proposition that I would score worse than most native speakers that qualify under many definitions as mentally retarded to be preposterous).
I'm highly skeptical of the site's claim that the 50% percentile knows 27k words. It seems to be from their own test takers, and I didn't find any references to other researcher's results, of which there are several.
On a side note: this makes me feel better about my 25k score ;-)
German, speaking English on a daily basis, had 8 years of English in school and 9 years of Latin, which probably helped me a bit. (Though I hope my teachers don't read this)
Does anyone know a similar test for grammatical proficiency?
that many
I only got an estimate of 11500 words, yet I can still correct you :). My native language is (Swiss-)German, spent a couple years in Montreal where you don't really have many chances to improve your english (at least not the vocabulary), otherwise only used the language in written form.
I've never lived in an English-speaking country, but English is so prevalent in Finland these days that I wouldn't be surprised if it gained some kind of official status within the next 50 years. Whether in formal meetings or informal bar encounters, people voluntarily switch to English if there's even one non-Finn present. In my field, this happens nearly every day.
I even use English to communicate with Swedes, even though Swedish is the second official language of Finland and I studied it for 6 years... There's no point in limping through the conversation with my childish Swedish, when it's 99% guaranteed that Swedes speak English.
My Swedish is atrocious. I can't recall a single time that I'd have had use for it outside school. In contrast, English is useful every day, though mostly in written form.
I wish I had been given language instruction of that caliber when I was in school.
I'd hoped to score at least median, disappointing. My exposure to English is pretty much limited to TV, 19th century literature, day-to-day conversations, and tech articles.
I scored about 33,000. I don't think the distributions are very accurate. I can't speak to where you should be, but I think you have an excellent score. I hope this point of comparison helps.
The difference between 10,000 and 30,000 is knowing the words "uxorcide" and "tricorn". Before this test, I've never seen the first one and the only time I've seen the second one is in reference to pirate hats.
Frankly, I'd rather have 10,000 word vocabulary in multiple languages than my current situation.
That and the amount of words I know RPGs in the quiz really surprised me. How often do you use a bludgeon, wouldn't you use a bat in stead?
Finally, there's a huge difference between active vocabulary --the words you actually use in speach or writing-- and passive vocabulary that is tested here. In my humble opinion, passive vocabulary should be tested in context of a sentence.
Most languages end up having about the same information bandwidth, but make different tradeoffs between the number of syllables per second and the entropy per syllable. Spanish and Japanese, for instance, tend to be spoken rapidly and have a more regular set of sounds, but languages like English or Chinese tend to be spoken more slowly and have more different sounds in each syllable. I'd imagine that low frequency languages tend to have large vocabularies that they use, but that's just a guess.
"To change your software so the user experience becomes overwhelmingly worse. e.g. Some critics say Microsoft committed uxorcide with the Office Ribbon"
My english comes mostly from watching movies and tv series long after I graduated, reading novels, articles, handbooks and later by chatting in english.
Due to movies and tv series, I adopted an us-american flavor of english.
Also keep in mind that this test didn't ask for domain specific vocabulary - if they would have made a section of computer related vocabulary, we probably would have all scored perfectly. :)
I got the feeling that that I could improve my score the most by thoroughly reading a 18th century novel.
I believe language knowledge is very influenced by the surroundings, the place where you live, the people you deal with.
Aware of the weakness, a few days ago I posted a call for suggestions on Google+; now the vocabulary test moved me to ask the same to HN: http://news.ycombinator.com/item?id=2772950
PS: I'm from Bari, Italy.
Beyond this, it's pretty much necessary to live abroad for an extended period of time, or be exceptionally good (driven) at languages and watch TONS of TV, use tons of online chat, etc.
My use of English is limited to daily-use for the most part. Apart from IT terminology, I rarely encounter field-specific words. Nor do I read literature.
I'm Russian by origin and have been living in Australia for the last 10 years. I think I read (and understand) a lot in English, but I still only got 13,000.
I got 5,490 words. May be my vocabulary is larger, but it's not broad: it targets specific fields.
I've been reading English since I was 13 or so - quite possibly the bulk of my reading has been in English (a quirk of mine).
I scored far lower than I would have expected, which although it hurts my ego a bit, I can easily dismiss because of the nature of this test.
In all the exams I had in my life, from 1st grade to BS in CS, I only had one multiple choice test.
E.g.: One of the words I encountered in the test was 'terpsichorean'. While I knew that Terpsichore is one of the Muses, I did not know which one and left the box unticked. Had there been multiple choices, I might have guessed the correct solution.
BTW, the word is 'deduce' ;)
To my defense, they are both derived from 'deducere', to lead away, and 'were not distinguished in sense until the mid 17th cent' according to my system’s dictionary
Google tells me it means dancing so I was way way off (maybe conflating turquoise and cerulean?).
But I have no way of knowing how many words that I feel comfortable defining are actually nowhere near correct, so to be any kind of accurate, they need to do some verification of correctness. All 'honesty' means is 'don't deliberately cheat' not 'don't be dumb'.
However, I think a multiple choice test could also inflate scores unless the definitions were very cunningly constructed.
I scored 75-80th percentile (32,800) which surprised me. It seems quite a lot of words, for one. For another I consider my vocab' to be very good and I don't think I'm being bigheaded in that. Ergo I expected to be ranked higher.
On the second page there was an entire column of words of which I recognised only three sufficiently to provide a guaranteed accurate definition. One of that column was terpischorean, another tatterdemalion.
Whilst looking up tatterdemalion I found little use of it after the 1930s except as a proper noun (a Marvel Comics character for example). What I did find however is that Google Books is useless for finding dates. One citation from an author Sir Edward Bulwer Lytton is given a date of 1999. That's a reprint date, the author died in the 19th century.
VOLUNTARY RESPONSE POLLS
As I commented previously when we had a poll on the ages of HNers, the data can't be relied on to make such an inference. That's because the data are not from a random sample of the relevant population. One professor of statistics, who is a co-author of a highly regarded AP statistics textbook, has tried to popularize the phrase that "voluntary response data are worthless" to go along with the phrase "correlation does not imply causation." Other statistics teachers are gradually picking up this phrase.
-----Original Message----- From: Paul Velleman [SMTPfv2@cornell.edu] Sent: Wednesday, January 14, 1998 5:10 PM To: apstat-l@etc.bc.ca; Kim Robinson Cc: mmbalach@mtu.edu Subject: Re: qualtiative study
Sorry Kim, but it just aint so. Voluntary response data are worthless. One excellent example is the books by Shere Hite. She collected many responses from biased lists with voluntary response and drew conclusions that are roundly contradicted by all responsible studies. She claimed to be doing only qualitative work, but what she got was just plain garbage. Another famous example is the Literary Digest "poll". All you learn from voluntary response is what is said by those who choose to respond. Unless the respondents are a substantially large fraction of the population, they are very likely to be a biased -- possibly a very biased -- subset. Anecdotes tell you nothing at all about the state of the world. They can't be "used only as a description" because they describe nothing but themselves.
http://mathforum.org/kb/thread.jspa?threadID=194473&tsta...
For more on the distinction between statistics and mathematics, see
http://statland.org/MAAFIXED.PDF
and
http://escholarship.org/uc/item/6hb3k0nz
I think Professor Velleman promotes "Voluntary response data are worthless" as a slogan for the same reason an earlier generation of statisticians taught their students the slogan "correlation does not imply causation." That's because common human cognitive errors run strongly in one direction on each issue, so the slogan has take the cognitive error head-on. Of course, a distinct pattern in voluntary responses tells us SOMETHING (maybe about what kind of people come forward to respond), just as a correlation tells us SOMETHING (maybe about a lurking variable correlated with both things we observe), but it doesn't tell us enough to warrant a firm conclusion about facts of the world. The Literary Digest poll
http://historymatters.gmu.edu/d/5168/
http://www.math.uah.edu/stat/data/LiteraryDigest.pdf
is a spectacular historical example of a voluntary response poll with a HUGE sample size and high response rate that didn't give a correct picture of reality at all.
When I have brought up this issue before, some other HNers have replied that there are some statistical tools for correcting for response-bias effects, IF one can obtain a simple random sample of the population of interest and evaluate what kinds of people respond. But we can't do that here on HN, nor can we for the online vocabulary estimation.
Another reply I frequently see when I bring up this issue is that the public relies on voluntary response data all the time to make conclusions about reality. To that I refer careful readers to what Professor Velleman is quoted as saying above (the general public often believes statements that are baloney) and to what Google's director of research, Peter Norvig, says about research conducted with better data,
http://norvig.com/experiment-design.html
that even good data (and Norvig would not generally characterize voluntary response data as good data) can lead to wrong conclusions if there isn't careful thinking behind a study design. Again, human beings have strong predilections to believe certain kinds of wrong data and wrong conclusions. We are not neutral evaluators of data and conclusions, but have predispositions (cognitive illusions) that lead to making mistakes without careful training and thought. Here, the conclusion "those other guys are cheating and that dragged down my vocabulary percentile score" is an example of a conclusion resulting from human predispositions.
Another frequently seen reply is that sometimes a "convenience sample" (this is a common term among statisticians for a sample that can't be counted on to be a random sample) of a population offers just that, convenience, and should not be rejected on that basis alone. But the most thoughtful version of that frequent reply I recently saw did correctly point out that if we know from the get-go that the sample was not done statistically correctly, then even if we are confident (enough) that HN participants are young or that their vocabularies are large, we wouldn't want to extrapolate from that to conclude that the users of any technology site are young, or that respondents to online surveys as a whole have large vocabularies.
On my part, I wildly guess that most HNers are younger than I am in part because this kind of poll recurs often on HN. I similarly guess that participants in online surveys of vocabulary size are likely to have larger vocabularies than average people in the general public because most people I meet find discussions of word meanings boring. But neither guess gives me a good quantitative basis for estimating how much users here differ from the general population.
I chose 'Canada' as my region since I'm from Montreal. My first language is English and I'm fluent French. I did the first half of elementary school in French. Firstly non-Quebec anglophones tend to have better grammar and larger vocabularies than anglophone Quebecers. Secondly it doesn't take into account that English can be a 3rd language . Most immigrants to Quebec are required by law to attend French language elementary and high schools (there are exceptions). Immigrant children who's first language isn't English or French (the majority) take on two new languages, English being their 3rd after French. English tends to be the social language for many.
Montreal has a strong tech industry employing bilingual/multilingual people many of which read HN and possibly took part in the survey. My gut feeling is that English speaking Quebecers are skewing the stats. More granular control over region will be useful; show some insight to this reality.
Note: I traveled through China and south east Asia last year and found the quality of English to be much better than I expected. Considering Indochina ruled by the French I didn't find a person who could speak it. To possibly classify any country as "non-English-speaking" is kind of silly. Every country is "other-English-speaking" but then again it's a subjective classification isn't it. Doesn't China have the largest English speaking population now...
Doesn't matter. Look at the density instead.
And when they see some egghead friend on Facebook has posted their vocabulary score and is challenging them to respond... they'll roll their eyes, and move on to their Farmville updates.
Don't get me wrong -- I love these things, and it came back with 37K for me -- but there's no way I'm posting that score, or even the link, to Facebook. I know how to maintain friendships, and saying "look how smart I am; I'm probably smarter than you" does not figure into it.
I think it would be a good idea to weight the survey with some test questions that ask if you know a definition to some of the less common words and then ask you to pick a correct definition from a list of 5 with 4 incorrect answers. At least this way they can approximate how much someone may exaggerate their knowledge.
However as someone who answered as honestly as I could (without spending the time to verify my definition of each word) it is cool to know what my personal vocabulary is.
Also this article does claim 24k-30k is the average for native english speakers: http://www.independent.co.uk/news/world/americas/english-lan...
It certainly seems unlikely that someone who can produce an excellent 500,000 word work of fiction (aside: thanks very much by the way,) in addition to reams of technical writing, has a vocabulary not in the 95th percentile of the population. OTOH, HP:MOR has fewer words in it than I expected, and even the upper bound of 14795 seems low. Maybe the working is wrong; it's shown below.
$ cat Harry\ Potter\ and\ the\ Methods\ of\ Rationality\ 1-72.txt |
tr -cs 'a-zA-Z' '[\n*]' | tr '[:upper:]' '[:lower:]' |
sort --uniq | cat - /usr/share/dict/words | sort |
awk '{count[$1]++; if (count[$1]==2) print}' | wc -l
12685
$ cat Harry\ Potter\ and\ the\ Methods\ of\ Rationality\ 1-72.txt |
tr -cs 'a-zA-Z' '[\n*]' | tr '[:upper:]' '[:lower:]' |
sort --uniq | wc -l
14795
[1] http://www.fanfiction.net/s/5782108/1/Harry_Potter_and_the_M... — Eliezer's amazing Harry Potter fanfiction in case you're missing out.Harry also presumably does some practical limiting of his vocabulary in conversation, because only shared vocabulary is useful if you're trying to actually communicate and don't want to stop to give definitions all the time.
It might be more interesting to compare the unique word count of HP:MoR against some of the "real" Harry Potter books, if you can get your hands on the text.
Total word count is 1,122,131 which is longer than HP:MoR by a factor of three. Plotting mean unique word count for the whole, halves and quarters of MoR gives a fit of uniques=168*length^0.3357, which makes sense given Zipf's law. That formula predicts about 18,050 words for a work of the same length as the original HP.
(Edit to add obvious test in the other direction.) The first 386,829 words of the original HP contain 12,255 unique words. The last 386,829 words contain 13,635 uniques. So, its comparable but perhaps slightly more varied (MoR had 12,685).
In light of those figures, is it possible Eliezer's vocabulary is less good than he thinks (Dunning-Kruger)? Especially as the Harry Potter book were written for children and presumably edited as such.
On the other hand, the fact that Eliezer seems to have used fewer words in his writing than you'd expect if his vocab was excellent doesn't mean that his known vocab is poor — he might just not use all the words he knows in writing.
Additionally, given the success of J K Rowling as an author, you might expect her vocabulary to be excellent, so it is conceivable that he's good and she's better.
† I have all the Harry Potter books on a shelf at home. Is torrenting the pdfs at work so I can word count them infringing copyright? I could have done it manually, it just would have taken longer.
Excellent work, by the way; thanks for the analysis.
And yet I look at the vocabulary used in the comment posts of those claiming high scores, and wonder how they ever scored so highly (accepting that comments are not neccessarily reflective of ones general writing or vocabulary). I believe that as this test is so open to cheating, that using responses to it as the corpus for determining median scores renders the entire exercise completely meaningless.
P.S., 36,700 .. I took this before it got a lot of general circulation, and my standings have improved considerably. I suspect this makes me more worthy of oxygen.
That said, I'm currently reading A Dance with Dragons and there are tons of words in this series (A Song of Ice and Fire) that I'm not familiar with. Most of the ones I missed are words I recognize from this series, although since I'm not 100% sure of them, so I left them unchecked.
Bollocks.
Maybe I fell behind my peers by only reading comp sci, math, business, and communications related stuff for 4 years. Guess I've got some catching up to do.
I scored a little over 40k on this, and did well on other verbal tests for the general population when I was still taking tests.
I attribute much of my facility to 1)reading fantasy and 2)looking up words I don't know. Since I loathe interrupting the flow of a story, I read with a pencil and make a list of words to batch learning later.
Best,
I also thought I had an unfair advantage early on because I knew a lot of words from playing Blizzard games all my life. Sorely disappointed :(.
Using big words make a good communicator not. The point of using a language is to communicate. Using esoteric words only a fraction of the population knows defeats this purpose.
If anything, given the same amount of practice but with a smaller vocabulary will make you a better communicator, IMHO.
edit: forgot to mention, while I am a native English speaker, both my parents tend to speak Polish most of the time, while I always reply in English. I have to wonder how much this had an affect on me.
Atypical language acquisition (e.g., as a second language, or through a non-standard channel like technology or fantasy literature) disrupts this extrapolation step. For instance, a German programmer that knows the word "polymorphic" via OOP is less likely to know similarly frequent and difficult but programming-unrelated words than British or American peers. So adding, say, 100 to the total would be utterly unfounded. Same thing for a science-fiction nerd: Acquaintance with obscure words from one domain doesn't extrapolate to other domains.
Unless they somehow control for domain specificity and atypical acquisition, let's not get too frustrated. (Disclaimer: Not a native speaker -- result around median.)
Your basic vocabulary will be the same as most other speakers because we have structured learning in our schools.
But the less-well-known words are generally learned in context... So if you haven't happened across them in a book, you probably don't know them... But there are probably many others that you do, instead. This test could easily hit a bunch you don't know and totally miss all the ones you do.
This test made me aware of that fact, regardless of it's scientific accuracy.
That being said, cognates between languages can give an artificially inflated score. I intentionally avoided any words with the same roots in English and Portuguese (my second language), which should hopefully also be true for most Romance languages. This way, you shouldn't be able to "guess" meanings you've never actually learned. However, it wouldn't surprise me if German speakers, for example, were able to "guess" an additional number of words correctly.
Also, the 60-80 words are only what you are tested on -- the word selection on pages 2-4 changes depending on your answers on the first page, so it attempts to narrow down your vocab knowledge "at the margin".
http://www.reddit.com/r/linguistics/comments/edlnv/reddit_i_...
This is a bewildering instruction to me, since I've learned most of my vocabulary through reading, and rarely look up the definition of a word, instead learning its meaning through repeated exposures to its use in context.
Take for example "garron" -- like many of us I've been reading GRRM lately, so I've seen the word used some 78 times in the past few months, and I'm sure I've encountered the word a few dozen times before. I know it's a slightly undesirable horse of some kind. Likely this means it's a gelding or a small pony-ish horse. Do I need to have looked up and remembered the three specific submeanings of the word, or that it's a specific breed of horse from Galloway to be able to say I am "exactly sure of" the word?
I don't think that's how language works, but it's how this test seems to want it to work. My score of 35,300 is suspect on multiple levels.
Similarly, your point has some validity. Most of the time, we don't acquire language through reading the dictionary; we acquire it primarily through context in the course of reading or conversation. This is why, when pressed to define words, we'll often reach for a string of synonyms, or else provide usage in sample sentences. I bet few people here, let alone anywhere, could render dictionary-acceptable definitions of 99% of the words they know.
(On the flipside, this is also why we forget most of the words we crammed in preparation for the SAT back in the day; we learned them completely out of context and in an artificial way).
I never really considered Hacker News to be full of overconfident people. If anything, to me being part of HN is a humbling experience. It reminds me that there are so many people out there that are smarter or more experienced than me, and I've seen other people say similar things.
I think one contributing factor is the structure of the questions. If you ask me if I KNOW at least one definition for the word mawkish for example, I'll choose no I don't know a definition of it. I do however have a good enough feel for the word, that I can almost guarantee I can get a SAT style analogy with it correct, or if you gave me a multiple choice selection of definitions I can probably pick out the right one. I don't consider either of those skills equivalent to actually knowing a definition of a word.
During a discussion of 'whether open source contributions were being overly important to job seekers' a while back, a surprising number of HN commenters automatically put themselves in the role of employer, peering dubiously over their glasses at me. Many of these commenters were hilariously under-qualified to be taking on that kind of role with respect to anyone.
I don't think SAT verbal is nearly as intensely focused on obscure vocab. I didn't do it, but did do a GRE verbal back in 1994 to get into grad school. For what it's worth, my score on this and my GRE verbal were quite consistent.
You inferred this based on user profiles?
Heh.
anyway, what i'm saying is that i suspect sat verbal is testing something quite different - isn't it much more aimed at ability to use the language rather than whether you can recognise obscure words?
[what surprised me the most was how graded it was - on both tests there was a pretty clear cut-off point where the words became unknown. i didn't expect it to be that ordered.]
Maybe you're thinking of the Dunning-Kruger effect? (http://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect)
End result: substantial inflation of scores relative to what you'd see if you gave the same survey to the general population (and assuming everyone is relatively honest).
But I created the site, less interested in absolute vocabulary size numbers (these vary widely depending upon methodology), and more in relative changes among age groups, SAT scores, etc. And hopefully, cheating would not be correlated to any of those...
But at the end of the day, this is not a controlled, scientific survey. It is a voluntary quiz, though I am doing my best to control for other factors.
Going by magazines, common forwards and so on it seems the average man loves a good self-quiz.
And in the back of my mind, I thought that the bulk of quiz magazines were sold to women, no? (Though I'm probably being too pedantic...)
ObScore: 36000, 32yo male US-english native speaker. There were a few more I'd seen but wasn't really sure enough to define. I maxed out the SAT verbal section (and got 1 wrong on math, which was enough to drop to 780/800) back in the 1990s.
The thing that always surprises me with this sort of test is how many words I half-know: I recognise the word, I have a (correct) general sense of what it means, but either I couldn't articulate an accurate definition or the definition I would give is an uncommon usage and I hadn't come across the more common meaning(s).
Following the rules strictly (so I didn't tick the half-known words), I came in slightly below average for my age, assuming the trend continues beyond the 32 years that their table currently gives.
For comparison, if I also included the words I half-knew, I gained nearly 4,000 new words and went up about 20%.
What IS interesting tho, from a language learners perspective, is the vocab size estimation. A metric a lot of us use as a rough benchmark of vocab needed for fluency in a foreign language is 10,000words. Comparing this with what an educated adult native speaker knows in their own language (using my own truthful score of 24k) is pretty interesting.
Would love to have something like this to quickly gauge my vocab in other languages!
I got 12900 (English is my second language) and it seems to be about accurate. I also feel that I'm fluent in English, and it confirms that 10000 words is sufficient for fluency.
The test is also missing professional lingo: where are such words as SQL, lisp, ai, startup, PG, HN and other that we all know so well?
Of course, I very quickly rationalized away the counter-reaction and took the test anyway, and then considered sharing it with my friends. What drives this?
People who lead revolutions tend not to be very humble, from what I can glean from history. So what you're suggesting might work.
I'm not sure that the people suggesting that the failure to correlate with the SAT adds much; I don't think the SAT really goes all-out of the more flowery bits of archaic vocabulary in the way that this test did.
My 3rd grade son got 10200, and enjoyed discussing the words he didn't get. I think every 3rd grader should know "mawkish". :-)
Memorizing trivia words is just something that has never interested me. Instead I keep a thesaurus and dictionary handy at all times :)
37,100 I'm ashamed I didn't do better. I'm considerably older than the average HN reader. I did degrees in 4 different subjects (mind you, I was classed officially as retarded at my high school - in the same classes as the arsonists).
So no-one should feel the score is that important. I'm a very mediocre programmer. I'd much rather halve my vocab score to double my maths ability.
I had a phase around 7th through 10th grade where I thought learning lots of vocabulary would make me smarter, especially words others didn't know well. (And so I'd use them in English essays for Extra Points since your grades are often determined by how little sense you make, because if the reader doesn't understand it obviously it's too smart for them!) I also had a general grammar nazi-ism.
Anyway, I think this exchange kind of tipped me over the edge to stop caring. (Of course that's led to forgetting a lot.)
William Faulkner, on Ernest Hemingway: "He has never been known to use a word that might send a reader to the dictionary."
Hemingway: "Poor Faulkner. Does he really think big emotions come from big words? He thinks I don't know the ten-dollar words. I know them all right. But there are older and simpler and better words, and those are the ones I use."
Of course, having some background in French and Latin probably helps for inferring a few words.
Coincidentally, I just happened to come across this quote in Peter Norvig's "Paradigms of Artificial Intelligence Programming".
I also just realized that taking the test primed me to write in a way more intelligently sounding manner than I usually do. Not an LOL in sight!
Also some definitions you can guess at - do these count as part of your vocabulary or not? E.g. Clerisy, can be easily deduced to be the class of clerics, but not sure if I'd be able to say if it was a real word or not.
I think this test is not very telling for NNSs as it doesn't consider specialist vocabulary which many of us have a lot of of, because of how we /really/ learned the language.
When I left school my English was very average. When I started communicating with email with people from all over the world, but mostly the US, it improved a lot in 1-2 years. When I first went to a congress in the US in my med 20's I was blown away by my aptitude to communicate in that language.
But these were all people from my field. What I'm saying is that the distribution of words pertaining certain subjects in my vocabulary is severely skewed by the field I work in -- visual effects (and IT). I believe this goes for many NNSs.
You clearly have a very good vocabulary IMO. However, if I may, I think it should be "pertaining to certain subjects". It can be used without the "to" but sounds a bit conceited to this native speakers ear.
For example see http://www.thefreelibrary.com/_/search/Search.aspx?By=0&.... All the literature examples written out use "pertaining to" in some form: "pertaining thereto", "pertaining only to", "to them pertaining", et cetera.
(My result:) http://testyourvocab.com/?r=35795
Also worth seeing: http://testyourvocab.com/hard.php
How about you? Were there words you could check because you encountered it often in programming context?
Thanks for all the participation, and comments -- I didn't submit this myself, so thanks, mike_esspe.
I've responded to a few points down below; there's a lot more info on the details of the test at:
http://testyourvocab.com/faq.php http://testyourvocab.com/details.php
I was intrigued by your estimation procedure and was thinking of a way of playing around with it:
(1) Create a random bit vector of size 45,000. (2) Pick 40 positions in the vector at random and use those positions to estimate the proportion of ones in the original bit vector (so far this is easy). (3) Select an additional 120 positions and use the more sophisticated procedure to refine the estimate from part (2).
I was wondering if you have code or pseudo-code for how you implemented part(3) specifically how you choose the 120 words for extra testing. It seems there are a lot of different natural ways you can do it. Did you write an academic paper about this?
Thanks!
I had 9 years of English in school. Because of my hobbies I read a lot of stuff in English. I also watched many TV shows and spent about 6 month living in Australia.
Yet I feel insecure even typing this. Knowing lots of words is one thing. But what makes it hard are all the subtleties you have to take care of when building sentences. I also think that I get grammar wrong most of the time.
Another thing is that my sentences are almost always way too long.
As someone else pointed out before: I don't want my former English teacher to read this, either.
Edit: And, also to add, I followed 2 criteria for whether I know the word or not.
1. What's the absolute definition?
2. And can I find the equivalent or meaning of it in my native language? (which is Tamil, an Indian language, if anyone cares.)
I'm Canadian born and raised, with English as my first language. Honestly, I'm surprised to be told I'm that far below the median and average.
Edit: I just had a look at the median word count for adults who took the survey. It's around 27,000. I wonder whether that's true or not.. it seems to me that I'm lacking.
Now with a low score of 17500 I wonder, if it isn't enough to completely endulge oneself in the language, what is?
Of course, watching the Simpsons all day won't teach me some of the rarely used words. But there must be some stepping stones. I still haven't read Wuthering Heights because I don't want to have a dictionary lying around just to understand the story. And looking up something, reading on and forgetting it at the end of the day is quite common for me.
Also I'm sure that 15 year old americans haven't read that many novels, still their vocabulary is supposed to be larger than most of the well read non-native speakers around here.
Sorry, couldn't resist. For a minute I thought I'd missed out on an alternate spelling. Then I realised that was unpossible ;0)>
Honestly I think the evaluation method is terrible. My collection of sci-fi/fantasy books alone probably contain over 100,000 headwords. A single biology text book might be as many as they claim the median person knows. Avid WoW players would similarly destroy the curve (if the test included the kinds of words they'd know instead of archaic religious words),
42,500 (http://testyourvocab.com/?r=37216)
Apparently the OED has 7 times more words I don't know. That's offal...
The thing I find a bit funny is that of all the words I didn't check I've seen almost all of them in books and articles. When I see them in a sentence and in context I do understand them fine but I can't give a definition for them.
I wonder if this is common when reading another language? It might be a better idea to look up the words in a dictionary when seeing them but I just can't be bothered, after seeing them in context a few times I can usually get a feel for their meanings. There are a few exceptions to be sure, adjectives are particularly bad at this.
Maybe working on that English Minor is panning out...
Quote:
Too limited. Words that are specifically American or British (in meaning or spelling), or slang, or scientific/medical, or anything labeled archaic, or anything else that isn't part of broad, general English. Also, no animals or ingredients, which depend too much on where you live.
So scientific words aren't general English? This suggests that the authors don't think anyone ever talks about science?
It isn't surprising that there isn't any technical vocabulary. Most technical vocabulary falls into one of the following categories: a) Acronyms b) Overloading of existing non-technical words c) Names and other proper nouns d) Phrases longer than one word
There's actually an argument for excluding highly specific vocabulary (some corpuses explicitly exclude textbooks for this reason) because knowledge of them doesn't correlate as well with overall vocabulary.
What difference does it make? The site doesn't say what it means in everyday life. I'm guessing if you exclude high achieving SAT vocab nerds, it finds the difference between people who care about the meaning of each word and people who will guess through context because they have no patience for a dictionary. Or people who don't read fancy texts, like the Scarlett Letter for example, after failing to read that I stopped reading books.
I wonder if there is some good web site that helps you learning new words. An iPhone application will also work for me, but I need one that is able to also tell me the sound of the word. I searched a bit in the past without good results.
When I was young, I thought that if I wanted to be a writer I should have a huge vocabulary... but now, when choosing words/synonyms I dismiss most options because they're much too obscure.
This is what I was thinking. I scored 34K, and rarely encounter a word that I don't understand in regular speech or reading. I also know several thousand jargon words, none of which were on that test. I know what I need to know. Memorizing another 16K words to reach Shakespeare's magical 50K (and feel good about myself) would be a waste of precious mental resources.
I'm certain that if I took this same test using a base of 'novel in fulltext' vs. 'list of all unique words in the novel', my recog would be FAR better on the novel.
I am not a native speaker and I was happy to get more than >10,000, really.
BTW: a nice thing about online dictionaries is they have sound files for pronunciation. e.g. http://www.thefreedictionary.com/malapropism With a plugin, you can highlight and right-click to open in another tab.
32800
I would love to be able to compare my score to what it would have been before moving to a non-English country and learning/speaking a new language. I definitely feel that a large part of my memory is now dedicated to Japanese and not English...
To put aside the ego matters, I'm curious if there are any interesting correlations for writings published in magazines. For instance, between the estimated vocabulary size and the average price for ads (I bet that there is a huge correlation.)
According to http://math.ucdenver.edu/~wbriggs/qr/shakespeare.html, Shakespeare used ~32K words in all of his works.
Quote:
No cognates or false-friends with Portuguese. This probably knocks out at least half the dictionary, since Romance languages have plenty in common with English. False friends need to be avoided as well, since a Brazilian beginner will see "pretend" and assume he knows it means pretender, which actually means "intend." Interestingly, the no-Portuguese rule leaves the test with a strongly pronounced short Anglo-Saxon flavor.
Funambulism is my new favourite word.
By rights a funambulist should be a John Cleese impersonator ...
Anyway, my score: http://testyourvocab.com/?r=36192
Not sure I buy the results though. I would think that the rate of increase would start to decrease quite significantly after high school/college but it appears to stay pretty much linear throughout the data.
Add Education (None,HighS,BS,MS,PhD) as a research statistic and compare it to the national average. I would bet that your sample is greater than 2 Standard Deviations from the mean.
"The companion Brazilian project can be found at howmanywords.com.br"
Which right now just says "Calcule o tamanho do seu vocabulário em inglês / Coming soon"
Edit: Oh. The Brazilian site is just a Portuguese version of the instructions for the English site? :(
Clearly my American public school education has served me well. >_<
I got a high score and went to public schools in the US, and didn't finish college either.
It's likely that education improves one's vocabulary, however there are a lot of variables there. People who go to good schools likely come from families where learning is important in any case, are wealthier, etc...
And at the upper end, it's more about personal predilections than anything else. Are rare words like little gems that you save and collect?
So, I beg to differ. Education can definitely help. Bravo to you, though, for building your vocabulary on your own.
I checked all the boxes on the first page but one, "vibrissae" (roughly, whiskers), and saw even more boxes on the second page and groaned. So I punted, and went back to the first page, reloaded it, and checked only one box: vibrissae.
After finding 17 words on the second page, I left them all blank, following the same methodology. My vocabulary size was estimated to be 20 words.
From this, I deduced that my total vocabulary size was all the words ever known to any English speaker anywhere, anytime - minus 20.
This made me very pleased with myself, even though I knew my assumptions were pretty terrible.
(Interesting that the spell check in my browser doesn't recognize vibrissae.)
I knew lampoon had something to do with criticism, so I checked the box, but I had no idea that the definition specified a public context. Does that mean I didn't know the word?
I suspect a problem with the test is that it's easy to know enough to figure out the gist of a word's definition without having any knowledge of the specificity of the definition, if that makes any sense.
Though I suppose, also -- IQ doesn't say much at all about what your vocabulary will be. If you don't read much, or mostly read popular literature, it doesn't matter how clever your are; your vocabulary won't be any bigger than the words you've encountered enough to form (or look up) a definition.
Though it's also worth noting that beyond a certain point, building a broader vocabulary isn't very useful; when you communicate, you generally need to limit yourself to vocabulary your audience knows. Likewise, when you're reading, it's pretty common that the authors using particularly obscure vocabulary are using it to muddy their meaning, not clarify it.
I am non-native speaker. I guess that's a excuse for such poor score. Need to improve :(
27 year old native speaker in the US. I talk good English. I even properize capital nouns.
You need a test where the definition is part of the test. Otherwise anyone can just check boxes.
I could use some help boosting my vocabulary...