Estimate the size of your English vocabulary (2014)
vocabulary.ugent.be
vocabulary.ugent.be
An example of something I said yes to on the second pass that I’d have said no to on the first: I guess superexcellent is a word, the meaning is very clear but I don’t know that I’ve ever seen it used before and would never consider using it myself.
Went back and also got a higher score by being less trigger happy with my right finger.
I got 87-0 being strict and 93-17 by including "maybes", and I'm not sure which is more reasonable.
Validity is not sanity; just conjulate the worbs, I guess
To be fair, I did also hit yes for words that I could work out the greek/latin roots for.
---
I know the intent of this test, but English has a special problem in that there is no real definition of what is and is not an English word. For example, there are a large number of words on this test that come directly from another language and are used as a term of art in a specific field.
See: mora, furuncle, cercal
Then you get transliterations like dacoit.
Are units of currency automatically English: piastre
Another genre is words that may be technically English, but would never be used - see delegable (delegatable would be used)
Then you have several alternate spellings from British English/American English/alternative transliterations.
As a native english speaker, answering for known or plausible words (with empty profile to not contaminate the test): 96% true positive, 23% false positive.
PS: Right-handed
Barras (French baras) n (Biography) Paul François Jean Nicolas, Vicomte de Barras. 1755–1829, French revolutionary: member of the Directory (1795–99)
Rigorous methodology this is not.
I have a feeling Scrabble players do well on this test.
>a phoneme which combines a plosive with an immediately following fricative or spirant sharing the same place of articulation
I wonder how tightly coupled this will be to education, gender, age, and especially handedness. Will righties rule and lefties drool, or will an unthinkable upset occur? I expect the U.K. to dominate because all the smart movies use U.K. actors, full of $2 words.
Also, in some cases I noticed that there was what appeared to be a real word, but it was slightly misspelled, so I'm sure that counted as an incorrect input.
Edit: oops replied to wrong parent
[1] https://en.oxforddictionaries.com/explore/how-many-words-are...
Meanwhile, a native speaker's everyday dictionary is typically on the order of ten to twenty thousand words. No way that their ‘reserve’ knowledge extends twenty times beyond this. I'd guess maybe several thousand infrequent and specialized words, with which the total knowledge would still clock in at low tens of thousands.
Meanwhile, the upper 10% of words includes highly specialized terms from niche sections of fields such as geology, chemistry and medicine, which arguably, are not words, even if you've seen them before. In some respects, a proper name isn't a word, and so, is a corporate invented product name really a word? Isn't that also a proper noun like "John" or "Texas"?
In this 10% grey area, I see a few words that I know to be words, but I cannot immediately tell you the definition, and if I were to try and use them in a sentence, I'd be able to apply the word class (noun, verb, adjective, adverb) within the structure of a sentence. But, I don't actually know the word. I only know that it's probably a word, but since I can't guess the definition or use it, I should dump it as an unknown, but more often than not, I'm getting away with some of these, mostly because after multiple passes, I can detect how the non-words are being generated, and the real words possess a more obvious validity of structure.
That said, I'm able to hover at or above 90% with occasional disembiggening penalties for over-cromulence.
Also a lot of the complex words are medicine/medical which makes sense since there are an awful lot of them but isn't necessraily representive of 73% (my score) of the words I will encounter outside of medical literature.
So basically I know 83% of the English words... that were asked in this test. If next year 1000 new words are introduced to the dictionary, that doesn't automatically mean I know 830 of them. In other words the minimum size of my vocabulary is 83 words. Nothing more. Nothing less.
They say this is fairly high for a native speaker.
I gave my mother tongue as German, so is this "fairly high" for a native German speaker?
It felt too easy.
This is pretty much what I expected, since as a non-native speaker; though even if you told me what stitchwort was in Finnish I'd still not know what it actually is besides a plant.
“Would you say you’ve read less than 10, less than 100, less than 1000, or all the books ever?”
I think it depends on the type of books one reads, not quantity.
A number of those were intuition, though.
(Scores for reference: 70, 77, 87, 84, 84, 84)
Which means nothing except that I've seen a lot of words and remember them.
I'm working on something similar for German at the moment.
However, it said that it is relatively high for a non-native speaker. How did they reach this conclusion?
So well above their estimates, but they don't extrapolate on the source of those numbers - they might be available somewhere at http://crr.ugent.be
My score:
You said yes to 63% of the existing words.
You said yes to 0% of the nonwords.
This gives you a corrected score of 63% - 0% = 63%.
I already knew I wasn't a wordsmith so that's not awful. About what I expected.
You said yes to 0% of the nonwords.
This gives you a corrected score of 79% - 0% = 79%.
This is a high level for a native speaker.
I'm impressed, I thought I would score way less since I'm not a native speaker.
My brain small. Sad!
I'd bet there is a bias from those reporting scores here -- a humblebrag, for science, of course!
You said yes to 76% of the existing words.
You said yes to 0% of the nonwords.
This gives you a corrected score of 76% - 0% = 76%.
Which I find oddly high.
"This is a high level for a native speaker."
I'm happy because I'm Swedish. I would like to see a histogram also by native language.
From the description I was hoping for a statistical estimate of how many English worsd I know...
So just multiply by 60k for how many you know. You've also got the info you need for stuff like stddev.
The posted test site does nothing of the sort.
Imagine the test only asked you a single 90% percentile word. vocabulary.ugent.be would tell you: You know 100% of the words. The process I described would tell you: You know 90% of the words.
Obviously both tests become more accurate as the number of words increases but the vocabulary.ugent.be test requires you to test all 60k words to get an accurate result. The second test will probably need less than 1000 words to be dead on.
http://testyourvocab.com/ seems to be a pretty good implementation of the approach I described.
> Obviously both tests become more accurate as the number of words increases but the vocabulary.ugent.be test requires you to test all 60k words to get an accurate result.
They could never get an accurate result for words known in a body wih this test - it needs a separate dataset of word frequency. That's not what they're testing.
Agreed that this seems like a measurement of something other than vocabulary.
I am curious about how common is having a ~75% score
83% (86% known - 3% false positive)