The most frequent 777 characters give 90% coverage of Kanji in the wild
japanesecomplete.com
japanesecomplete.com
98% is closer to what you need in order to read a text and have an idea of what's going on. See this article:
https://www.sinosplice.com/life/archives/2016/08/25/what-80-...
I would guess (based on not-too-rigorous anecdotal experience) that the optimal level of word familiarity for most rapidly improving listening comprehension when read-to 1:1 is closer to 90% than 98% (but with the caveat that books should be re-read several times). The big difference is that having someone to talk to makes it very quick and easy to ask what words mean and discuss other aspects of the story.
When read to for an average of 1+ hour per day, young children improve at listening comprehension very quickly.
Any references on this? Interesting.
You could start with a google scholar search like https://scholar.google.com/scholar?q=storybook+OR+%22picture...
Was wondering if that was research based or just personal experience. The hypothesis makes sense, but I'd love to see it substantiated.
I suspect that the role of pictures (or any other context) is to help correct your mental state of the narrative world so that the sentences you don’t understand don’t present an impassable barrier, preventing you from getting to the later sentences that will help you learn.
Sure, reading isolated sentences with only 90% coverage is really hard sometimes. But usually you're reading whole blocks of text, so you have context. Also, maybe you don't know some kanji but it has similar radicals to other, so you can make educated guesses (only sometimes of course).
I mean it's just like English! There are radicals/etymological clues, there are context clues, there's so much that you can work off of to figure things out. And exercising "figuring things out with partial information" is a valuable skill that can compensate for a lot of missing kanji knowledge.
An aside: I've heard that when learning a language, you should be able to understand roughly 80% of what you're reading/listening to, because then you'll have enough of a base to not feel overwhelmed, while still having new stuff to absorb.
I have no clue what a "N2" is and I'm not entirely sure what you meant by "Kanji/readings". There are only six other sentences of context and they didn't really help to understand the first one.
From a quick Google I gathered that there are 5 levels of proficiency you can certify with. They are N1 to N5, N1 is the highest rating.
Kanji: The Chinese characters in question
N2: The second highest level of the Japanese Language Proficiency Test (JLPT). N1 is by many seen as a very high bar and many Japanese employers demand it if you'll be using Japanese as your main working language. N2 is generally attained after 1-2 years of Japanese studies.
Ideographic, as in trying to portray ideas graphically, not idiomatic, as in using expressions like a native speaker
https://www.japantimes.co.jp/life/2018/10/29/language/ghost-...
From my own experience, this would only be true in the most favorable of the situations. 2 years studying fulltime while living in Japan sounds about right. 1 year studying as a hobby few hours a week, no way.
For context, it took me 3 years of university (outside Japan) to pass N2.
(I'm just amused of all the % thrown around I this thread)
This is a dangerous and corrosive mindset that did a lot of harm to me and other people I know. It took me maybe until 25 to actually start to accept my lack of knowledge and incompetence and start to do something with it. I sometimes wonder what my life would be like if I had the emotional maturity to challenge myself at a young age instead of pursuing destructive perfectionism.
And allow me to claim that you've actually proved your parent's point! I have no idea what an "N2" is either but from the context was able to accurately guess that it's some kind of proficiency test.
Maybe whether one can make a reasonable guess depends on more than the context - perhaps background, experience, and point of view. I know I've missed points that are obvious to others because of our differing perspectives.
If I hear something that amounts to "you idiot, get the schmonck out of my way", I'll understand even if I have no idea what "schmonck" really is.
Funny as this is how I learned (and have similarly described) to read language X and it felt ___almost___ like zero resistance. I learned first just to identify each character by name, then somewhat how they were pronounced, then vowels, then how there were pronounced, then special rules for X, then special rules for Y.
All of the learning material clump them together from start, a few consonants, a few vowels, some of their properties, and rules on how to combine them. All at once, this made me ignore writing for years until I tried it my way which I think worked well (although it took some time, still).
We were a class of 6-8 students, in pairs making our guesses of what the words that we didn't "understand" meant. We all got pretty advanced ideas and all got to similar conclusions. The teacher would tell us afterwards how those were made up words, and we were all in awe at how much we still understood and agreed upon the meaning of the words.
https://en.wikipedia.org/wiki/Jabberwocky
Sidequestion, is this poem taught in schools in the US by any chance?
I also read/type Japanese (my writing out of lack of continuing to write on a daily basis has gone to the dogs). With four years in Japanese universities and 20 years living in Japan I had the opportunity to experience what it is to get reasonably proficient in a second language.
The synaptic connections you can make by understanding 'most' of a compound kanji (multiple characters joined together) along with parts of the kanji makeup for a specific kanji you do not know puts you in an incredible position to figure out what the word is likely to mean.
Regardless, once you Japanese gets to a point, you begin to realise it is less about the meaning of words and more about the context and usage of language. You find yourself understanding how to use some words you have not studied without having said to yourself 'so what is this word equivalent in English' - and at times there is not a word equivalent in English. That is when you realise you are thinking in Japanese not just translating.
Synaptic connections, context and knowing what the relevant responses to how you want to respond are the key to fluency in my opinion. This is also the case for writing emails etc.
This also explains why people who do homestays where they get to watch people interact with each other progress faster than someone who goes to live with their significant other half. If you are involved in half of the interactions you get less opportunity to mimic the common responses.
Sorry for digressing a little on this one. Once you get that 777 kanji mark you would be well on your way to being able to contextually understand much more for the kanji/sentences you do not fully comprehend - which in turn leads to more kanji being 'gut-felt' understood if not fully formally learnt.
An additional benefit once you get some proficiency is that you can hear a word you do not know - but you can guess the half of the kanji being used in the compound word (and confirm by asking) and that gives you some context of what the word means. Very much like using latin/greek roots in English to breakdown words.
Time to go back to my coding before the day gets away from me. I just needed to respond to this one because I thought Diego's opinion on language learning was too far off the mark from my experience with it.
Edit: Doh, the OP link actually includes this statistic.
If you have 90% you can more often than not guess the uncommon word...
If you recognize 90% of the kanji and know what they mean, that's not the same as knowing words.
No, I'm taking about the premise of the article: if you know 90% of the words most widely used (that is, the above 777 characters for example). Not 90% of the words in general.
If you know the most likely words, you can more often than not guess the unknown words in a phrase from the context.
My point is that those characters aren't words.
Well, some of them are, sometimes.
For example if you just understand the word "car" and the name of a place, it most likely is related to traffic. If that's coming from someone arriving late you can bet he is finding excuses ;)
That's a reason why you shouldn't bad mouth foreigners if they can hear you based on the premise that they don't understand the language. They may not be able to converse with you but they can still get the general idea.
TFA muddles this somewhat, but the research referred to has to do with coverage of kanji characters in a corpus text, not words. Someone who can read 90% of the kanji in a Japanese text might understand significantly more or less than 90% of the words, depending on the text.
Also notable: the research behind this was based on newspaper text corpora, which one could expect to have a lot of kanji and kanji compounds that are uncommon in casual usage.
Also, "understanding" a Kanji is an ill-defined term. Most Kanji have multiple meanings based on context, and many different readings. So for each Kanji you're not learning a single character, but possibly a lot more. Especially the Kanji that correspond to more abstract concepts you cannot learn by themselves. They don't have a concrete meaning like "eat" or "drink". You essentially have to memorize all of their word compounds, a single character does not help.
EDIT: I also started out learning Japanese by memorizing the top 2000 Kanji using Spaced Repetition. While it definitely helped, it wasn't nearly as useful as some of these marketing-driven sites want you to believe. Kanji meaning are too complex to be captured like that. Even if you "know" all Kanji in a word you'll likely not understand the word's meaning unless it's something simple and concrete. I think you are much better of memorizing and studying word compounds. Over time you will automatically "pattern match" the Kanji you see often to their abstract concepts.
Example from 1min of browsing a JP text: Take "可能" which is a very common word that usually means "possible". Knowing the two Kanji (tolerant and ability) it not going to help you. It could mean dozens of other things based on those simplified Kanji meanings. This is not an exception, the majority of words are like this. On the other hand, let me give you a bunch of words containing "可": 許可、可決、可能性、不可欠 (permission, approval, possibility, essential) and you start to pattern-match that "可" corresponds to something like "positive possibility", but it's hard to translate.
Have you ever been around non-native students of Japanese? This is the one word that is pretty much guaranteed to be known by all with an interest in anything remotely Japanese.
Your point is valid, but kawaii is perhaps not the right example.
Most non-native students of Japanese will almost always encounter that word in hiragana, katakana, or even Romaji--encountering that as a kanji is actually remarkably rare.
The fact that the kanji for kawaii is an ateji makes it one of those odd ducks.
Interestingly, it's very common for students of Chinese. The word 可爱 (kě'ài, "cute") is dirt-common, but it isn't native to Chinese -- it originates as a loan from Japanese.
This isn't clear to the Chinese themselves, who use 卡哇伊 (kǎwāyī) if they want to refer to the Japanese word.
There's another modern word for cute, 萌 méng, which is also a loan from Japanese, though I think popular awareness of it as a weird loanword is higher, since its literal meaning ("sprout") is so far removed from the concept "cute".
Which is the trend with native Japanese people too.
You obviously learn it by itself but as Chinese words are mostly a combination of 2 characters, you immediately also have to learn e.g. 可以 (can, may, be able to), 可能 (maybe), 可爱 (cute) etc.
So someone who's learning characters in order to get 90% coverage (or whatever) would not simply learn characters but learn actual words. Learning characters in isolation would not be that helpful, indeed.
When you don't know a word (i.e. a combination) but you know the individual characters it is much easier to learn the new word either by guessing or checking.
In context, the meaning of 可爱 would be fairly straightforward to guess, for example. Even in English 'lovable' is a synonym of 'cute'.
To be honest this whole thread about 可愛い is more or less bonkers, because it's an ateji. The word's meaning doesn't derive from the characters, the characters got arbitrarily attached to an existing word because they were similar in sound and meaning.
As such, the whole thing is about as meaningful as talking about how easy it is to guess that 珈琲 means "coffee"...
In this case, though it does seem that the characters where chosen at least partly because of their actual meaning.
It seems that it is both an ateji and a jukujikun [1] because the word does not come from the characters but the characters do have the correct meaning.
[1]https://en.wiktionary.org/wiki/%E5%8F%AF%E6%84%9B%E3%81%84#J...
Sure, didn't I say they were in my post?
The point was, in a discussion of how well X predicts Y, it's not very useful to examine a test case where Y came first and X was chosen post-hoc to match it.
No, quite the opposite actually ;)
I do take your point that using that word in the discussion above, which is about Japanese was not the best example. On the other hand, it is a good example in chinese.
(B) By arbitrary here I mean that there is no linguistic connection. "Arbitrarily chosen because they are similar" => "chosen for no reason other than their similarity".
Nearly every word involving the kanji 可 in Japanese has something to do with permission: impossible (不可能), possible (可能), permission (許可), approval (可決), etc.
That said, the kanji-centric view of the world is ineffective, and one primarily adopted by beginner learners who have not spent any significant amount of time studying words in Japanese. It is always better to just learn words.
In Chinese, it can be further compounded into many common phrases:
可悲(pathetic, 可+'sad')
可气(annoying/irritating, 可+'anger')
可怜(pitiful, 可+'pity')
可恨(resentful, 可+'hate')
What would be a more suitable example would be 可(ke)乐(le) in my opinion, in which case it means Coke in Chinese, it is a transliteration.
But it could also be 'a thing that can be easily loved'. Rr simply, 'pleasing'. For example, 可口 means tasty. So, 可爱 could be understood in both ways actually.
You don't just load single words up into it and study them on their own, you load words, compounds, sentences, sentence fragments, radicals, etc. etc. and everything else you can, and the SRS system helps you study and remember it. I can understand SRS not being as useful to you if you were only using it for single characters.
I don't think anyone passingly familiar with Japanese thinks otherwise. That is, I think you're arguing against a position here that nobody actually holds. The argument for memorizing kanji is that it makes it easier to learn compound words, not that you'll just know them without learning them.
If you talk to people who don't know much about Japanese they often believe that memorizing the characters is the difficult part and doing so will help you understand a large fraction of written text. I think it's quite a common misconception.
The "777 Kanji for 90% coverage" figure is probably more relevant for illiterate native speakers. Based on a corpus of 167,281 Japanese sentences I had segmented and lemmatized some time ago, you'd need 2,685 words for 90% (up to "故郷", "birthplace", usually written "ふるさと"), 6,564 for 95% (up to "すらり", "slender, smooth"), 14,098 words for 98% (up to "鼻先", "tip of the nose") and 20,657 words for 99% (up to "恐慌", "panic"). Obviously the exact numbers depend a lot on the diversity of the corpus, so don't mistake them for fixed targets to aim for.
Even at the beginning, learning the simple kanji, having the vocabulary I knew/heard living in Japan associated to the kanji made it very useful from the start.
For anybody learning kanji with anki, etc. and failing, I recommend giving wanikani a try. I am just a happy user. Disclaimer: I live in Japan, work in a mixed English/Japanese environment and speak simple Japanese w/ my wife every day
If you are speaking/writing, much easier to form a sentence and search for the one word you don't know than search how to generate an entire sentence structure.
(Here I make the assumption that you'd know all the readings and all but obscure or archaic words for a kanji to count)
With a bit of context it's usually easy to figure out the missing parts
As a counterpoint: “son, if you go to the jawn tomorrow morning, don’t forget to pick up some jawn”
https://www.atlasobscura.com/articles/the-enduring-mystery-o...
Vaya mañana a la vaina esa y me trae la vaina.
Sabe qué, dejémonos de vainas.
"¿Yo dónde dejé esa chimba?"
"Está muy chimba su carro"
"Esa chimba estaba muy chimba"
I don't see anyone claiming: 'learn this 777 characters, and then STOP LEARNING'.
You learn the most frequent first. Then you start reading. Much easier than when you only understand 10% of the words.
So, you get to 90%. Then you read some text and you start to find new words from time to time. That's when you start using a dictionary.
Heck, I'm an avid reader and I am still finding new words (or old words with new meanings) in my native language!
You are better off trying to cover specific domains. I was able to read technical articles related to hardware and software topics pretty easily.
I have no particular reason to disbelieve the headline, it sounds reasonable enough, but this page in specific is doing nothing to persuade me of its main point.
https://researchmap.jp/YOKOYAMA_Shoichi/%E8%B3%87%E6%96%99%E...
It looks like the article author used the kanji from the research as a starting point, but then randomly added in words endings or compound characters, perhaps as their own study aid.
Some critique of using the linked list: That is based on a newspaper corpus from 93, which explains why day/sun is at the top, which they won’t be for list created from Wikipedia or television subtitles, both sources that are arguably closer to “in the wild”
If you know 777 kanji in some way (like associating them with meanings, through your native language) and you haven't crammed on any vocabulary, you absolutely will not be able to read a thing.
In fact, even if you continue that way and memorize over 2000 kanji, and recognize every single one in a given document, you still won't be able to read anything without vocab.
The broadened knowledge will help support vocabulary building, though.
So when reading kanji, how do you look up a word (picture) you don't know? Since there isn't a minimal set of characters, the notion of "alphabetical order" seems impossible. Weird that I've never thought about this until now, but I'm honestly baffled.
Better to have instant dictionary lookup if you're on the phone or laptop than a regular paper dictionary.
Also note that on the site linked to common kanji have a red background. Many of the characters are obscure so radical + stroke count narrows done the choices to very few kanji.
Firstly, unusual words tend to have furigana, which makes it trivial. Next, if the unknown kanji isn't the first one in the word, it's possible to do a prefix-based dictionary lookup. In many cases, also, I've been able to guess a reading by common structure. E.g. both 赤(red) and 跡(traces, remains) have a "seki" reading due to a common element, and 根(root) and 痕 (traces, remains) have a "kon" reading, also due to a common element. You might be able to guess at 痕跡 (konseki) by thinking of 根赤.
How we can find 跡 in the Kokuko Jiten's appendix is by counting the strokes first: 13. Then in the 13 stroke section, of the appendix, we find the subsequence of 13-stroke kanji that have the 足 seven stroke radical. The radicals are sorted by stroke count also, so we can find this subsequence fairly quick.
I can recognize quite a number of words that have at least one kanji which doesn't occur in any other word that I know. I've never studied the kanji in isolation, but I can recognize it in that word.
Jim Breen's KANJIDIC has frequency information.
http://www.edrdg.org/wiki/index.php/KANJIDIC_Project
> The 2,501 most-used characters have a ranking which expresses the relative frequency of occurrence of a character in modern Japanese. The data is based on an analysis of word frequencies in the Mainichi Shimbun over 4 years by Alexandre Girardi. Note: (a) these frequencies are biased towards words and kanji used in newspaper articles, and (b) the relative frequencies for the last few hundred kanji so graded is quite imprecise.
I’ve listed them in a google sheets together with a few other corpora
https://docs.google.com/spreadsheets/d/1yb5dq4ahdwc_g0aQTL3Y...
Choose the jimaku tab for subtitles to see how big the variation between corpus can be.
According to other comments here, it appears that OP list is based on a newspaper corpus from 1993.
I would love to attempt to segment a bunch of Japanese subtitles into words and then do frequency analysis. My interest is in increasing my listening ability, so I want to put the most frequently spoken words into SRS/Anki, and perhaps even break it down by anime.
Alternatively, has anyone already done this?
I don’t have the source code on me, but I scraped it from a website that publishes subtitles. The scraping was easy, the cleaning not, and I believe this spreadsheet is generated from my first attempt at cleaning.
A lot of sources in Japanese nlp and linguistics have a bad habit of changing url often, so it bitrots easily. Sorry.
Take me as English learner for example. I would say I was only able to understand everyday English without too much of a hassle, after I acquired like around 10k words, which as I just checked had a coverage about 98%+.
Noted, it is still NOT enough, actually far from enough. Right now I believe I master around 15k to 20k words, by various estimates, and navigating English on the internet is like a charm, very little context switch in between with my native language.
Still, reading literature is huge undertake for me. I would still need to pardon myself about every once a page that if I stumped upon certain unknown words/phrases and can't move on before fully understands it, my pleasure of reading would be ruined. Such comprise frustrates me still, to this day. On the other hand, I will never have a second thoughts reading most cryptic novel in my own language, understanding might still be a challenge, but unlikely due to my insufficient vocabulary.
The downside of this is, when someone asks while they are reading, "hey what does [word] mean?", I often can't tell the precise definition, because I never looked it up, but I can usually read the sentence and understand it with context.
Those are not set-in-stone science, so it is just an estimate, but I take about 10+ over those years, and the the number from those estimates seems consistent.
Coverage, there is some paper to track, I just googled it.
No idea about its accuracy, but as a native English speaker it's telling me my vocabulary size is 30k words, which sounds roughly correct.
http://kanamozi.org/hikari959-04.html
> If you really want to be native-level Japanese, kanji are essential
There are visually impaired people who have difficulty learning kanji but speak Japanese fluently. Language is not just for people who can read and write, let alone reading and writing complex characters, or spell things "correctly".
Richard Feynman:
If the professors of English will complain to me that the students who come to the universities, after all those years of study, still cannot spell "friend," I say to them that something's the matter with the way you spell friend
Disclaimer: native Chinese speaker, non-fluent Japanese speaker, English sufferer
edit: disclaimer and grammar...
In consequence, this horror-show of a writing system, with its crippling memorization burden for students and malign impediment to progress in science and industry, is the focus of so much intellectual investment and cultural pride that getting rid of it is out of the question. Intolerable though it is, it will continue to be tolerated – leaving English, with a spelling system that positively stinks, smelling almost like a rose.
https://www.chronicle.com/blogs/linguafranca/2016/01/20/the-...
Don't get me wrong, I'm generally in favor of made-in-China products, but those characters to me were the worst made-in-China product I wasted so much time on (and unlike video games, I wasn't even having fun with them.)
I always wonder about something that might sound anecdotal: why isn't there any equivalence of "spelling bee" in Hanzi? Or maybe there is?
> malign impediment to progress in science and industry
What? Do we have any evidence for this?
Even as a non-native reader, reading 100% kana makes my eyes bleed — I wouldn’t wish that on anyone who actually knows kanji.
Or you can watch one of the let's play videos with narration here:
https://m.youtube.com/watch?v=F_UrqsO2JQ0&list=PLC4EWNG6GsuY...
Have you read pages and pages in Hiragana? Or even paragraphs? It's not fun, and it's take a lot of concentration.
It would look awkward initially, but that's just because people are just used to the status quo which is kanji-kana mix.
Reading hiragana (especially without space as word boundaries) is totally different. Reading long hiragana by speaking aloud help, but there is still a problem of ha/wa and he/e.
Also, the very reason there are so many homonyms in written text in the first place is because of the kanji (over)usage - that is, because people think there are visual cues they are less careful about choosing words that are also understood easily by people listening to the words.
When people speak, at least if they are a competent speaker, people tend to avoid the overuse of homonyms (mostly kango).
I can't find the references right now but my memory from reading up on it at one time is that users were very happy with the system, that there is a fair amount of material transcribed into Japanese Braille, that it's quite easy to learn and that children become proficient readers faster than child learners of the regular writing system.
Quoted from my other reply.
I notice this every time I see comments on HN about a topic I actually understand in depth. Quite a few comments are just outright incorrect, most are misguided, but all are confident.
On the other hand, the most confident pronouncements usually come from people unaware of their lack of understanding.
Most written communication is typed. Which involves typing each word phonetically (and then picking the correct kanji from a list). This is orders of magnitude easier than writing a kanji by hand. You can get by without knowing stroke order or without knowing exactly how to write it. As long as you're even vaguely familiar with a kanji, you'll be able to pick the right one from the list.
Give this another generation or so, and I expect the coverage of those "777" will increase to ~100%, and the number of kana-only words will increase dramatically. National pride and linguists be damned, convenience always wins out in the end.
Something similar will probably happen to Kanji. They'll become more and more useless. And, eventually, education will adapt. Give it another generation or so.
[1] I realize that whether or not studying Latin (or Kanji) is a waste of time is a contentious point. If it's your cup of tea, go for it, but don't force it on others. Force them to learn maths instead.
I don’t know how useful these are for independent reading by 5–6-year-olds, but anecdotally they are great material for reading to 2-year-olds, better than most picture books. (Note: some of the recent readers are garbage marketing gimmicks with movie tie-ins, ranging from boring to incomprehensible; skip those.)
This works passively as well, if you see a sign and you understand 90% of what's written on it but there's one word you don't understand, you will much more easily remember that word later because you can place it in context.
The marketing materials for language learning resources tend to make full use of this vagueness, like this one does. I wish these resources instead did more to enlighten their prospects as to what one can actually expect to achieve and in what kind of time frames.
Languages are endlessly deep. "Native" is not even close to the top. Even amongst "native" speakers, skill with and understanding of language is enormously varied. Compare the wedding toast of a skilled public speaker with that of an average one. Compare a literature scholar's understanding of a classic novel with that of an ordinary high school graduate. It's night and day.
IME, 777 kanji wouldn't get you very far in a newspaper and certainly not a novel. It would likely be enough to understand 90% of ordinary emails and text messages.
So many great resources to learn Japanese with these days; this vocabulary list is not one of them.
I made a free website to memorize Kanji that works offline: https://core.cards/. Initially I did maintain a list of the top 100, top 500, and top 1000 (approx) if I recall correctly, extracted from Wikipedia lists, to learn Japanese Kanji. But now I've switched it to just follow the JLPT because they were almost the same.
How many words do you need to comprehend for daily competency? Would the 10,000 suffice?
How many words do you need to be able to watch anime aimed at children (eg Bono Bono) or teenagers (Boku no Hero Academia)?
Putting aside the argument of whether removing all Hanzi from Japanese text would actually be more efficient or not, the question to me is: why stop at Hanzi? Why not romanizating all the Japanese literature? Surely almost all the reasoning in favor of getting rid of Hanzi can also apply here?
edit: grammar
But even then they still use Hanja to disambiguate sometimes.
edit: obviously Kana has already been created
https://ja.m.wikipedia.org/wiki/%E3%83%AD%E3%83%BC%E3%83%9E%...
I'm personally fine with both romaji and kana (learning kana is not a big cognitive burden anyways).
I'm even fine with kanji (or Latin or anything) if people learn it voluntarily. What's not ok is kids being forced to learn them in school.
The difference is that removing kanji from Japanese profoundly changes it, in that tons and tons of words that were previously distinct become indistinguishable. Hence people arguing about getting rid of kanji are arguing about the benefits of the language being easier to learn, vs. the drawbacks of having huge masses of distinct words that are all written the same way.
Whereas, using kana vs. romanization isn't really an important distinction. Nobody argues which is better because they're basically equivalent; anyone who knows one could learn the other in a matter of days or weeks.
If you're going to rote memorize something, I'd probably start with the radicals.
1) Does anyone have a resource for Japanese subtitles (in Japanese/kanji, not English)?
2) Does anyone have good frequency { word => frequency } lists? Especially if they are topical, eg. school-related, anime-related, industry-related.
3) What are the best programs for segmenting Japanese text into words reliably?
4) Does anyone have a vocabulary set for any given manga, anime, or film that you could study before watching?
5) In addition to Anki and Wanikani, what are good SRS apps or programs?
6) Does anyone use Skype (or similar) to practice with native Japanese speakers? How is it? How did you find people to practice with?
7) What is the inflection point (in terms of raw # of vocabulary) to being able to understand Japanese anime or drama? What JLPT level does this correspond to?
8) How many new words do you acquire per day of study? How long have you been studying? Have you taken any of the JLPT tests?
Not long after this HN thread, this thread popped up on Reddit. It answered a lot of my questions ([1], [2], [4], and [7]), and I found it immensely useful:
https://www.reddit.com/r/LearnJapanese/comments/crlsqj/googl...
I'm a native Japanese speaker. I agree that 777 kanjis are contained in common sentences at the rate of 90%, but it doesn't mean you can complete 90% of Japanese. Actually, these sentences are very easy to read for Japanese people, although it's hard for foreigners because of the difference of vocabulary. I also reviewed the list. I believe「経(ふ)」and「格別空」are odd as a word. In addition, 「恬然」and「整復」 are not frequently appeared so I think there was a bias to choose sources.
Here are examples of native Japanese speakers. Usually we can achieve 20,000 easily. My result was 36,000 words.
http://burusoku-vip.com/archives/1798300.html https://b.hatena.ne.jp/entry/s/www.arealme.com/japanese-voca...
I chose Japanese because I watch and read a lot of stuff from Japan. It's also very different from English, which is fun.
p.s. just confirmed that the list is not enough for filing a tax in Japan. They don't have words like 所得 (income), 控除 (deduction) or 医療費 (medical expense).
No Kanji at all, english words and some romaji like "tsu" which, can also be kana.
I was disappointed by the article.
Memorize the 625 most used words to jump-start your language learning and then you can move to grammar and other stuff.
https://japanese.stackexchange.com/questions/11735/how-many-...
Normally Zipf's law refers to the frequency being inversely proportional to rank - i.e. the 3rd most common element would be 1/3 as frequent as the first, not 1/4th.
Your argument only holds if Kanji was designed to be optimal in compressing information, which is of course not how Kanji came to be.
That doesn't mean you are wrong, but your argument is invalid.
https://simple.wikipedia.org/wiki/Wikipedia:Simple_English_W...
Another example of "using only the ten hundred words people use most often" is Randall Munroe's "Up Goer Five":
It may be helpful to think of "characters" as representing some middle ground between words and alphabetic letters, a little like word stems.
But the list isn’t at all comprehensive. There are a considerable number of kanji in regular use that aren’t on the list, and when you include place/person names, it grows massively.
It’s possible to memorize all the standard kanji, but crack open a history book and you won’t recognize half of the words.
It's not hard - I've done this before using a corpus made up of work documents. There's a part-of-speech analysts tool called mecab that gives the word stems, and makes it easy to find word boundaries (since Japanese doesn't use spaces).
The output went into Anki and it didn't take long before I was reading emails and documents at work fairly easily.
https://en.wikipedia.org/wiki/Pareto_distribution#Applicatio...
It's really any distribution with the CDF of the form x^(-a)