Show HN: Neural Japanese Transliteration
github.com
github.com
These is also a reset for "Conversion Learning" under the Keyboard > Input Sources > Japanese that I've just found thanks to the hints from the thread on this feature in iOS.
Might be selection bias but I mostly notice people using the 10-key click one
The real problem is transforming the phonetic transliteration into the correct word in either kanji (for most Japanese words) or katakana (for words with foreign origin).
This problem is akin to disambiguating between two homonymes (which are much more frequent in Japanese). In some cases it is easy by looking the previous words, but in some it is heavily context dependent.
Nowadays, most japanese typing system will propose a list of kanjis as you type that corresponds to the most frequent writting of your transliteration, but sometimes for unusual kanjis o(or people's name) you have to dig deep into the list.
I can see how such a system could improve typing speed in Japanese.
But none of the Japanese I know use the hiragana directly. They told me it is mostly old people who use it. Almost every Japanese knows romajis now, there is no additional cost of learning a new alphabet.
Given a Japanese sentence (that uses kanji), figure out the proper reading for each Kanji character, using a neutral network.
I know there are already hardcoded analyzers, like kuromoji, but they produce incorrect answers in a lot of edge cases.
The kanji -> kana direction should be considerably easier than the kana -> kanji direction. There are many fewer sources of ambiguity, and the space of possible answers is smaller.
Not sure how well this model performs, but the task is not novel.
P.S. Mandarin has a fair few homophones as well (yes, tones and all). English has tonnes of phonemes, and still we have loads of homophones (and in some of the most common words, no less!). Japanese, to my elementary-level ear, doesn't sound an order of magnitude more ambiguous than English.
I don't have any citations, but subjectively I disagree. One can think of English words that are homophones, but in JP the challenge is more to think of words that aren't.
On one hand, if you include uncommon words (as a keyboard's corpus would) then practically any medium-length word will have multiple possibilities, which is not the case in English. But much more important is that, where an English homophone typically has two or possibly three interpretations, a kanji jukugo might have 3-4 everyday meanings and a bunch more uncommon ones.
And all this is on top of the matter of word boundaries. If the user enters し, there could be 10+ possible transliterations of that character as a standalone word/particle, above and beyond whatever readings are possible together with the characters before and after.
So I don't know what an order of magnitude would mean in this context, but I think the whole matter is significantly more ambiguous than English.
But a stream of romaji furigana with no spaces is quite ambiguous—since there's nothing to indicate word boundaries, any substring of the input might turn out to have actually intended to be e.g. a katakana spelling of a name.
If CJK IMEs expected and required people to hit the spacebar between the "words" (lexer tokens) of the provided input for matching, they'd have a much simpler job. But as it is, they fall over quite badly when you type multiple words into the IME input box, and are mostly only usable if you resolve single words at a time (which, sadly, throws away a lot of the inter-word context that would otherwise be available for matching.)
This is (surprisingly) not true. People do not pause between words, however when listening to a language that they understand, they do perceive pauses between words; even though such pauses do not exist.
I'm not a linguist; I don't know what the proper name is for the thing people do between each pair of spoken words—that they don't do inside words—but I do know that there is something people do there. I would call it "a pause" because that is the function it serves. It's an overlap of lesser "terminal" sounds that forms something that is detectably a semantic gap—like the pause between crossfaded tracks on a gapless record, or between movements in a concerto.
Whatever it is, it is there, because speech-recognition systems use it to detect spoken word boundaries regardless of language. (This heuristic does screw up sometimes; spoken language does often "slur" particular word-pairs together. But it's rare enough that these can be trained as specific exceptions to the rule, rather than needing to throw out the rule.)
Seems more productive (and enlightening to all) than the agree/disagree dialogue here.
The best way to see this is to try listening to a language you do not understand, and try to identify word boundries.
Indeed, the paper I link argues that some phonetic cue must exist because babies can recognise word boundries.
[1] https://www.sissa.it/cns/Articles/94_doInfPerceiveWordB.pdf
A monotonal, monotempo sound would not be able to make that difference audible
The Japanese disambiguate word boundaries in spoken language using the pitch accent as the primary clue. Tokyo Japanese has a phenomenon called initial rise, which differentiates the pitch between the two first moras of an accent phrase – either the pitch rises or steeply falls.
Here's an example - upper case: high pitch, lower case: low pitch.
KYOu, kaINI iKIMAshita
today, to buy I went
KYOu KAini iKIMAshita
today, to meeting I went
kyoUKAINI iKIMAshita
to chuckh I wentAlso intonation, which is not captured by the written system at all. Japanese isn't strongly tonal in the way Chinese is, but it has a regional prosody, like Swedish, which helps in disambiguating meaning.
Edit: your P.S. made me remember the "ma-ma-ma-ma..." mouthful the Chinese language students I studied in parallel with discovered (apparently well-known - something about a horse and a...mother?). If tones are not represented in romanised Chinese, things seem to get tricky, indeed.
https://en.wiktionary.org/wiki/%E5%AA%BD#Chinese 媽 mā 'mother'
https://en.wiktionary.org/wiki/%E9%BA%BB#Chinese 麻 má 'hemp' (sometimes 'flax')
https://en.wiktionary.org/wiki/%E9%A6%AC#Chinese 馬 mǎ 'horse'
https://en.wiktionary.org/wiki/%E7%BD%B5#Chinese 罵 mà 'scold'
It's also cool that you can see that 媽 is made up of "semantic 女 + phonetic 馬", where 女 means 'lady' and 馬 sounds like "ma", so the character was meant to suggest "a word relating to ladies that sounds like ma".
https://en.wikipedia.org/wiki/Chinese_character_classificati...
https://en.wiktionary.org/wiki/%E5%97%8E#Chinese 嗎 ma 'question particle'
Apparently the phono-semantic derivation for that is "mouth ma" (maybe because a mouth is used to ask questions?).
https://en.wikipedia.org/wiki/Lion-Eating_Poet_in_the_Stone_...
(roughly: "you dare to scold my mother's horse?")
Anybody familiar with the history of this project?