How does Chrome decide what to highlight when you double-click Japanese text?
stackoverflow.com
stackoverflow.com
Assuming Blink is using the same technique for text selection as V8 is for the public Intl.v8BreakIterator method, that's how Chrome's handling this-- Intl.v8BreakIterator is a pretty thin wrapper around the ICU BreakIterator implementation: https://chromium.googlesource.com/v8/v8/+/refs/heads/master/...
Chrome uses v8's Intl.v8BreakIterator which uses icu::BreakIterator, which, for Japanese text, uses a big long list of Japanese words to try to figure out what is a word and what isn't. I've worked on a similar segmenter for Chinese and yeah, quality isn't great but it works in enough cases to be useful.
and then according to the LICENSE file[1], the dictionary :
# The word list in cjdict.txt are generated by combining three word lists
# listed below with further processing for compound word breaking. The
# frequency is generated with an iterative training against Google web
# corpora.
#
# * Libtabe (Chinese)
# - https://sourceforge.net/project/?group_id=1519
# - Its license terms and conditions are shown below.
#
# * IPADIC (Japanese)
# - http://chasen.aist-nara.ac.jp/chasen/distribution.html
# - Its license terms and conditions are shown below.
#
It's interesting to see some of the other techniques used in that engine, such as a special function to figure out the weights of potential katakana word splits.[1] https://github.com/unicode-org/icu/blob/6417a3b720d8ae3643f7...
Firfox is just not a very good macOS citizen, sadly.
If I could set up a custom keyboard shortcut for the service, it would work great—but support for custom keyboard shortcuts also doesn't work in macOS Firefox[1]! (Unless you set a combination which Firefox doesn't already use for something, and it already uses everything sensible, so you have to set something long and annoying.)
It's not like English highlighting does complex grammatical analysis to make sure "Project Manager" gets highlighted as one chunk and "eventUpdate" gets highlighted as two chunks, most implementations just breaks at spaces like the user expects.
Consistent behavior > inconsistent behavior, almost always. If I can anticipate what my computer is going to do, I can plan around it, even if it means a bit of extra manual work.
In general, whitespace characters have no place or significance inside a Japanese sentence, and most of the whitespace in Japanese typesetting is built into punctuation marks.
Furthermore, even most punctuation is optional in Japanese. The full stop 。 and comma 、 are mostly a matter of preference, sometimes spaces are used in place of full stops or commas.
Firefox's approach is fairly useless in this regard. Even if it's predictable from a technical perspective, it's not predictable for a reader who naturally processes semantic breaks rather than technical ones. Unlike in English, where a space is both semantic and visual, hiragana-kanji boundaries often don't mean anything. As a result, for me at least, Firefox's breaks feel a lot more random than Chrome's, which, while dodgy, are often fine.
Having used Firefox as my main browser since 2006, I remember when I discovered this feature in Chrome, and being shocked at how much of an effect that minor improvement had for me. It's not a deal-breaker, certainly, but it's become my one big annoyance with Firefox.
Machine-readable strings can still just be an ASCII byte array, but we need to keep the two separate.
[Edit: Ah, it might not be common knowledge but you can drag on the double click to select more words. You can also hold alt while editing text to jump words, and even combine alt+shift and alt+delete. Windows and Apple differ on the boundaries wrt whitespace, but I'm fairly sure they're both alt. Wordjumps while editing is an unappreciated superpower imo.]
Also FWIW Safari also supports this.
I don't know Alt + Shift/Delete action, how is it useful?
Now consider another anglophone arbitrarily; how likely does it seem to you that they would share the exact same set of expectations of the above with you? Because if they don't, then one of you is going to have their expectations violated.
Now give it a try: do the highlights match your expectations? (Moreover: are those expectations consistent with where you placed the word boundaries?)
If your expectations of what double-click highlight does for English is consistent with what actually happens when you double-click, it is deeply unlikely that the causality goes word boundary -> highlight expectation; if they do, it's more likely that double-click highlight has influenced your sense of word boundaries.
But it's even more likely that you your expectations of double-click highlight don't actually match up with your intuition for word boundaries, and you just have a pretty good model of how double-click highlight behaves independently of where you think the word boundaries are.
(The notion of word boundaries in this context is fraught anyway. Even setting aside the problem of compound words, the phonological word boundaries of Japanese don't match up with the lexicographical ones, in both directions.)
For me at least, double-clicking is like fuzzy string matching. It's an imprecise behaviour by nature, but should be designed to be useful. Firefox's behaviour in the case of Japanese, but especially Chinese, often isn't useful, not because it doesn't perfectly reflect linguistic expectations, but because it usually selects too much--meaning you can't correct your selection by moving the cursor back and forth, and instead, have to stop and select using a different method.
Take the following passage:
もうこんなことは終わってほしいと願うばかりだ
Different people may break apart もうこんなことは differently, but it's hard to argue that it's one word. Chrome (and Word) separate it like this: もう こんな こと は, and whether or not you agree that each one of those parts are words, simply splitting もうこんなことは into parts semi-logically makes it smoother to edit (for me, at least! I understand most people might not be as dependent on double-click selection).
The same behaviour for Chinese is so useless that it's not even worth mentioning, since you'll always select either a large part of a sentence, or an entire sentence at once. I mean, I love the word 匹配的近似度用如下方法来度量.
I actually work with Japanese tokenization a lot - I took over maintenance of the most popular Python Mecab wrapper last year, and I have another Cython-based wrapper that I maintain.
Word boundary awareness for Japanese is a pretty uncommon feature in applications, so I was surprised to see the feature had been in Chrome all along, even if it's buried and the quality has issues.
Anyway, thanks to everyone who tracked down the icu imlementation and the relevant part of Chrome!
Each unicode character has certain properties, one of which is whether that character indicates a break before / after itself.
I've done extensive research on this for my job, but unfortunately don't have time to do the whole writeup here. Here are several resources for those who are interested
Info on break opportunities:
https://unicode.org/reports/tr14/#BreakOpportunities
The entire Unicode Character Database (~80MB XML file last I checked)
https://unicode.org/reports/tr44/
The properties within the UCD are hard to parse, here's a reference if you're interested:
https://unicode.org/reports/tr14/#Table1
https://www.unicode.org/Public/5.2.0/ucd/PropertyAliases.txt
https://www.unicode.org/Public/5.2.0/ucd/PropertyValueAliase...
Overall, word / line breaking in Unicode in no-space languages is a very difficult problem. Where the UCD says there can be a line break isn't where a native speaker would put one. In order to do it correctly you have to bring in Natural Language Processing, but that has its own set of complexities.
In summary: I18N is hard!
function tokenizeJA(text) {
var it = Intl.v8BreakIterator(['ja-JP'], {type:'word'})
it.adoptText(text)
var words = []
var cur = 0, prev = 0
while (cur < text.length) {
prev = cur
cur = it.next()
words.push(text.substring(prev, cur))
}
return words
}
console.log(tokenizeJA("今天要去哪裡?"))
still seems to parse just fine. so most likely just using the passed input to parse.This is where there are sometimes discrepancies between how a given browser or device would output this data, as it could be working off of an outdated version of Unicode's data.
Some devices even overwrite the default Unicode behavior. There are just SO many languages and SO many regions and SO many combinations thereof that even Unicode can't cover all the bases. It's all very fascinating from an engineering perspective.
Your first link also says:
> To handle certain situations, some line breaking implementations use techniques that cannot be expressed within the framework of the Unicode Line Breaking Algorithm. Examples include using dictionaries of words for languages that do not use spaces
[0] posted in another top-level comment: http://userguide.icu-project.org/boundaryanalysis
* with a few exceptions, of course :)
- Korean segmentation is way easier than Chinese and Japanese because it uses spaces between words (they are thus distinctive, not like Vietnamese which use it on syllable boundaries. Vietnamese consequently also requires segmentation)
- Chinese and Japanese segmentation are hard NLP problems that are not fixed, so they are in no way "easier" than the same take for other languages
- The limited valid combination of characters that form words in Chinese doesn’t mean segmentation is easy because there is still ambiguity in how sentence can be split. There is still no tool that produce "perfect" result
- difference in scripts is indeed used in some segmentation algorithms for Japanese, but that doesn’t solve the issue totally
- the phonetic/non-phonetic parts of Chinese characters, have been used in at least one researcher paper (too lazy to find the reference again, it didn’t worked well anyway) but are not in state of art method. So contemporary Korean not using a lot of Hanja anymore has no influence on the difficulty of segmenting it
https://chinese.yabla.com/chinese-english-pinyin-dictionary....
It does basic path finding, and then picks the best path based on the following rules:
1) Fewest words
2) Least variance in word length (e.g. prefer a 2,2 character split vs a 3-1 split)
3) Solo Freedom (this is based on corpus analysis which tags characters with a probability of being a 1 character word. For example 王家庭 (this is either "Wang Household" (王 家庭) or "Prince's courtyard" (王家 庭) and we split as Wang Household, because Wang 王 is a common name that frequently appears in isolation, and 庭 is less likely to be in isolation. It is interesting that solo freedom works better than comparing the corpus frequency of "Prince" 王家 vs "Household" 家庭.
It works reasonably well. A surprising number of people use it every day.
What? 王家 doesn't mean "prince". Or at least, there is no such dictionary entry in the ABC dictionary or in the 汉语大词典. I would expect 王家 to mean "prince's household", in the same way that 皇家 means "imperial household".
汉语大词典 has two glosses for 王家:
1. 犹王室,王朝,朝廷。 [Equivalent to "royal family"/"royal court".]
2. 王侯之家。 [An aristocratic household.]
There's no problem with the concept of the phrase 王家庭 meaning "prince's courtyard", since a home can easily contain a courtyard. But the phrase should arguably be segmented 王-家-庭. (Or not -- there's very little to distinguish the idea 'one word, "the prince's household"' from 'two words, "prince"/"household"'.) Regardless of that choice, the courtyard is being associated with a household, not a person.
"King" is also a decently well-known English surname, so we can draw a pretty close analogy to the distinction between 王家, the royal household, and 王家, the 王 family. Compare the English sentences
1. That's the king's house. [The king lives there]
2. That's the Kings' house. [The Kings live there]
I have been told there are quite long sentences that humorously can be segmented two ways and have completely different meanings, but I can't find any. My favorite in English is expertsexchange.com had to change the domain to experts-exchange.com :)
No need; as an example of segmentation, the one you've presented is fine. The biggest problem, mostly irrelevant, is that "prince's courtyard" is pretty archaic. I was objecting to the translation-in-passing of 王家 as "prince". It's defensible to segment the "prince's courtyard" sense 1-1-1, but it's also defensible to segment it 2-1.
I don't have an example of the type of you're looking for to hand, but I will see about finding one.
I find this line hilarious for some reason. Reminds me of the line about being a tourist in France, "French people don't expect you to speak French, but they appreciate it when you try"
Because I was a bit surprised about that and made me wonder if opening this JSFiddle on Safari would work at all (I’m on a phone so I can’t test).
Still, I think my original reading is correct, because I don't think there is any issue with the "quality" of v8 inside of jsfiddle. While imagining Chrome doing its best to identify real words in long strings of Japanese text and failing spectacularly just made me laugh again.
TypeError: Intl.v8BreakIterator is not a function. (In 'Intl.v8BreakIterator(['ja-JP'], {type:'word'})', 'Intl.v8BreakIterator' is undefined)
I only knew a handful of phrases though, so anything off script and I was pretty lost.
[0]: http://www.solutions.asia/2016/10/japanese-tokenization.html...
"Viet Nam" is also, actually, the "official" English way to write it. (Check how the UN puts it on all their stuff.) However, most Europeans don't do that in their languages, so it usually gets written as Vietnam even by Vietnamese when they're writing European languages.
In both cases, some liberties have been taken with notation to intentionally encourage silly mis-readings; It happens much less often in ordinary text.
すももももももももの内
"Japanese plums (sumomo) and peaches are both kinds of peaches"
In speech intonation would make the word boundaries here clear, but in writing it looks odd.
This isn't a tongue twister, but tokenizers often fail on it:
外国人参政権は難しい問題だ
"Voting rights for foreigners is a complex problem."
外国人 (foreigner) / 参政権 (gov participation rights) is the right tokenization, but a common error is to parse it as 外国 (foreign) / 人参 (carrot) / 政権 (political power).
The really interesting question is how does Chrome decide what to highlight when you double-click Thai text? It's a non-breaking, (baroquely) phonetic script and the training set is much smaller.
Many all-kana texts employ spaces. All-kana with no spaces is a nightmare.
On one hand, language has always been an amorphous thing and has changed throughout history. On the other could some changes change the very nature and heart of it? No one agrees on those trade-offs, and I argue likely never will.
This has been always the case. I mean, think of how modern latin letter forms developed. Think back to the Roman Empire, when the letters were etched in stone or wax tablets, with only capitals (before the invention of minuscule), drawn with straight lines that are easy to chisel. Think of the later periods when many books were written by scribes using pen and ink, resulting in changes to letters so they could write entire words in one hand motion. Later still, we have the printing press and its demands that each letter be a discrete unit, doing away with the wavy flow of words. Today most people write using print letters simply because the majority of writing you would encounter is in print, so that's what's easiest to read. Those changes probably don't register to you as egregious as the one parent proposes, due to the fact that the person enforcing the change and the person developing the technology are one and the same - a native speaker modifying his own language to fit the tools at hand. Which I think brings us to a resolution - the native speakers of the language are the ones to decide how to change it to suit their everyday needs, including making it easier to produce and consume with the tools of the day.
Agree with you here, it's well outside my rights as an English speaking American to have an opinion of any merit. However I would still struggle to call handing it over to native speakers a resolution. Many native speakers have incredibly strong feelings on both sides of the argument; and often for very good and valid reasons.
I'm glad the Unicode Consortium exists, because I am CERTAIN I don't want to be the decider or facilitator of these discussions. Way above my pay grade.
To choose some random capital Roman letters, how about... SPQR? Those got inscribed all the time.
Roman letters don't really show any bias towards straight lines. Don't confuse the fact that we split the Roman letter V into two letters U/V with the idea that they didn't carve curved letters. Check out this awesome plaque from the year 90: https://www.timesofisrael.com/in-a-bronze-inscription-a-remn...
Don't leap to conclusions about the benefit of breaking up the text. For example, reading Mandarin pinyin is much more difficult than reading characters, but neither of those involves a mixed script. (And heck, the pinyin might have word spacing! [1])
The reason is pretty obvious - characters contain a lot more than just phonetic information. That makes things harder on the writer, and easier on the reader.
[1] Pinyin produced for foreign learners usually has word spacing. Pinyin produced for Chinese children usually doesn't.
I think having spaces is useful, too, for learning speed and expression.
Early Japanese on computers also used half-width kana but now it's all properly fullwidth.
Don't mistake technical limitations for choice.
This cannot actually be true; the spoken sentences do not indicate word breaks, but grade schoolers (and toddlers) can parse them just fine.
The most obvious way to see this would be to think about synthesized speech, which almost never uses natural prosody. Think of the voice of Stephen Hawking, or just a splicing of some prerecorded options. ("At the tone, the time will be / TWO / TWENTY / THREE / and / TEN / seconds.")