e.g., Hann-Tzong Wang's (王漢宗) free font collection[1] includes two typefaces with phonetic pronunciation guidance. These are wp{0..3}10-05.ttf and wp{0..3}10-08.ttf [2] As you can see from the filenames, there are actually four different font files for each of these two typefaces. The font files numbered {1..3} are for 「破音字」, characters with alternate pronunciation.
When a user types a word like 「給予」 (ㄐ一ˇㄩˇ/jǐ yǔ) for which there is an alternate, less-common pronunciation (ㄐ一ˇ/jǐ instead of ㄍㄟˇ/gěi for 給) they simply change the font for just the affected character to the variant with the correct pronunciation.
In the case of this Cantonese Font, the authors distribute a single .ttf (alongside a “phrasebook” .ttf whose purpose is not clear to me) and indicate in the Roadmap section of the website that ligature support must be enabled. If alternate pronunciations are common in Cantonese, then I suspect that they must use some ligature-based method. I would have to imagine there must be cases where this could be ambiguous, but I don't know how you would resolve those.
(In practice, just swapping the font on a single character works fairly well.)
[1] https://code.google.com/archive/p/wangfonts/
[2] https://dywang.csie.cyut.edu.tw/dywang/download/pdf/sample-o...
This is the phonetic system that is used most commonly in Taiwan.
Typically, phonetic pronunciation guidance is used only in educational materials. For native speakers, this means materials only for very small children. However, in Taiwan, it's not uncommon to see 注音符號 guidance to indicate when a word should be said with a non-Mandarin pronunciation. You'll see this, for example, in shops whose names contain a pun when using the Taiwanese Hokkien pronunciation.
There are other fonts that include pronunciation guidance in hànyǔ pīnyīn 漢語拼音, the phonetic system used most commonly in China: e.g., http://fonts.mobanwang.com/200909/5832.html
I don't think there are any fonts with pronunciation guidance in any of the other phonetic systems (e.g., 通用拼音, Wade-Giles, 國語羅馬字) but these have almost all fallen into disuse and appear only in old signage, historical place names, or in people's names.
(Presumably, if you are born in Taiwan, you get to pick how you want your named romanized… especially since you may want the spelling of your name to match that of your relatives. But are you allowed to pick any transliteration you desire?)
The page says they do handle variations:
Pronunciation in the Cantonese Font adapts to the context. Based on what comes before or after, the Jyutping romanization changes to the right one. The magic behind this is a careful curation from 100,000 contexts where the pronunciation differs from the standalone character.Setting aside for a minute the question of whether you _should_, I wonder how far you can take this? I.e. what limits are there on how much context you can take into account, etc.?
This is all off the beaten path, so I suspect the answer is no one knows. Font tables have a limit of 65k characters, but this ceiling can be busted in whacky ways using multiple lookups, useExtension... Practically, font building tools / operations crash (mysteriously), stalls (mysteriously), or slows to a crawl (indistinguishable from stalling), and the Cantonese Font about pushes the limit.
----
Conceptually it is simple: 1. assign a default (most likely) sound for each character, 2. loop through contexts, extracting words (char-combos) where the sound is different from the default ("alt-word") 3. create SVGs + font-paths (fallback for incompatible systems) for every char and every alt-word 4. assign a ligature to substitute each char-sequence that forms the alt-word (e.g., "when 乾 隆 appears adjacently, replace with `uniF1234` (the codepoint for the alt-word 乾隆")
It is not perfect, but I didn't expect this to work so well, and was stunned when the testers report high accuracy. I have always believed that bespoke computation with word segmentation (with some 1M frequency attached library) and large data-bank (100k+ words) was necessary.
----
Practically it was horrific, tedious, mind-numbing, gawd-awful set of "why this doesn't work": 1. SVG automation that works for 10^3 breaks with 10^5 2. what worked for Latin breaks for unicode 3. what worked for unicode breaks for PUA 4. what worked for monochrome breaks for color 5. what worked for single glyphs breaks for ligatures 6. what?! The assignments in the database is wrong?? 7. [...]
As I was trying to coerce the system to do what it wasn't designed to do, many of these breaks are undocumented, pretty mysterious to solve, and some steps just got manually gritted through. (And each of the 15k+ glyphs got gritted through about five times.)
It does look pretty elegant at the end ;)
The history of digital fonts added a great deal of complexity to font formats, and without him writing such a concise yet comprehensive guide, I would have been stuck for even longer.
> Unfortunately, without being able to do proper word segmentation, this will remain a limitation.
Can the user manually add a zero width space to help?
(For everyone else wonder what ackfoobar is proposing: let's take the phrase (if you don't read Chinese, just treat them as shapes) 香港地少人多, properly segmented, is 香港.地少.人多. The font treats this incorrectly, because "香港地" is a commonly used fragment, the 地 in the fragment have a special sound, and parsing as 香港地.少.人多 gives a mistaken sound for 地.
Ackfoobar is absolutely correct that we can coerce the correct reading by going 香港[ ]地少人多 --- where the [ ] is an invisible spacer. My contention is that most users don't know how to do that in their favorite word processor.
Someone is probably thinking, could you add "香港地少" as a fragment? Purist says it's not pretty, but I'm a pragmatist, so I did do many of these patching. Doing this or not relies on some acumen as a native speaker, and there were hundreds of these decisions made. This language knowledge would be necessary if someone were to do Mandarin (or Thai or, ...))
I notice you're using OpenType-SVG here; have you investigated whether it would be possible to implement this using COLRv1 (which would potentially result in a lighter-weight font, I suspect, and eventually wider support)? Or are there technical limitations in COLRv1 that make it impossible?
But I did try to make it into COLRv1 (as well as COLR/CPAL). The only tools that build COLRv1 right now are the tools from the Google Fonts team; I remember them stalling for hours before saying completion, yet the output was broken (I can't remember how it was broken).
I personally would love to see a COLR/CPAL version, and have some idea on how that could happen. But I probably should be working on some revenue-generating product instead ;)
* 行: xíng or háng
* 的: de or dì
* 长: cháng or zhǎng
(plus I'm sure many more that I can't think of just right now)
Example
觉得 juede, to think 睡觉 shuijiao, to sleep
Here the same character is pronounced jue or jiao depending on context
- 说服/說服 Mandarin: shuì fú Cantonese: seoi3 fuk6
- 说话/說話 Mandarin: shuō huà Cantonese: syut3 waa6
Another large class class comes from vestiges of derivational morphology in Old Chinese: https://en.wikipedia.org/wiki/Homograph#In_Chinese For instance, the character 度 in modern Mandarin can be pronounced dù (when used as a noun) or duó (when used as a verb), both of which derived from Old Chinese /daːɡs/ and /daːɡ/, respectively.
With Simplified Chinese characters, some of them come from the merger of originally different words that had similar, but not exactly the same pronunciations. For instance, both 髮 (fà) and 發 (fā) were merged into 发.
The 3/6 tones in Cantonese and ˋ (4th, falling tone) in Mandarin are the "departing" tone, which comes from the departing tone in Middle Chinese, which I believe comes from the -s ending in Old Chinese.
The idea is that each honzi has exactly one meaning is a misconception.
With respect to ligatures, if by that you mean the length of the same word across different Sinitic languages, that depends on the specific language and its phonology. Mandarin, for instance, has lost a large number of finals over the course of its evolution which has resulted in words generally being longer and requiring extra syllables to resolve the phonetic ambiguity. The Sinitic languages that have retained more finals (and sounds in general) tend to have more of shorter words. Cantonese is one of them albeit not the only one.