I'm sorry, but I fail to see how "This visually displays as a single unit" could ever differ from "Display size in a monospace font" or "Thing that gets deleted when you hit backspace".
I'm sorry, but I fail to see how "This visually displays as a single unit" could ever differ from "Display size in a monospace font" or "Thing that gets deleted when you hit backspace".
* Coding ligatures often display as a single glyph (maybe occupying a single-width character space, or maybe spread out over multiple spaces), but are composed of multiple glyphs. The ligature may "look" like a single character for purposes of selection and cursoring, but it can act like multiple characters when subject to backspacing.
* Similarly, I've seen keyboard interfaces for various languages (e.g., Hindi) where standard grapheme cluster rules bind together a group of code points, but the grapheme cluster was composed from multiple key presses (which typically add one code point each to the cluster). And in some such interfaces I've seen, the cluster can be decomposed by an equal number of backspace presses. I don't have a good sense of how much a monospaced Hindi font makes sense, but it's definitely a case where a "character" doesn't always act "character-like".
For example, when == is written, connect them to be a 2 column wide = instead.
Or when === is written, display a three column equals sign, but it's three bars instead of two.
Some clusters are going to be multiple characters wide.
> thing that gets deleted when you hit backspace
Some clusters are meant to be composted of multiple keystrokes and a natural editing experience would allow users to delete the last stroke.
Look into how Korean works.
As for "display size in monospace font", emojis and CJK characters are usually two units wide, not one (although, to be honest, there's a fair amount of bugs in the Unicode properties that define this).
[1] https://android.googlesource.com/platform/frameworks/base/+/...
Discussion on HN: https://news.ycombinator.com/item?id=31858311
Seems like the right answer for codepoints vs graphemes, unfortunately, is dependent on the context.
edit: emojis are filtered by HN
I guess it's more ambiguous for some languages that can have long ligatures though.
Where some layouts may require this method for some characters, another keyboard layout may have the same character on a dedicated key.
The program receives the combined character as one unit, and does not need to be aware of different keyboard layouts.
Which ones? At least the French and German ones don’t work like that: there is no composing, just separate keys for all the characters with diacritics that appear in the language.
It's worth mentioning of course that there are no-dead-keys variants of the keyboard layout but this has been pretty much the norm on Windows since the 1990s I think.
That being exactly the way “floating diacritics” in ISO 2022 (or properly one of its Latin encodings, T.51 = ISO 6937) work, amusingly. I wonder which came first. (Yes, I know that a<BS>` came first, the ASCII spec even says that this should give you an accented character IIRC. Or perhaps it was one of the other “don’t call it ASCII” specs—ISO 646? IA5?..)
A美C
would take up the width of four ASCII monospace characters, the “美” being double-width.Similarly, for composed characters like say the ligature “ff”, you may want to backspace as if it was two “f”s (which logically it is, and decomposes to in NFKD normalization).
Latin: Katakana
Full width: カタカナ
Half width: カタカナ
(How that fixed width text looks in a web browser is anyone’s guess though. On iOS none of the Japanese kana stay on the fixed grid.)𒐫
𒈙
꧅