Breaking our Latin-1 assumptions (2017)
manishearth.github.io
manishearth.github.io
If your system can handle these it can probably handle most global text.
Cursive has existed alongside print letters going all the way back to antiquity.
Both forms existed before print, during the inception of print, and long after print.
It does sound rather snarky
Your comment on the other hand… /snark
it wasn't an accident, the printing press was banned by the ottoman empire in 1485.
Do you really think "those Arabs and what-not" did not have printing before Unicode condescended to "accommodate them" in 1988? Traditional printing could very well do Indic scripts, and so could typewriters. And so could computers: in fact, Unicode for Indic scripts takes inspiration from ISCII, an Indian standard extending ASCII for encoding scripts on computers, that dates back to 1983.
I downvoted you for this: our guidelines require you to assume the strongest interpretation of the comment you are responding to. From my perspective, you’re comment is needlessly reactionary to a thoughtful, well stated question.
Sorry to include this lecture, but it helps my conscious to avoid drive-by downvoting. I’m very much open to a rebuttal.
(This is a different part of the guidelines than "respond to the strongest plausible interpretation" BTW: the comment very explicitly suggests non-Latin scripts not having had "much printing press and typewriter use", so I don't think there's much scope for interpretation there. And in fact, the idea is inherently a condescending and colonialist one—of Western superiority over unsophisticated "those" people—even though it may well be held innocently and stated as such. But I do take the point about not responding to the person, and responding kindly to the idea instead. Will try harder, as I think I usually do.)
(BTW "reactionary" means something specific…)
• U+0915: DEVANAGARI LETTER KA (क)
• U+093F: DEVANAGARI VOWEL SIGN I (ि)
gives कि where the vowel sign not only attaches to the letter, but (in this case) happens to be to the left of it.Similarly, the sequence
• U+0915: DEVANAGARI LETTER KA (क)
• U+094D: DEVANAGARI SIGN VIRAMA ( ्)
• U+0937: DEVANAGARI LETTER SSA (ष)
• U+0941: DEVANAGARI VOWEL SIGN U ( ु)
gives क्षु where the shapes of the "characters" are indeed modified based on which characters appear next to them.This is kind of touched upon in the second section in the original post titled "Indic scripts": https://manishearth.github.io/blog/2017/01/15/breaking-our-l...
This is handled in traditional printing (and in modern Opentype fonts) by simply having separate pieces of type for every possible combination of consonants (there are a few hundred, see e.g. https://en.wikipedia.org/w/index.php?title=Devanagari_conjun...), and then having separate pieces of type for vowel signs (which have to have different lengths depending on the width of the consonant cluster, but a handful of lengths for each will do), the same way any good Latin-script font will include ligatures like fi, and will also include separate glyphs for the accent marks (things like ´ ` ¨ ˆ ˜ which can be placed over letters to give é à ü î ñ and so on).
Incidentally, Gutenberg's very first Bible itself used nearly 300 letter forms, including ligatures and variants of several letter (see e.g. https://www.oldbookappreciator.com/libraryofhistoricaltype/e... )
(About "cursive" vs "separate" letters: conceptually it's not such a big difference, because there are only finitely many letters/characters in the script, so as long as you carefully specify the ending position of each glyph for each possible choice of following glyph, you produce the appearance of everything being joined. Indeed that's how traditional printing and modern fonts do it too. And for that matter, cursive fonts in English: try playing with https://fonts.google.com/specimen/Cedarville+Cursive which if it had been made with slightly better kerning rules would likely modify shapes of letters depending on adjacent letters. Maybe one of these fonts do it? https://fontsgeek.com/search/?q=cursive )
Basically allow everything except some separators, most control chars, and some lookalike characters (which have to be updated as more characters are added to Unicode). It's not as clean as I'd like, but it's at least manageable this way.
https://github.com/kstenerud/concise-encoding/blob/master/ce...
[At least in Chinese] Characters represent parts of words. Some basic words consist of just one character, but others consist of several characters. Characters impart a mood to the word -- the set of words in which a character appears generally have similar meanings. I think the closest thing metaphor would be to describe characters as latin roots.
Useful as long as "represent" means internal in-process use only. You're supposed to never save them into documents or send them between systems.
These are valid C-style identifiers if you extend those rules to Unicode, and editors have no problem with them:
fish_3
рыбы_3
So are these, despite the combination of left-to-right and right-to-left characters:
سمك_٣ (Eastern Arabic numeral for three, note how it appears to come first to an LTR reader.)
سمك_3 (Western Arabic numeral for three)
Problem is, editors really have trouble dealing with mixed-direction words like this. The caret stops pointing to the correct part of the word, the home and end keys become useless, etc. It's an odd situation where the lexer can handle these edge cases well but I cannot think of how you'd actually get them into the compiler.
English is almost universally spoken among western college graduates, and it is the same character set. But I can’t imagine the difficulty of learning programming compounded by it being in a language you barely understand with a character set you are unfamiliar with. Did Chinese language based programming languages appear? Or is everyone eating the bullet (or perhaps I am overstating the difficulty, hard to get an idea of how much english the average chinese college graduate is exposed to).
Don't use capital ß. While it may be found in a few places on the Web in posts of enthusiasts, it's not used in real life. (While it's based on an old proposal, it has been introduced only recently and it's supported only by recent OSes. At the same time – or rather, even before this –, orthographic reforms have replaced the lower-case variant by "ss" for most use cases. So it's anachronistically lost in retro-futurism, but bare of any retro-futurism charms.)
Fun fact: "eszett" is a denomination specific to Germany, in Austria it's "scharfes s" ("sharp s", or rather, acute s), and the Swiss got rid of it altogether in the first half of the 20th century (1938), generally replacing it by "ss".
Interesting. So maybe Germans will eventually experience a similar double take as I do when reading older (English) texts which used the 'long s', though that may well take a similar ~ 150 yrs.
Regarding the "long s", this has been in use in German writing until the early 20th century, as well. One of the best things about "ß" is that there is no consensus what this actually is. Historically, it's a ligature, probably of a long s and a round s, and it can be found in renaissance Italian cursive (e.g., in samples by Palatino). Also, in 17th century type setting, it can be seen as a compositum of long and round s. This also explains, why there is only a lower-case form, as there is also (mostly) only a lower-case long s. So "eszett" or "ß" are probably historically wrong and it was only in Fraktur (non-Latin broken letter type) that "ß" became split into "s" and "z", and ever since, there is this ambiguity. Both "SZ" and "SS" are viable capitalizations, with the former being more authoritative in the mid-20th century and a strong tendency towards the latter (which is about the only form used nowadays).
(Does the ambiguity matter? Not at all. Mind that probably only a minority knows that the ampersand (&) is a ligature of "et", or that "@" is literarily "at" and that you can write "it" the same way. We're perfectly able to use these things without knowing what they are.)
Regarding letter forms, there are plenty sources of worry in German writing: There's Fraktur in books, mostly from the 18th century until 1940 (contrary to common belief, Fraktur was not the preferred typeface of Nazi-Germany, rather they abolished it), there's Schwabacher typeface, which was Latin characters, but still with broken forms, there had been long hand and short hand Kurent cursive, various national forms and epochs of Latin hand writing, even after WW II, etc. Fun!
The correct abstraction for working with Unicode text is the code point. UTF whatever is an implementation detail and not something to get hung up on. Python is one of the few languages that gets this right.
Languages like Swift give you the ability to iterate over encoded bytes, graphemes, or code points. You can access these properties if you really want to, however, the point is that these properties might not be useful in practice.
I kind-of disagree with the author that grapheme clusters meaningfully solve the problem. I think it's just another level of kicking the can down the road, with its own set of problems. But I think that it's at least closer than code points.
> The correct abstraction for working with Unicode text is the code point.
I am not aware of any scenario where the code point is meaningfully what you want. Feel free to inform me of a case.
> UTF whatever is an implementation detail and not something to get hung up on.
Grapheme clusters are not part of the transmission format.
The code point is the most meaningful thing. I'm confused as to what you think you're parsing when you parse text?
> Grapheme clusters are not part of the transmission format.
Absolutely. The intent of my comment was to point out how most programming languages expose strings as some specific UTF encoding rather than exposing them as a collection of code points - exposing strings as UTF anything is a leaky abstraction. An exception could be made for lower level languages, like C, which define strings as pointers to the encoded data.
Interestingly enough modern tajik is written using the russian cyrillic script, which was sort of forced on them in the post-1915 era. But there is a major resurgence in the use of the Farsi/Dari script and alphabet in modern Tajikistan. It's 95% mutually intelligible with what the Dari that's spoken in Kabul. Just a weird accent.
Urdu is of course a huge deal as it's the default language for Pakistan. People whose first language is something else like Balochi, or Pashto, or other will almost certainly learn modern standard Urdu at school. In addition to Pakistan's extensive use of English, of course.
other fun things: the letter "P" or peh doesn't exist in arabic, but does exist in Farsi. So a pizza would be a bizza, and so on. The farsi alphabet is obviously derived from arabic but has some key differences. There's a stanards body in Iran that has defined the 'normal' farsi keyboard layout and unicode information.
https://en.wikipedia.org/wiki/Pe_(Persian_letter)
https://en.wikipedia.org/wiki/Persian_alphabet
in much more "recent" times than the historical person empire, much of the historical extent of the mughal empire resulted in things like the right-to-left persian alphabet and farsi derived language you see in modern standard urdu. urdu is absolutely chock full of farsi words.
https://en.wikipedia.org/wiki/Mughal_Empire#Language
https://en.wikipedia.org/wiki/Persian_language_in_the_Indian...
as to how this might impact software, forms, database fields and such: there's now a VAST population of people who might prefer to either write their info, name, fill out forms entering urdu into a text entry field, or write stuff out in English text if that's the default language they use on the Internet in modern Pakistan. Or some combination of the two. And both are totally valid. You might have somebody's name written out phonetically in English in a text field and their street address and other details are in Urdu or Farsi. Or the other way around.
So for example, Azure SQL Database is forced to us-english, “dmy” date formatting, and UTC time zone. This cannot be changed.
Just give up on living outside the US or using local time.
So I think you underestimate the sheer scope and bloat needed to accommodate all possible languages and cultures equally well. And if we abandon that idea, perhaps it makes sense to keep things simple and support customization instead. If on the other hand Azure doesn't even allow installing any non-English FTS engine that would be questionable of course.
Microsoft just decided to pretend that globalisation is not a problem, but proceeded to market this platform to foreign enterprises as a good migration target for legacy databases and associated applications.
Similarly if I want a task to run at 6AM local time every day
If a time is in the future you usually want to store it as local time because we don't know what the offset between local and UTC will be.
Yet the box-ticker drones from IT procurement are content with vendors who assume we live in a world without diacritic marks, I guess. Using 8-bit character sets in 2022 is a "brown M&M's" [1] indicator for me. If they can't be bothered to use Unicode, what else they don't care about?
It was like that up until very recently for some gargantuan banks.