Text normalization in Go
blog.golang.org
blog.golang.org
in NFC form, "base characters and modifiers are combined into a single rune whenever possible"
the interesting detail is "whenever possible": since NFC works by first decomposing, and then recomposing... there're some cases in which if you run NFC normalization on it, the characters will remain decomposed
an example is 𝅘𝅥𝅮 (U+1D160) which its normalized composed form is made of 3 different codepoints
I tried to look at the algorithm for generating the composition table, and it seems it's generated from the decomposition table... if that's so, I can't understand how it could happen that some code points have an NFC form longer than 1
more details: http://stackoverflow.com/questions/17897534/can-unicode-nfc-...
does anyone knows the cause behind this?
2. It's not compression, it's normalisation. So it's not compose everything you can. I cannot tell you exact the algorithm off the top of my head, but:
the reason for U+1D160 — it's in CompositionExclusions list.
http://unicode.org/reports/tr15/#Primary_Exclusion_List_Tabl...
> When a character with a canonical decomposition is added to Unicode, it must be added to the composition exclusion table if there is at least one character in its decomposition that existed in a previous version of Unicode. If there are no such characters, then it is possible for it to be added or omitted from the composition exclusion table. The choice of whether to do so or not rests upon whether it is generally used in the precomposed form or not.
fmt.Println(strings.Replace("multiple cafe\u0301", "cafe", "cafes", 1)) // multiple cafeś s = "We went to eat at multiple cafe\u0301"
"We went to eat at multiple café"
s.replace('cafe', 'cafes');
"We went to eat at multiple cafeś"
Interesting thing is when the text is copy-pasted backspacing first deletes the accent. At least in chrome.Node.js - https://github.com/walling/unorm YMMV, but looks good.
It can also serve as a polyfill for the eventual http://people.mozilla.org/~jorendorff/es6-draft.html#sec-str...
This is useful for example, to ensure that users don't try and spoof each other's usernames. Simply create and store a skeleton string for each username, and keep a unique constraint on it
I'm coming at this from a comment spam point of view, not usernames, btw.
In javascript:
"ß".toUpperCase().length !== "ß".length;
Does weiss == weiß ?But \u00DF appears to be a special case, as there's no uppercase for it. If I had to guess, I'd say it should return \u00DF. I mean, if I uppercase "+", do I expect something else back? Doubtful.
• Special Casing: Lowercase: 00DF [ ß ] Uppercase: 0053 0053 [ S S ] Titlecase: 0053 0073 [ S s ]
• NamesList: = Eszett • German • uppercase is "SS" • in origin a ligature of 017F and 0073 → (greek small letter beta - 03B2) → (latin capital letter sharp s - 1E9E)
(in origin a ligature of 017F and 0073 is not undisputed)
U+1E9E (LATIN CAPITAL LETTER SHARP S ẞ) is not officially allowed in German orthography • NamesList: • lowercase is 00DF → (latin small letter sharp s - 00DF) • Designated in Unicode 5.1
Yes and no. The swiss would write the former, other German speaking (writing) countries would write the latter. It is incorrect in Germany (after ie, au, eu, ... you must not write ss, unless it's a name, such as the city Neuss)
The upper case of weiß would be WEISS. But it's hard from the upper case WEISS to determine if the lower case is weiss or weiß. (This is why one should never write people's names in bibliographies in small caps.)
Technically, Unicode has a capital sharp s since 5.1.0, so we could write WEIẞ.
And I am glad that U+1E9E (LATIN CAPITAL LETTER SHARP S) is not official part of German orthography.
2) Ligatures (e.g. ffi as ffi) are deprecated in Unicode;
3) weiss ≠ weiß in any sense
Edit: 4) x.toUpperCase().length ≢ x.length, upcasing can change length;
5) length in JS (in 100000 other languages) count codepoints (at best), it's useful for nothing here
You need a case folding function/method to check for this.
For eg. in Perl, see the fc function - http://perldoc.perl.org/functions/fc.html
fc("weiss") eq fc("weiß"); # true <link rel="alternate" type="application/atom+xml" title="blog.golang.org - Atom Feed" href="http://blog.golang.org/feed.atom"/>