Blindly calling .ToUpper() on anything is a typical anglo-centric mistake. Just don't use .ToUpper(), shoutcase is ugly anyways ;)
See also: one of the many "100 fallacies programmers assume about natural written language" documents or such.
Blindly calling .ToUpper() on anything is a typical anglo-centric mistake. Just don't use .ToUpper(), shoutcase is ugly anyways ;)
See also: one of the many "100 fallacies programmers assume about natural written language" documents or such.
Speaking of surviving Fraktur ligatures, I’m sorry that a couple of others like tz didn’t make it to Roman. It makes poor ß appear lonely.
[0]: https://www.icao.int/publications/Documents/9303_p3_cons_en....
Bringing the thread back to the topic of this comment section: the ICAO document also calls the digits 0123456789 “Arabic” even though their shapes are closer to the original Hindi (Devanagari) forms than to actual Arabic digits — another “Hindi/Turkey” situation
† Although it’s Germany and of course there exists an obscure Verwaltungsvorschrift according to which you can write the non-machine readable field of the Personalausweis/Pass in lowercase, exactly for this use case. I didn’t know that last time but I fully intend to make some poor civil servants life a slight hell the next time I have to renew.
As long as there is no unicode SS character, we are into the "what color are your bits" problem or tolower needs to be language and word aware.
In .NET the uppercase and lowercase functions are culture aware (with defaults to system settings, which breaks more software than you might think) but not word aware AFAIK.
It turns out there is such a unicode character -- ẞ/ß -- although based on other comments here it looks like it was added fairly recently.
Upper/Lower case stuff just seems to be at an annoying intersection where it has cultural and also programming significance. Or at least, people will use toUpper when they really want some case-insensitive sortable version of the string.
(based on some googling, probably localeCompare is the way to go in javascript at least).
Yes, one that you might make if you were for example, trying to make English text uppercase. Which is why it would be daft for anyone to suggest that their country has two different English spellings depending on the character case.
JavaScript actually seems to be the smart one here - its default .toUpperCase() uses the "locale-insensitive case mappings in the Unicode Character Database".
Thanks for the correction!
I don't think most Java and C# software is desktop apps? Surely in most cases it's the locale selected for the server or VM, which should be consistent?
(I'm not saying it's good coding practice, mind you, but it probably ends up accidentally working in a lot of cases.)
> I'm not saying it's good coding practice, mind you, but it probably ends up accidentally working in most cases
Fully agree. It's still bad practice and I high-five every linter that automatically flags it.
I believe the actual name is Eszett.
By this measure, the English name of “W” would be wrong because it’s not actually a “double-U” but a “double-V”. But at the time of the letter’s formation, U and V were not yet separate letters.