It's going to be fun to watch developers wrap their heads around that one...
It's going to be fun to watch developers wrap their heads around that one...
"Türkiye".ToUpperInvariant()
"Türkiye".ToUpper()
TÜRKIYE "Türkiye".ToUpper([System.Globalization.CultureInfo]::GetCultureInfo("tr-TR"))
TÜRKİYE > "Türkiye".capitalize()
UniquenessError("Ankara is already the capital")Also, "Nauru".capitalize() should succeed as it has no official capital.
(tongue firmly in cheek ;)
"Switzerland".capitalize()
"HELVETICA"> Due to its linguistic diversity, Switzerland is known by a variety of native names: Schweiz [ˈʃvaɪts] (German);[note 5] Suisse [sɥis(ə)] (French); Svizzera [ˈzvittsera] (Italian); and Svizra [ˈʒviːtsrɐ, ˈʒviːtsʁɐ] (Romansh).[note 6] On coins and stamps, the Latin name, Confoederatio Helvetica – frequently shortened to "Helvetia" – is used instead of the four national languages.
* https://en.wikipedia.org/wiki/Switzerland
That's also why the ccTLD of Switzerland is .ch.
What you're seeing may be an 'artifact' of that.
"Switzerland".capitalize()
𝔈ℑ𝔇𝔊𝔈𝔑𝔒𝔖𝔖𝔈𝔑𝔖ℭℌ𝔄𝔉𝔗
Maybe my locale is messed up?One of those two is Switzerland, which has the official Latin name "Confoederatio Helvetica", leading zinekeller joke that the capital form was "HELVETICA". See https://en.wikipedia.org/wiki/Name_of_Switzerland .
I pretended my implementation of that non-extent programming language generated "Eidgenossenschaft", see https://en.wikipedia.org/wiki/Eidgenossenschaft . As that's a German word, I decided to use a blackletter typeface, specifically, the Fraktur in Unicode which is meant to encode mathematical alphanumeric symbols, then imply that my locale the reason I got a German word.
Edit: I now understand this to be a joke. But.., given the craziness that I've seen in the past, this doesn't seem out of the realm of plausibility!
And, in theory, two countries could share the same capital (as a condominium which is equal joint territory of both, or some kind of third territory belonging to neither), although I'm not aware that's ever actually happened. But, at the subnational level, it has happened in India – Chandigarh is the joint capital of two Indian states (Punjab and Chandigarh), each of which has it as its capital despite it being in neither of them, as well as being a union territory (and hence also capital of itself).
> "Türkiye".capitalize()
PrivacyError("That's nobody's business but the Turks'") > 'Türkiye'.toLocaleUpperCase('TR')
'TÜRKİYE'
[1] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... "Türkiye".toLocaleUpperCase("tr-TR")
'TÜRKİYE'Blindly calling .ToUpper() on anything is a typical anglo-centric mistake. Just don't use .ToUpper(), shoutcase is ugly anyways ;)
See also: one of the many "100 fallacies programmers assume about natural written language" documents or such.
As long as there is no unicode SS character, we are into the "what color are your bits" problem or tolower needs to be language and word aware.
In .NET the uppercase and lowercase functions are culture aware (with defaults to system settings, which breaks more software than you might think) but not word aware AFAIK.
It turns out there is such a unicode character -- ẞ/ß -- although based on other comments here it looks like it was added fairly recently.
Upper/Lower case stuff just seems to be at an annoying intersection where it has cultural and also programming significance. Or at least, people will use toUpper when they really want some case-insensitive sortable version of the string.
(based on some googling, probably localeCompare is the way to go in javascript at least).
I believe the actual name is Eszett.
By this measure, the English name of “W” would be wrong because it’s not actually a “double-U” but a “double-V”. But at the time of the letter’s formation, U and V were not yet separate letters.
Yes, one that you might make if you were for example, trying to make English text uppercase. Which is why it would be daft for anyone to suggest that their country has two different English spellings depending on the character case.
Speaking of surviving Fraktur ligatures, I’m sorry that a couple of others like tz didn’t make it to Roman. It makes poor ß appear lonely.
[0]: https://www.icao.int/publications/Documents/9303_p3_cons_en....
Bringing the thread back to the topic of this comment section: the ICAO document also calls the digits 0123456789 “Arabic” even though their shapes are closer to the original Hindi (Devanagari) forms than to actual Arabic digits — another “Hindi/Turkey” situation
† Although it’s Germany and of course there exists an obscure Verwaltungsvorschrift according to which you can write the non-machine readable field of the Personalausweis/Pass in lowercase, exactly for this use case. I didn’t know that last time but I fully intend to make some poor civil servants life a slight hell the next time I have to renew.
JavaScript actually seems to be the smart one here - its default .toUpperCase() uses the "locale-insensitive case mappings in the Unicode Character Database".
Thanks for the correction!
I don't think most Java and C# software is desktop apps? Surely in most cases it's the locale selected for the server or VM, which should be consistent?
(I'm not saying it's good coding practice, mind you, but it probably ends up accidentally working in a lot of cases.)
> I'm not saying it's good coding practice, mind you, but it probably ends up accidentally working in most cases
Fully agree. It's still bad practice and I high-five every linter that automatically flags it.
>>> "Türkiye".upper()
'TÜRKIYE'
>>> import locale
>>> locale.setlocale(locale.LC_ALL, "tr_TR.UTF-8")
'tr_TR.UTF-8'
>>> "Türkiye".upper()
'TÜRKIYE' # Expect "TÜRKİYE"
>>> locale.resetlocale() import locale
from contextlib import contextmanager
@contextmanager
def use_locale(*args, **kwargs):
locale.setlocale(*args, **kwargs)
yield
locale.resetlocale()
Contextlib is one of the under-appreciated gems in Python: https://docs.python.org/3/library/contextlib.html#contextlib...Even despite the:
> Absolutely.
My absolute least favorite response to this is:
> It's not that hard to write your own though:
I strongly disagree with that sentiment as often these simple things have subtle gotchas, and/or subtle differences in possible net effects creating much bigger problems than their diminutive implementations would suggest.
Or use a programming language that doesn't rely on libc for locales.
> There is no way to perform case conversions and character classifications according to the locale. For (Unicode) text strings these are done according to the character value only, while for byte strings, the conversions and classifications are done according to the ASCII value of the byte, and bytes whose high bit is set (i.e., non-ASCII bytes) are never converted or considered part of a character class such as letter or whitespace.
% env - LC_ALL=en_US.UTF-8 PATH="$PATH" python3
Python 3.8.12 (default, Nov 13 2021, 10:49:08)
[Clang 11.0.3 (clang-1103.0.32.62)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> from icu import UnicodeString, Locale
>>> s = b'the Republic of T\xc3\xbcrkiye'
>>> s = s.decode()
>>> s
'the Republic of Türkiye'
>>> lc = Locale("TR")
>>> s = UnicodeString(s)
>>> s
<UnicodeString: 'the Republic of Türkiye'>
>>> s = s.toUpper(lc)
>>> s
<UnicodeString: 'THE REPUBLİC OF TÜRKİYE'>
>>> s = str(s)
>>> s
'THE REPUBLİC OF TÜRKİYE'
>>> s.encode()
b'THE REPUBL\xc4\xb0C OF T\xc3\x9cRK\xc4\xb0YE'
>>> s = UnicodeString(s)
>>> lc = Locale("CN")
>>> s = s.toLower(lc)
>>> s
<UnicodeString: 'the republi̇c of türki̇ye'>
>>> s = '壹,貳,參,肆,伍,陸,柒,捌,玖,拾,佰,仟,萬'
>>> s.encode()
b'\xe5\xa3\xb9,\xe8\xb2\xb3,\xe5\x8f\x83,\xe8\x82\x86,\xe4\xbc\x8d,\xe9\x99\xb8,\xe6\x9f\x92,\xe6\x8d\x8c,\xe7\x8e\x96,\xe6\x8b\xbe,\xe4\xbd\xb0,\xe4\xbb\x9f,\xe8\x90\xac'
>>> s = UnicodeString(s)
>>> s
<UnicodeString: '壹,貳,參,肆,伍,陸,柒,捌,玖,拾,佰,仟,萬'>
>>> s = s.toLower(lc)
>>> s
<UnicodeString: '壹,貳,參,肆,伍,陸,柒,捌,玖,拾,佰,仟,萬'>
* https://pypi.org/project/PyICU/edit: Oh neat there's a hansfin keyword that I had not known about.
https://unicode-org.github.io/icu/userguide/locale/#keywords
https://stackoverflow.com/questions/6224177/how-to-convert-e...
Interestingly, `"türki̇ye".title()` does _not_ work correctly, returning `"Türki̇Ye"`, presumably because the "title-case" algorithm incorrectly detects \xcc\x87 as punctuation. Not sure if this has been fixed in 3.10 or 3.11.
Python 3.9.13 (main, May 24 2022, 21:13:51)
[Clang 13.1.6 (clang-1316.0.21.2)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> s = 'TÜRKİYE'
>>> s.lower()
'türki̇ye'
>>> s.lower().upper()
'TÜRKİYE'
>>> s.lower().title()
'Türki̇Ye'
>>> s.lower().encode()
b't\xc3\xbcrki\xcc\x87ye
Edit: It turns out that this behavior is documented [0], and the more-correct routine is `string.capwords` [1]:> The algorithm uses a simple language-independent definition of a word as groups of consecutive letters. The definition works in many contexts but it means that apostrophes in contractions and possessives form word boundaries, which may not be the desired result ... The string.capwords() function does not have this problem, as it splits words on spaces only.
[0]: https://docs.python.org/3/library/stdtypes.html#str.title
[1]: https://docs.python.org/3/library/string.html#string.capword...
Python 3.8.10 (default, Mar 15 2022, 12:22:08)
[GCC 9.4.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> short_name_lower = "Türkiye"
>>> short_name = "TÜRKİYE"
>>> short_name.encode() ## The Ü and İ are UTF-8 chars
b'T\xc3\x9cRK\xc4\xb0YE'
>>> short_name_lower.encode() ## Note only the ü is special, i is just ASCII
b'T\xc3\xbcrkiye'
>>> short_name.lower()
'türki̇ye'
>>> short_name_lower.lower()
'türkiye'
>>> short_name.lower() == short_name_lower.lower() ## Looks the same, but it isn't
False
>>> short_name.lower().encode() ## The i has extra \xcc\x87 here
b't\xc3\xbcrki\xcc\x87ye'
>>> short_name_lower.lower().encode() ## The i doesn't have the extra here
b't\xc3\xbcrkiye'
>>> short_name_lower.upper().encode() ## So this is wrong too, since its just an ASCII i to start
b'T\xc3\x9cRKIYE'
Edit: Formatting> In Turkish, the character “i” becomes “İ” when capitalized, while the “ı” (a Turkish-specific character) becomes “I” (which looks just like the Latin upper case “I”).
> The out-of-the-box capitalization method implemented by developers or by localization tools by default is often the standard ‘toUpper()’, which doesn’t follow language-specific rules and will convert the “i” into an “I”. As for the lower case “ı”, it will simply fail to capitalize it at all. This will result in a very strange looking text in the game with uncapitalized characters and wrongly capitalized ones.
https://news.ycombinator.com/item?id=32076177
https://unicode-org.github.io/icu/userguide/transforms/casem...
<img src="/img/turkiye.gif" alt="Turkey" />
No, it will not. Nobody gives a damn. Nobody will implement it.
I am still trying to get people to correctly denote beginning and end of a day (much more useful, practical). Everybody I work with seems to be bent on using 23:59:59 as the end of the day rather than start of next day. Explaining that there is no 1s delay between end of one day and start of the next isn't helping either.
When you say "Thursday" you mean entire 24h period. It is not a point in time, it is a label for a span of time.
But when you say "1 pm" you don't mean an entire hour, you mean a point in time that is more or less precisely 1pm.
It seems people extend the first model to a lot of cases where it is not correct. And so for many people their mental model of 23:59:59 is a label for a span of time that is one second in length.
If your model is that a day consists of 86400 "seconds", each second being a span of time of length of one second with a label like 12:37:28, then 23:59:59 is the last second of the day and 00:00:00 is the first second of the next day.
If there's one person in the queue to the checkout, then span-wise, that person is the start of the queue and the end of the queue. But if the start and end are the same, what is the length of the queue?
(If you still insist it's one, you'll not be able to answer the question where the start and end of an empty queue is. There is no first or last person!)
Just to spell it out, some days have 25 hours or 23 hours due to daylight savings, or 24 hours and 1 second because of leap seconds, etc.
```TSQL select cast('2022-07-12 24:00:00' as datetime) ``` >>> The conversion of a varchar data type to a datetime data type resulted in an out-of-range value.
```python from datetime import datetime d: datetime = datetime(2022, 7, 12, 24, 0, 0) ``` >>> ValueError: hour must be in 0..23
As someone whose last name contains a character from the same Unicode subset (ć), it's often a white square or just flat-out removed from my last name completely.
A realistic scenario is an Australian writing French poetry while on holidays in Turkey. Now you have an OS GUI with en-GB as the language, licensing and date/number formatting as per the AU region, French spelling dictionary, a US-101 keyboard layout, Turkey as the location, and GMT+3 as the time zone.
It's a rare piece of software that can handle this. Few vendors have staff that have even heard of such exotic places.
There are still people... many people... that deny the existence of places outside of the United States of America. Such filthy, heathen locations are surely a thing of myth, or legend!
Places where dates are formatted with the days before the month, followed by the year in some sort of weird, unnatural order.
Nations that have fallen into the trap of some sort of mass hallucination, or shared dream of common measurement units. Some sort of... metric, for space, time, and matter. Maybe they've been watching too much Star Trek!
Multicultural countries where strange unions of races are commonplace, and couples may want to watch Netflix in one language, but have subtitles in a different language. Neither of which are English for the hearing-impaired! Surely, nobody but the deaf are unable to comprehend the universal English language! Not to mention that such taboo couplings are, thankfully, still banned on this stream of holy virtue and shall not be permitted by the data scientists that have declared: "Your union is a statistically negligible!"
Two fun anecdotes my mother has recounted when we were all in Denver for ~5 months in Denver 1996:
1. Asked where we were from on one occasion (after hearing an Australian accent): “Australia.” Response: “Oh, did you come by bus?” Now it is possible that there was a misunderstanding there, but mum doesn’t think so.
2. Of those that didn’t already know, only one person successfully identified where we were from, and that person was deaf. (It’s fascinating to try something like this on video without audio where there’s nothing obvious static to identify nationality: I find I can identify both Australian and American correctly at a rate considerably better than random, without ever having made a study of it.)
Yes, I could begin typing "Unite.." but we all know that ends up on United Arab Emirates, and then when I continue on to "...d Sta..." and it doesn't work, only to find out this country picker uses "USA" or something similar so it breaks the autocomplete muscle memory.
Clearly, if we're the center of the known universe, we could use our power to make it a little easier to enter my billing and shipping information.
Some of use have moved beyond “falsehoods programmers believe about X” to “falsehoods the people in charge believe about X.”
23:59:59.500 is after 23:59:59 when interpreted as an instant, but it is part of 23:59:59 as a duration.
It was because of the way they write/parse integer using dots as separators. (Yes, the real problem was me having forgot to force server and client to use the same locale settings :) )
An old article talking of it : https://blog.codinghorror.com/whats-wrong-with-turkey/
https://unicode-org.github.io/cldr-staging/charts/latest/sum...
PS. It's much easier to type umlauts on Mac. Just hold 'u' to see the variants.
On a Mac now the alphabetic and digit keys do not autorepeat when held but instead pop up little menus of variants. Shifted alphabetic keys get different variants appropriate to their letter. Sadly for TÜRKİYE the single dot capital "i" is not one of them.
At least for my settings, none of the digits get options, nor do they autorepeat.
A completely useful overloading of "long press" of a keyboard key, but completely undiscoverable unless you make long strings of "vvvvvvvvv" as a pointer or some such, then you get a disappointment instead of what you wanted, but you will be enlightened.