For text files, maybe, but various APIs like the Window API and the Java String API still use UTF-16.
UTF-8 dependence is also a major pain for many where the local character set conflicts with UTF-8. For example, there's still a lot of Japanese files out there in SJIS that need to be decoded accordingly. The country of Myanmar officially switched to unicode less than two years ago so if you still need to operate on older data, you're going to need to support their old character set.
UTF-8 as a fixed encoding only works if you manage to write mappers from and to alternative character set for practically any language outside US English. Instead of breaking compatibility with most libraries, python3 would have broken compatibility with most libraries and a few countries instead.
Just like the rest of the world has to deal with three countries refusing to switch to metric, python3 needed to deal with countries refusing to switch to UTF8.
Huh? I'm using UTF-8 exclusively for string data since around 20 years in C and C++ and never had to deal with language specifics (also true for non-European languages, we need to deal with various East Asian languages, and Arabic for instance). You need to convert from and to operating system specific encodings when talking to OS APIs (like UTF-16 on Windows), but that's it (and this is not language specific, code pages are an "8-bit pseudo-ASCII" thing that's irrelevant when working with UTF encodings.
When dealing with "vintage" text files with older language-specific encodings, you need to know the encoding/codepage used in those files anyway, and do the conversion from and to UTF-8 while writing or reading such files. Those conversions shouldn't be hardwired into the "string class".
From a European perspective, this sounds very unlikely. Sure, you may have to deal with deprecated _encodings_, but I’d like to hear about mainstream languages with writing derived from the Latin alphabet, that aren’t supported by UTF-8
Or they could have fixed setdefaultencoding or give us a way to set the default encoding https://stackoverflow.com/questions/3828723/why-should-we-no...
I don't buy the "discouragement" part there, if anything they could have made it mandatory or at least it set it to UTF-8
> For example, there's still a lot of Japanese files out there in SJIS that need to be decoded accordingly.
Yes but you would have to work on those cases anyway and ASCII would have made it blow anyway. But convert it to UTF-8/16 and it works.
EDIT: the reason is apparently that "(setdefaultencoding) will allow these to work for me, but won't necessarily work for people who don't use UTF-8. The default of ASCII ensures that assumptions of encoding are not baked into code"
Really. I can't explain my anger at how this is such an idiotic excuse. Yes, your program will fail if you use Latin-1 encoding, duh. Configure your environment correctly and it will work. Sounds like the kind of pedantry that made Guido quit over the walrus operator
Either way, it's no big deal. There are no excuses for Python 3. Some people are just stubborn.