Python 3’s Unicode support has its issues, e.g. try using `len()` on a string containing surrogate pairs. But it generally works and it’s internally consistent, which is far more than can be said for Python 2’s unholy shambles where str and unicode could be unsafely mixed for all kinds of exciting data corruption and errors. Python 2-to-3 was a move in the
right direction. I switched early, and the number of encoding-related problems I’ve had in my Python 3 code has been a fraction of those under Python 2.
Article author’s clickbait-y hysterics are a ridiculous overreaction to the usual IO hassles when working with outside data in arbitrary legacy encodings. But Python 3.1 was released a decade ago and we’re now up to 3.8, so it’s not like folks who still deal with legacy data haven’t had plenty time to file tickets and patches on Python’s famously grotty stdlibs. e.g. Adding optional `path_encoding` and `data_encoding` parameters to `ZipFile()`, similar to `open(f,encoding='…')`, would easily address the described problem, no sturm-und-drang required.
Or heck, in the time it took him to write his post, he could’ve easily knocked out a Python 2 script that rewrites all his legacy .zip files to use modern UTF8 instead. But some folks just prefer complaining, I guess.
Obligatory: https://i.ytimg.com/vi/tJ-LivK4-78/hqdefault.jpg
(Honestly, from the HN post title I thought it was going to about something Python 3 really did fuck up, like the 2-to-3 migration which has been an absolute sucking swamp of unnecessary complexity and make-work for years, and a fine demonstration of how not to manage lifecycle.)