There's no optimal or blameless solution here but I think in particular when it comes to the scientist use-case mentioned above, an entire decade is enough to address what is likely a relatively small codebase, it's not just a fault of the language, but also neglect or lack of resources in fields where programming is secondary in nature. (also btw detrimental to the quality of science, as recent cases have shown where programming errors basically invalidated entire studies)
Not really. Just introduce an optional version declaration at the top of the file, which guarantees the old behavior. Then feel free to tinker with the design as much as you want. It's a solved problem and it works well. Sure, you still break people's code if they don't put a version declaration in - but now the fix is a single line.
The fatal flaw in Python's approach was cleaving apart the Python2 and Python3 worlds, making it impossible to mutually import modules. Python2 and Python3 are practically speaking different languages - different syntax, different semantics, different libraries, and a different userbase. As far as the computer is concerned, they are separate entities, with the feeble "2to3" program being the only thing linking them.
You can use Cython to compile your python 2 code into .so objects and then include them in Python 3. Then you can progressively migrate code file by file to Python 3. I didn't do it on a larger project, so maybe there are some gotchas, but it seems to work as far as I could test it.
Regarding the differences in languages Python 2.7 was actually a bridge to migrate to Python 3. That's why all features of 2.7 were backported from 3 (this was also the reason people saying: "why should I migrate to 3, it doesn't offere anything new". It didn't because everything new was backported.
Anyway thanks to this it was actually possible to write code that worked on both Python version. I did it several times. It was actually the easiest when you wrote the code for Python 3 and only used features that 2 had. Add all the `__future__` imports and remaining parts filled with the `six` module.
Yes, and tools and infractstructure should account for that: Python 2 and Python 3 should use different file extensions, and libraries written for python 2 should not just adapt to Python 3 and break backward compatibility, but come _at least_ as a new major version - better with a different library name, as Go lang requires.
And still it would have been possible to support Unicode in python 3 without breaking backwards compatibility. It was possible to support in Common Lisp which was defined as a ANSI standard in 1994, and Common Lisp implementations like SBCL or Schemes compile to far faster code than C Python (while providing many of its flow control affordances, like list comprehensions, or a numeric tower). Why should it not be possible for Python? Apart from the developers not wanting to do it?
A lack of resources for porting and maintaining code to new language versions is totally typical for most of science. There are simply no jobs or incentives to do that, especially if you want to have any kind of career in science.
Further, python developers seemed to think that the bulk of code and maintenance work is in the language implementation and that the many small libraries and scripts are not important. But this is not the case, similar to that in trees, a rather large part of the biomass is not in the trunk, but in the roots and leaves. And the small application scripts and libraries used in science change much more slowly than the main language revisions.
You could argue that this does not matter because most of python's applications are not in science, at least not if you look at the respective amount of money moved. But this fails to take into account that a lot of python's success was so far based on its use in science, especially numerical python and data analysis.
Of course python could have dealt better with backwards compatibility, but the python code I wrote was transfered to version 3 pretty quickly.
Perhaps some chose instead move to other languages.
The thing I think they did better though, is they didn't give people 10 years to fix their sh*t. You either moved on or you stayed with an unsupported version, and people did move on.
For Python 3 the major shifts were in 2015 and 2020. The first year was when PSF announced they won't add any new features to 2.7, and that's when libraries started adding Python 3 support. And 2020 was EOL for Python 2.7.
It doesn't matter how much time you give people to migrate, they will always do it at the very last minute.
Also, 1.8 marked the creation of complete compliance test suite (as Ruby 1.8 became an ISO Standard), and that test suite has been expanded since then allowing for alternative implementations.
Because the test suite is your target, not a hard to interpret 2kLOC spaghetti code C function.
Ruby was then using a version numbering scheme similar to the old Linux version numbering scheme, where odd minor numbers were development releases and even minor numbers were stable releases. In that scheme a odd minor version is equivalent to a SemVer major release. (Ruby switched to SemVer-ish versions with 2.x)
It would be nice to have an alternative to Python that took backwards compatibility seriously. Might be easier to start from Python2 than to go from scratch. You might be able to add a REPL to perl5, but its syntax isn't the most beginner friendly I've ever seen.
Things these languages do have in common with Python3 are:
- support for keyword arguments and optional arguments
- support for objects and interfaces (but not necessarily for inheritance)
- handling of names in scopes and name spaces
- closures and lambdas
- list comprehensions
- arbitrarily long integers, complex numbers, a numeric tower
- an empty sequence or empty string is logically false (with a few nuances)
- low-level bit operations (often down to a very low level such as popcount in C)
- easy to call into C code (or Java code if it runs on the JVM)
- full unicode support
- conditional expressions like (a if boolval else b)
- type annotations
- string formatting forms a mini-language
- print is a function
- hexadecimal, octal, and binary literals
- tools for iteration and first-class functions, like partial application
- support for OOP
In addition to this, the members of the Lisp family usually have as well:
- strong support for functional programming
- full and efficient garbage collection
- persistent data structures
- compile to machine code (native code generation or JIT compilation)
- a powerful macro system
- capability to define custom control constructs
- support for pattern matching to members of compound objects, similar to
(a, b) = name[:2]
- often very good support for parallelismWhat exactly would be suited best depends IMO on the use case.
I think Racket is ideal for easy learning, a platform-independent GUI, and comprehensive documentation, and is also useful fir science applications. It can also interact well with C libraries (or, more generally native libraries with C interface), and it is good for scripting.
Clojure is ideal for server applications with a lot of concurrency and a high performance. Its concurrency support is brilliant and best-in-class, much better than Python.
Common Lisp / SBCL is somewhat under-hyped in comparison to Clojure, yet it has an excellent performance, integrates well with C libraries, has a highly interactive repl, a comprehensive library system (quicklisp), a standard build system (asdf), comprehensive documentation for beginners (common Lisp cookbook), it is very usable for scripting, has standard pthreads-like concurrency primitives, a powerful object system, and then some more.
It's also the reason I never started in the first place. It wasn't in any way clear which the right option was during the period of schism, and it went on for an awfully long time.
With that being said I'm starting to tinker with Python now simply because it has a great ecosystem for NLP and ML. I certainly need the former, although how much of the latter I'll need (possibly none) is unclear.
Regardless, it seems to be the de facto standard language for anything that falls under the broad category of data science to the point where for me to pick up a different language would feel very much like swimming upstream.
(Related, but separate:) Would you have preferred if Python 3 was released under a different name that didn't include "python"?
In python 2 you could specify a unicode literal with u''. So a natural evolutionary option if you had a unicode addiction and wanted python 3 to be unicode native, would be to treat u'' as unicode literals in python 3 as well.
This let's the code run under python 2 still, and then you get unicode in 3. And if you are 3 only, just "blah blah" works as well - nice.
But they didn't do that. They purposely took u'' out to make it hard to support 2 and 3. That is absolute stupidity. In other words, they BROKE compatibility for no benefit - but just to break it.
They did the same thing with byte handling, jumped the io abstraction in - these things all destroyed performance and use cases.