There's no optimal or blameless solution here but I think in particular when it comes to the scientist use-case mentioned above, an entire decade is enough to address what is likely a relatively small codebase, it's not just a fault of the language, but also neglect or lack of resources in fields where programming is secondary in nature. (also btw detrimental to the quality of science, as recent cases have shown where programming errors basically invalidated entire studies)
Not really. Just introduce an optional version declaration at the top of the file, which guarantees the old behavior. Then feel free to tinker with the design as much as you want. It's a solved problem and it works well. Sure, you still break people's code if they don't put a version declaration in - but now the fix is a single line.
The fatal flaw in Python's approach was cleaving apart the Python2 and Python3 worlds, making it impossible to mutually import modules. Python2 and Python3 are practically speaking different languages - different syntax, different semantics, different libraries, and a different userbase. As far as the computer is concerned, they are separate entities, with the feeble "2to3" program being the only thing linking them.
You can use Cython to compile your python 2 code into .so objects and then include them in Python 3. Then you can progressively migrate code file by file to Python 3. I didn't do it on a larger project, so maybe there are some gotchas, but it seems to work as far as I could test it.
Regarding the differences in languages Python 2.7 was actually a bridge to migrate to Python 3. That's why all features of 2.7 were backported from 3 (this was also the reason people saying: "why should I migrate to 3, it doesn't offere anything new". It didn't because everything new was backported.
Anyway thanks to this it was actually possible to write code that worked on both Python version. I did it several times. It was actually the easiest when you wrote the code for Python 3 and only used features that 2 had. Add all the `__future__` imports and remaining parts filled with the `six` module.
Yes, and tools and infractstructure should account for that: Python 2 and Python 3 should use different file extensions, and libraries written for python 2 should not just adapt to Python 3 and break backward compatibility, but come _at least_ as a new major version - better with a different library name, as Go lang requires.
And still it would have been possible to support Unicode in python 3 without breaking backwards compatibility. It was possible to support in Common Lisp which was defined as a ANSI standard in 1994, and Common Lisp implementations like SBCL or Schemes compile to far faster code than C Python (while providing many of its flow control affordances, like list comprehensions, or a numeric tower). Why should it not be possible for Python? Apart from the developers not wanting to do it?
A lack of resources for porting and maintaining code to new language versions is totally typical for most of science. There are simply no jobs or incentives to do that, especially if you want to have any kind of career in science.
Further, python developers seemed to think that the bulk of code and maintenance work is in the language implementation and that the many small libraries and scripts are not important. But this is not the case, similar to that in trees, a rather large part of the biomass is not in the trunk, but in the roots and leaves. And the small application scripts and libraries used in science change much more slowly than the main language revisions.
You could argue that this does not matter because most of python's applications are not in science, at least not if you look at the respective amount of money moved. But this fails to take into account that a lot of python's success was so far based on its use in science, especially numerical python and data analysis.
Of course python could have dealt better with backwards compatibility, but the python code I wrote was transfered to version 3 pretty quickly.
Perhaps some chose instead move to other languages.
The thing I think they did better though, is they didn't give people 10 years to fix their sh*t. You either moved on or you stayed with an unsupported version, and people did move on.
For Python 3 the major shifts were in 2015 and 2020. The first year was when PSF announced they won't add any new features to 2.7, and that's when libraries started adding Python 3 support. And 2020 was EOL for Python 2.7.
It doesn't matter how much time you give people to migrate, they will always do it at the very last minute.
Also, 1.8 marked the creation of complete compliance test suite (as Ruby 1.8 became an ISO Standard), and that test suite has been expanded since then allowing for alternative implementations.
Because the test suite is your target, not a hard to interpret 2kLOC spaghetti code C function.
Ruby was then using a version numbering scheme similar to the old Linux version numbering scheme, where odd minor numbers were development releases and even minor numbers were stable releases. In that scheme a odd minor version is equivalent to a SemVer major release. (Ruby switched to SemVer-ish versions with 2.x)
It would be nice to have an alternative to Python that took backwards compatibility seriously. Might be easier to start from Python2 than to go from scratch. You might be able to add a REPL to perl5, but its syntax isn't the most beginner friendly I've ever seen.
Things these languages do have in common with Python3 are:
- support for keyword arguments and optional arguments
- support for objects and interfaces (but not necessarily for inheritance)
- handling of names in scopes and name spaces
- closures and lambdas
- list comprehensions
- arbitrarily long integers, complex numbers, a numeric tower
- an empty sequence or empty string is logically false (with a few nuances)
- low-level bit operations (often down to a very low level such as popcount in C)
- easy to call into C code (or Java code if it runs on the JVM)
- full unicode support
- conditional expressions like (a if boolval else b)
- type annotations
- string formatting forms a mini-language
- print is a function
- hexadecimal, octal, and binary literals
- tools for iteration and first-class functions, like partial application
- support for OOP
In addition to this, the members of the Lisp family usually have as well:
- strong support for functional programming
- full and efficient garbage collection
- persistent data structures
- compile to machine code (native code generation or JIT compilation)
- a powerful macro system
- capability to define custom control constructs
- support for pattern matching to members of compound objects, similar to
(a, b) = name[:2]
- often very good support for parallelismWhat exactly would be suited best depends IMO on the use case.
I think Racket is ideal for easy learning, a platform-independent GUI, and comprehensive documentation, and is also useful fir science applications. It can also interact well with C libraries (or, more generally native libraries with C interface), and it is good for scripting.
Clojure is ideal for server applications with a lot of concurrency and a high performance. Its concurrency support is brilliant and best-in-class, much better than Python.
Common Lisp / SBCL is somewhat under-hyped in comparison to Clojure, yet it has an excellent performance, integrates well with C libraries, has a highly interactive repl, a comprehensive library system (quicklisp), a standard build system (asdf), comprehensive documentation for beginners (common Lisp cookbook), it is very usable for scripting, has standard pthreads-like concurrency primitives, a powerful object system, and then some more.
It's also the reason I never started in the first place. It wasn't in any way clear which the right option was during the period of schism, and it went on for an awfully long time.
With that being said I'm starting to tinker with Python now simply because it has a great ecosystem for NLP and ML. I certainly need the former, although how much of the latter I'll need (possibly none) is unclear.
Regardless, it seems to be the de facto standard language for anything that falls under the broad category of data science to the point where for me to pick up a different language would feel very much like swimming upstream.
(Related, but separate:) Would you have preferred if Python 3 was released under a different name that didn't include "python"?
In python 2 you could specify a unicode literal with u''. So a natural evolutionary option if you had a unicode addiction and wanted python 3 to be unicode native, would be to treat u'' as unicode literals in python 3 as well.
This let's the code run under python 2 still, and then you get unicode in 3. And if you are 3 only, just "blah blah" works as well - nice.
But they didn't do that. They purposely took u'' out to make it hard to support 2 and 3. That is absolute stupidity. In other words, they BROKE compatibility for no benefit - but just to break it.
They did the same thing with byte handling, jumped the io abstraction in - these things all destroyed performance and use cases.
They were spending the time doing other important things. They were making the rational decision to work on something that was important and WASN'T already working. The python 2 code was working fine, it didn't have bugs and it didn't need any new features. Other code either didn't exist and needed to be written or had bugs that needed to be fixed.
It makes sense that changing working code never made it high enough on the priority list to get worked on.
And this old code is also written in a style (usually global variables and mixed concerns and years of sub-ideal bug fixes and patches) which makes it very hard to change or modernize (forget adding threads with all of this global mutable state).
The only point being that old code may superficially work, but in many cases (in my experience) and regardless of language, it all has technical debt and needs work to keep it relevant.
If the situation changes and the original code suddenly becomes a basis for further development then that's a different thing. But cleaning up the code without clear benefit is a waste of resources.
Hats off to original developers. Making code that serves its purpose for 25 years is quite an achievement.
All code rots. If you are lucky to never need to fix it or change it, that's great, but that is (in my experience) extremely rare over time spans measured in decades.
Sometimes a code base runs for decades not because it was particularly robust or well-written, but because it is too painful to replace it due to its complexity or the long ago loss of the knowledge needed to work on it.
Some Python users understand compiler flags, but there's also a big community of scientific developers who just want a working language.
Python 2 -> 3 created a lot of problems for them without making their lives easier or more productive.
Perhaps Python 3 should have been called something entirely different, and P2 could have been handed over to a new set of maintainers who weren't going to EOL it.
“The Fortran standards committee generally refuses to break backward compatibility when Fortran is updated. This is a good thing (take that, Python), and code written decades ago can still be compiled fine today.” -- http://degenerateconic.com/backward-compatibility/
But seriously:
- C never did such change, it doesn't make as much sense for it to do it since it is low level - C has a type static typing, so if they did, they could make compiler refuse to compile until types are right. Python has dynamic typing so these kinds of things show up when running, you can use something like mypy, but unlike C it is optional.
BTW: I saw people using Go as an example that there was no need to add unicode support like in python, and simply store everything as UTF-8. Those people are forgetting that Go already has clear distinction between text and bytes (via string, bytes, runes types), how the language stores text is just an implementation detail, since normally there shouldn't there be access to the internal representation of the text anyway.
Sure, there is some amount of C code from 25 years ago that works unchanged, but there's also a good bit of Python code from 25 years ago that works unchanged. (I just tried a few of the examples in the Python 1.3 tarball, released in 1995, and it seems like most of them will work with very minor changes.)
- Warnings recommending the use of parens in code like `if (a || b && c)`. The original code is (probably) not broken, it's just that GCC thinks that it's error-prone to write conditionals that way (I agree). There's also a similar warning for code like `while (a = foo())` where it asks you to add parens to make sure that you didn't use = instead of == by mistake.
- Warnings for unmarked fall-throughs in `switch`. Again, not broken if the fall-through was intentional. That's a fairly new warning, so a lot of all code will trigger it.
- Warnings for some const to non-const pointer conversions (in particular for literal strings in C++). There you could argue that the original code is bogus but as long as the pointer is not written to it's fine. Const-correctness wasn't really a thing in C and C++ for a long time, a lot of legacy code plays fast and loose with const qualifiers, but that doesn't necessarily mean that the code is broken.
- Warnings about unused variables, in particular variables being written to but not read. As compiler get smarter they tend to catch more of those.
Glibc and GCC introduce incompatible changes all the time.
Writing some code and not have it executable a dozen years later is a terrible terrible situation. And it wouldn't be this way if Python 2 wasn't EOL.
I'm going to always need to keep a copy of Python 2 around. Always. And if I'm going to change old, working code, I'm probably going to rewrite it in a new language.
But yet I get why things are the way they are. Heck, it was only last year that I switched to using Python 3. I use it so rarely that it was just easier to stick with the Python 2 which was already installed on my system.
Ultimately, I stopped writing Python: these days I write Common Lisp for personal projects and Clojure/JavaScript for everything else.