My main difficulty was having targets that ran python 2 and wouldn't get an official python 3 package unless a miracle happened. Meanwhile most systems supporting python 3 also had a python 2 package. Ended up porting a few scripts to a statically linked c++ binary, those should keep on working until we get 128 bit systems that drop support for 64 bit binaries.
In my experiences trying to deploy software at scale (10^4~5), statically linked binaries are the only acceptable form of software.
We're talking about this versus the alternative of potentially not being able to deploy anything. Then you're dead.
In both cases you needed a build pipeline anyway...unless you're operating in a world where "it works on my machine" is a good enough answer.
Curiously Linux, despite having an extremely stable user space interface, suffers from a high degree of user space incompatibility. This starts with glibc and its components and ends with graphical toolkits. Go seems to be very successful because it ditches all of that crud and goes for the stable interface instead.
In some spaces (ML) you're going to have to ship some Python, but what I was saying was that I would rather not.
My role the past few years has been some mix of deployment engineering & systems engineering and the conclusion that I quickly came to was that deploying scripting languages like Python/Ruby/etc is the realm of assholes. ...which was a funny lesson because mainly through my career I've worked in smaller scale shops using almost exclusively scripting languages.
If 13 years isn't enough warning to migrate your use case, no amount of warning will help you.
At some point you have to stop saying its to hard to migrate.
Anyway, how hard would it have been to have a "from __past__ import old_strings" that could have worked for the first few releases to allow the single biggest issue to be smoothed over universally and then fixed file by file under the Python 3 environment? With that in place a lot of shops could have just migrated on day 1 and then iteratively worked to finish the job rather than delaying for so long.
Personally I still have a sour taste of the urllib migration. All the imports were moved around for no reasons, completely breaking all usage of urllib and any "import urllib", with no way to fix in sight.
The solution came with six a million years later, adding a hundred alias in six.moves.urllib.somefunction that dispatch to the right place.
And this attitude and lack of empathy ("just use the automated tool, lol") is a big part of why the blasted thing took so long.
At the end of the day, they had anywhere from 8 to 12 years to migrate. In the tech world, that's a millennia. Most products and services will give you months when they deprecate or sunset.
There's 0 difference between the states that were putting a call out for COBOL programmers to deal with their disaster of an unemployment system and the professor running some python script from 2003. Their lack of responsible ownership is their own fault.
Python decided instead that their customers were of no value, and all the millions of lines of existing Python shouldn't continue to run, even though it would be trivial for them to continue supporting 2.x syntax along with 3.x. I think it is one of the most insane decisions ever made by a mainstream programming language.
So a systems programming language, 13 years younger than COBOL, or 10 years, given that the work in C started in 1969.
http://cm.bell-labs.co/who/dmr/chist.html
> C came into being in the years 1969-1973, in parallel with the early development of the Unix operating system; the most creative period occurred during 1972. Another spate of changes peaked between 1977 and 1979
But lets make you a favour and consider 1972, it makes C a 48 years old systems programming language, only surpassed by NEWP (1961) as oldest systems programming language still in use in 2021.
Maybe it is about time to start talking about C the same way people talk about COBOL.
My last interaction, which felt similar, and put me off commenting for a while, was with someone who claimed that "most concert pianists and serious competitors have absolutely gigantic hands" then moved the goalposts around so they could be Right.
https://news.ycombinator.com/item?id=25174394
I guess I should just let vague claims lie. But, well, when something sounds wrong, it sounds wrong.
And could fit a dictionary definition on the comment to prove my point, but then we would really be moving goal posts by then.
But that means that any language revision needs to work hard at making the transition easy and incremental. 2to3 was and is laughable; at no point was it a reasonable solution.
It's okay to say "we need maintenance money". But no one has an unlimited budget. People complain about the 2->3 transition because (1) it was extraordinarily steep, (2) was not justified, (3) its huge costs are repeatedly denied.
Oh wait, we pay computing grad students literally 1/10 of their potential salary ($10Ks vs $100Ks). Why could there possibly be a shortage?
-_-
More importantly, it shouldn't be prohibitively expensive to port academic Python code.
The low pay is in exchange for the ability to do research, not to be a discount software engineer.
But graduate schooling is still a big opportunity cost, and not to go all Mark Twain, but sometimes that can get in the way of your education.
* Very low pay.
* Good benefits in a nation with poor safety nets.
* Tuition waivers along the lines of $10-100K/year.
* When the Dr. says jump, you ask how high.
If you view education as an investment, it isn't necessarily bad compared to an ordinary job. But it's kind of like a FAANG company; your experience depends on who you report to.
But seriously, you're right. Grad students do grunt work, that's how it goes. And if an academic Python2 library is widely-used, porting it is important grunt work.
Surely, no serious researcher would let an important tool rot, right?
What is important to the grad students is to produce research papers and to fulfill their mandatory obligations (teaching, project deliverables). And most grad students, even in CS, are not professional software developers anyway. Good luck convincing capable grad student candidates to join your group to do boring software maintenance for horrible pay and no job security.
What is important to the professors, who decide what the grad students will work on, is again to produce research papers, fulfill their mandatory obligations (teaching, project deliverables) and to continually file for grants. Spending grad student time on porting and maintaining libraries does not help with that. In the worst case your grad student is spending their time maintaining a tool that a competing group's grad students are using to churn out papers, beating you to publications and grants.
What is important for the funding agencies is flashy new research in the current hot topics. I never saw a funding agency that would even consider paying a grad student, let alone a full software engineer salary, to port an academic tool from Python2 to Python3 or do all the other maintenance you need to do on production codebases---nor do most universities even have salary classes and positions for that.
As a result, in the many years I spent in academia, I saw many important research tools rot (both software and large hardware testbeds). The solution is not grad students, but to have fully paid software engineer positions in academia. But realistically that is not going to happen.
Rot is a very big part of what happens to a lot of information.
It should be no surprise that such a poorly engineered process produces awful results.
Keeping current with industry is the only way they can stay relevant in the modern world, particularly in IT and related fields where anyone with a computer at home can do the same research as someone at a university.
“ It's kind of disappointing that developers aren't self-aware enough to understand that they are essentially screwing themselves by not moving quickly in dropping "legacy" support.
"Python 3 will be available in 2008 but we understand that you won't use it for at least 5 years."
There's an entire class of developers who won't upgrade until they absolutely have to. There's also a class of administrator that won't upgrade their current PC browser from IE6-IE8 until they absolutely have to. Developers are basically screwing themselves by not drawing a firm line.”
My shop was interested in migrating long ago, but we had to wait until languages and frameworks (such as Django) supported 3. On projects where that wasn't a risk, we moved to Python around 2014 or so. But: I totally get that for some shops, the feasibility of migrating was low.
Python 2->3 was a poorly managed update that did not follow normal upgrade rules, so the "normal" rules don't apply.
It is still often very extremely expensive to convert python2 to python3. Normally you could upgrade in small steps, maybe a file or library at a time, instead of changing everything simultaneously including all transitive dependencies. That problem continues to be denied, so python2 continues to be used. I say that as someone who has converted code from 2 to 3. Python3 is fine, it's the huge unnecessary transition cost that is not. It's gotten a little better, but not a lot better.
This makes it worse. Now it's even harder to incrementally update, making it even harder to switch to 3.
Just some hack around that would probably already help.
I also supported things like obj.to_string("abc.gz") to get the record in gzip-compressed form.
I also had a from_string() -> obj functions.
You can see the problem. I had to change to_string() so it didn't allow compression (breaking backwards compatibility) and add a to_bytes() for that functionality, and add type dispatching in the from_string() code to support either byte or strings, with different code paths.
And change all the open() calls to use "b", and to add type checks on user-passed-in file objects to insist on reading bytes (if not isinstance(user_file.read(0), bytes) raise "Must be open in binary mode") because all of the underlying parsers are in C and the needless decode/encode step adds overhead.
Oh, and re-write the C extension so it handles both Python 2 and Python 3.
That was a really boring 6 weeks.
If of course you spent an extra decade producing obvious technical debt whose fault is it?
In many organizations there was never a time where they could start writing new code in Python 3. They needed to write code that was compatible with their existing python 2 code, and the only way to do that is to continue to write new code in python 2. Rinse, lather, repeat.
This is why the failure to provide a gradual transition was so bad. When I write new code in Python I use python3, but that assumes that there are python3 modules available that I need.
If you have infinite money this is not a problem. But I think we should be sympathetic to the people who do not have infinite money and have never been given a realistic upgrade path from 2 to 3. The 2to3 program is not a workable solution for many.
That's when the interpreter finally got the minimum support required to make code compatible with both, and linters improved enough, and some libraries started being ported.
Which meant you could not usually convert, since ALL dependenices had to be converted. The chance of doing that successfully then with 300 libraries (including transitive dependencies) was approximately (0.5)^300, which is practically 0.
It takes time between the initial release and enough fixes/additions to become usable, and the documentation/stackoverflow to cover it.
I joined an AI company on 2018 which was building their whole prototype with python 2.7 since 2017. I had to spend 2 weeks and do the migration myself, do a merge request out of the blue and give it to their lead engineers, otherwise i am pretty sure they would have a meeting tomorrow 24/1/2021 to see how they are going to migrate.
Just because industry has a habit of rewriting the whole stack every five years on account of make-work job security doesn't make foundational scientific algorithms change.
If this academic code isn't well understood or well tested, it's probably not as valuable as you might think.
That 5% that requires manually fixing up is the sticking point. You still need to audit every line of the codebase, and each line that gets missed is a guaranteed bug introduced to your codebase by the conversion. This is not much of an issue for small scripts or tiny programs. It is an issue for big applications. This migration really highlights (yet again), the dangers of using interpreted languages at scale. With no compiler to pick up errors, no typechecking by default etc., identifying all of the remaining faults is a huge task.
Like it or not, this is a huge risk to a business. There is a risk of introducing vast quantities of bugs, and there a huge developer cost to performing the migration.
For the record, I have migrated several medium-sized codebases with 2to3 and python-modernize. Because these were internal tools with defined inputs and outputs, it was trivial to validate that the behaviour was unchanged after the conversion. But for most projects this will not be the case.
The 2 to 3 conversion will be a textbook case of what not to do for many decades to come. For the many billions it will cost for worldwide migration efforts, the interpreter could have retained two string types and handled interpreting both old and new scripts. The cost would have been several orders of magnitude less.
That is not at all my experience, nor is that experience of most other people I've ever talked to.
Now do that for 199 other modules and it becomes much more work than “just run 2to3”.
It's perhaps also worth noting that I did this in late 2017. Your experience likely varies depending on when you attempted it.
That remaining two percent had a lot of painful things (truly, I have some stories), but "the overwhelming number of use cases" was trivial.
Plenty of code written in this time frame is liable to depend on unsupported things by that point.
if its not being used, then does it running with the latest tools even matter?
This was absolutely possible. Via the path of upgrading your code (or your dependencies code, in any arbitrary order) to be compatible with both py2 and py3 (via, say, six) and then once all code was compatible in either direction, flipping the switch.
I can think of exactly one language who bungled the upgrade path worse and it’s Perl 6, which they finally renamed after 19 years of stringing people along like it would be the next big thing.
One option could have been to spread the brraking changes over multiple versions, but given the bad state of version/dependency management in Python, this would have likely been a clusterfuck too.
Best option would have probably been to have a longer period of RCs and only release 3.0 with the performance regressions fixed. The myths around the slowness of 3.x (especially for scientific libraries) stuck around for a very long time after they were fixed.
I suspect many python 2 shops will maintain the old and move to something else for the new. Maybe the "new" thing will be Python 3, but I expect this will give other languages opportunity, because people do have emotions, rational or not.