Why I'm Making Python 2.8
naftaliharris.com
naftaliharris.com
What a load of bollocks. For new projects this only matters if libraries aren't ported, which they are for the most part. For old projects, either you're in a situation where you can spend time porting your code to Python 3, or you don't; but as TFA mentioned pep-404, the writing has been officially on the wall ever since 2011 so at that point you have to admit you did choose to incur tech debt and do nothing about it, so the claimed loss of productivity is on you.
> Unlike 2.7 code, Python 2.8 wouldn't be able to guarantee exact 3.x compatibility, since there are some python scripts that will run under both Python 2.7 and Python 3.x but produce different output, and Python 2.8 chooses the 2.7 behavior in these cases.
What a terrible, terrible situation. Now you'll have "python" code that will neither run on 2.7 nor run compliantly on 3.x. As for the latter, please explain how that will alleviate anything on the following point, since behaviour at runtime will be subtly different:
> adding these remaining Python 3 features would greatly simplify running code targeting Python 3, and allow people to use Python 2.8 to run a mix of Python 2 and 3 code.
I don't know what recourse the PSF has but maybe they should even go all in and defend the "Python" name so as to prevent confusion and stop a potential community fracture. Just call it anything else but "Python 2.8" is not Python.
This! A thousand times! I love open source and free software. I absolutely love the fact that you can fork the code and adapt it to your needs. If you find others who like it great! But please, don't use the Python name! It will create more confusion than help. This fork with a different name is 100% fair in my opinion.
While I wish Naftali well in his efforts - I have a private Python-derived language myself! - this is not "Python 2.8." For trademark purposes, "Python" is only what is released or endorsed by the PSF.
We have already reached out to Naftali and asked him to change the name of his project and update this blog post accordingly.
Obviously, though, this is someone who cares a lot about Python, so let's be sure not to rain down on him with a lot of scorn; I admire that he was willing to sit down and 'scratch his own itch.'
Source: I am the General Counsel of the PSF.
- The Author (who isn't me)
https://github.com/naftaliharris/python2.8/issues/47#issuecomment-266240525I call BS (to counter your "bollocks").
Whether the "writing was on the wall" or not, doesn't change the fact that people had to actively port their old code if they wanted it to run on 3.
Sometimes that code could run into the tens of thousands (or even millions for large companies) of lines.
And why would they do it (and at a great cost and time effort)? For the marginal improvements Python 3 brings?
The "writing has been on the wall" is not an excuse, it's mostly blackmail ("port or else you wont run on 3, and we'll stop the 2.x line"). And most people didn't (and shouldn't) fall for that.
However, at a certain point it is worth your time to move forward instead of doing nothing, you gain a little time savings now and you run into few moments of "Oh @#$%^!!" later. Real world example: You don't bother updating ssl to deal with weak DHE and suddenly chrome users can't see your payment site.
My approach has always been to try to front load the work instead of doing it in crisis mode later. It sorta sucks but that's just how software is right now.
If IE7 was slower than IE6, you can't blame people to not move over.
As for Python 2... well, there are still people signing petitions for Microsoft to bring back VB6. Last one was this year, I think.
There are however non backwards compatible changes, like unicode by default, IE7 was also non backwards compatible so the comparision still holds. (with the exception that IE had a compatibility mode if you sent some magic http headers)
Would you also call the RHEL life cycle a blackmail? I'm using version 5 now and the normal support ends in March 2017. My options now are "port or pay extra for extended life cycle or else my RHEL will be without security fixes". And like Python, major RHEL versions break backwards compatibility.
Honestly, that 10 years later we're still having this conversation is ridiculous.
The RHEL life cycle is based on real business needs (and a real business need to balance between newer releases/features and stable environments).
Not on some decree from above that "you should use this new thing".
If you e.g. can't be bothered to do continuous integration or automated testing, then you might consider RHEL with it's life cycle to be an acceptable alternative. Which is fine. Just be ready to pay for that service.
Similarly, if you wanted continued Python 2 support, you could have donated time or money towards that goal. I would be surprised if anybody complaining did that. There's just not that much business value in dragging legacy Python further along.
But to keep it running, you don't really need Python 2.8 with new features, right? You need extended support for Python 2.7 - basically, making sure that it keeps working with updated versions of other software (like OSes), and that bugs are fixed.
Those systems are not just sitting there untouched.
Heck, not even 70s COBOL systems are "just sitting there" (they are hooked to newer systems, get new forms, have alterations, etc. all the time), and those Python 2.7 systems have been written 10-15 years before or less.
And they continue to get new subsystems, new features, alterations, etc. In 2.7.
So, yes, people would very much like to get not just "extended support for 2.7" but also the ability to keep running it in newer versions, and be able to take piecemeal adoption of new features to make their life better and eventually organically refactor in their own timeline.
So let me get this straight.
1. A bunch of people you've never met and probably have never paid or financially supported,
2. Gave you a high-quality programming language, for free, to use for any purpose you liked,
3. And then when you and they disagreed about the best way forward in a new version, you claimed their refusal to continue supporting and adding new features to the old version for you, for free, essentially forever, constitutes "blackmail" on their part.
Do I have that right?
1) You frame this as some single random individual on HN is the only one that is concerned with the switch.
2) You seem to have missed that companies and individuals that do dislike the switch have contributed to the Python ecosystem, from employing core developers in the past, to creating frameworks, libraries etc that helped Python succeed.
3) You have missed the fact that some (a lot? most?) of the concerned people have actually donated to the PSF through its PayPal donate link (as I've done in the past, and I've used Python since 1998).
4) You seem to think that an open source community project is pretty much "anything goes" and end users be damned. And then the team can complain about "lack of adoption" for the new version.
Do I have those right?
Python 3 adoption has been rising for a couple years now as people realize that A) Python 3 is a quite nice language, B) porting to Python 3 is not as hard as people keep claiming it is, and C) Python 2 is going to run out of zero-dollar-cost support one day as the number of people willing to support it without being paid for their trouble diminishes.
If someone does want to commit to supporting Python 2 + backported Python 3 features, they are of course welcome to do so provided they observe the license and trademark terms (not terribly hard to do). But I suspect it won't last very long, at least not as a small-team zero-dollar-cost project. Between Python 3 gaining steam and people staying on 2 in order to avoid work and expense, I just don't think it's going to work out on the kind of decades-long horizon the Python 2 die-hards seem to want.
Here we have a person (the author) who has rejected the path that an open source project has taken, and invested the time and energy to move the source along a path they prefer.
In this particular case, there is a natural constituency of people who share that desire but are unable or unwilling to put in the effort to push the source down the path.
When there is critical mass, that group forks off and begins to bring other people along to the alternate path.
At that point the people who endorsed the change in direction come out in force to yell at these people who aren't doing what they are supposed to and threaten them and implore a higher power to emasculate their effort.
Sometimes that works, sometimes it doesn't. But it always results in massive amounts of confusion when someone new comes to the community and sees these two different paths for the same thing and can't really figure out why they are different.
Further because there is no mechanism for "righting" the ship as it were, the diverging paths lead to a lot of wasted time and effort on everyone's part. This happens to be a Python fork but its happened to window systems, video codecs, graphics libraries, data bases, hell even C compilers.
The nice thing about a Cathedral is that the Pope keeps the Cardinals toeing the one and only line.
Reminds me a lot of the Perl 5 / Perl 6 debates.
They own the Python trademark, so they can make him stop using it.
In fact, I really hope that they do. This project does no good.
I don't expect litigation because that's not how the community rolls. I expect a polite message from gvr to the author asking to change the name.
If that fails, the PSF has no choice. It's use it or lose it.
And he has - https://github.com/naftaliharris/python2.8/issues/47#issueco...
Or license it.
When Python 3 was released, it offered Python users a trade: In exchange for a productivity loss (porting your Python 2 code), you'd get a productivity gain (new features in Python 3 and removed cruft). Some projects and companies thought this was a good trade, and have upgraded over the years, and many have not, and haven't. The interpreter I've been working on tries to improve on the terms of that deal for people who have not switched to Python 3.
> What a terrible, terrible situation. Now you'll have "python" code that will neither run on 2.7 nor run compliantly on 3.x.
That's the point, yes. Obviously any interpreter that's backwards compatible with 2.7 but includes new features from 3.x is going to let people write code that doesn't run under 2.7 or 3.x. But what does it matter if your code doesn't run under interpreters that you aren't using and don't intend to use?
> Just call it anything else
I'll change the name.
Except when they aren't. And then what?
I've run into this multiple times. Sometimes there's a branch of the project for 3 that's underway, and I sit and wait. Other times it means dropping the project or committing to reimplementing a library.
Python 3 is not only the future but the current version of Python. It is the version kids learn in School (in the UK kids do some CS from the age of 6 or 7, starting on scratch and then normally Python), it is the version colleges teach.
However there are lots of reasons enterprise users may want to use a legacy codebase. It is not like Python 2.7 is about to stop working! When a section needs a major re-write, then consider porting it. I don't see how this is different to any obsolescence problem. I know an enterprise software company that wrote a lot of stuff in VB6. Some of it is still in VB6 and they have to manage everything that means (especially around 64 bit architecture problems), when they do major updates they use .net. How can we be in the technology game and not just except that life moves on!
I found it's easy enough to add the line:
# -- coding: utf-8 --
to the top of my Python 2 files so I can get UTF-8. That, and Python 2 has the bindings for GTK (which I like to use). Both versions of the language have their usage, and to each his own. There's no sense in bickering about it.
Text handling is the same. The thing is, many people just deal with ascii compatible English so they don't realize this is a problem for other people and aren't motivated to change. The reason both sides can't just do their own thing (i.e. Python2) is because of libraries and shared code makes it miserable for people using other character sets (most of the world or any company growing bigger than a certain size).
For the record, that's not what anyone's talking about when they mention Python 3's unicode support.
But the runtime is still supported - in fact, it ships with the OS! If you have any non-ARM version of Windows, up to and including Win10, around, check the file named msvbvm60.dll in C:\Windows\SysWOW64 - that's it ("MS VB VM").
And because it ships as an OS component, the official support policy is the same as the rest of the OS, which is at least 5 years of mainstream support (longer if there's no successor release), and then at least 5 years of extended support. This is even clarified specifically for VB6:
https://msdn.microsoft.com/en-us/vstudio/ms788708.aspx
Since VB6 was first released in 1998, this marks 18 years of continued support to date; and if it's not dropped from the OS within the next 2 years, it has a chance of hitting 30 years...
For what it's worth, PSF also has a fairly generous (especially for a non-commercial OSS project) support policy for Python 2.7 - it had already extended the end-of-life date for it once to 2020:
They are only trying to run out the clock on other people's interest- an effort to kill Python2 so it doesn't evolve.
I don't say you are wrong, but I am afraid this situation looks "terrible" only to people who do care about 3. If someone doesn't care about it and thinks that he can survive with never porting to 3 or start using it, for those people the situation isn't that terrible...from that perspective, his 2.7 language evolved to next step, and he know that new features can be used if he upgrades from 2.7 to 2.8.
The majority of Python code is old projects, just like every other established language. This might change in the future as I now see people starting new projects in Python 3, but if your company is older than 5 years old, then there is a good chance that you started with Python 2 simply because at the time of creating your codebase a whole lot of libraries weren't ported to Python 3.
We would look very closely at a Python 2.8, if it existed.
Also, I don't know what you mean by "Amazon only supports 2.7" because boto (the main client for Python) has supported Python 3 for 2 years now. Perhaps you mean Lambda?
Genuinely curious as i thought nearly all of the main ones were ported now
Curious. What do you mean by that?
"Starting fresh" is something many (most?) companies simply never does, over decades. (And if they attempt it is often an all out disaster..)
2011 is fairly recent in this context, and many popular libraries were not available on Py3 until much more recent than that, even if you have the rare luxury of starting fresh.
Only recently 3.x has become a viable alternative. I for one welcome this 2.8 fork.
I also think it's pretty irresponsible of the author to call this Python 2.8, because it may cause confusion to developers unfamiliar with the history and come from a tutorial that is still in Python 2 (it does show up on the first page of Google for me). It's also especially irresponsible and hubristic to attempt to make a language that is seemingly compatible with both Python 2 and 3, because 1) I trust that if it was possible Guido and the other developers would have made it, and 2) it can cause significant confusion when code doesn't work when it hits an edge case, and then the whole tooling around it can't be guaranteed to work. The last thing I'd want in my programming language is unaccounted for ambiguity.
It is possible actually, that's kind of the point! The interpreter I've been working on passes the 2.7 unit tests (i.e. those in Lib/test/), and as well as unit tests for the new features that have been backported from Python 3.
Even if you don't believe me, it's interesting to note that, e.g., while Python 3.0 was being developed, function annotations and keyword-only arguments coexisted with tuple unpacking. I built the code and ran it myself, in fact: https://twitter.com/naftaliharris/status/784421498291310592. Tuple unpacking was actually removed later, introducing the backwards incompatibility after the new functionality had been added. Timeline:
Oct 2006, keyword-only arguments.
Dec 2006, function annotations.
Mar 2007, removing tuple unpacking.
There was also a promising backport of keyword only arguments to CPython 2.6 (!) that was never merged, (http://bugs.python.org/issue1745), due to lack of follow-through.
Why isn't it a valid Python compiler?
To me, the whole morass about trying to end-of-life Python 2 is a bit silly. People have gotten emotional about the situation.
On one side, people like Zed Shaw are calling the Python maintainers 'evil' and claiming conspiracy.
On the other side, people are calling companies using Python 2, 'lazy' and claim they're a threat to the ecosystem.
Yet elsewhere, C is still being written in all of its various year-specific formats, and people end up using 'old' versions simply because they join pre-existing projects or need to totally interface with something that's written in an 'old' version.
Python is an extablished language, it's likely that 10 years from now there will still be Python 2 codebases going strong.
My biggest gripe with this project besides calling this Python is that it's seemingly ambiguous with its code compatibility. I don't mind ambiguity in programming languages, but generally the ambiguous cases are explicitly defined with cases to explain them, and I don't see anything of that nature here, only something that vaguely says that if there's something that works in both the Python 2.7 way will be the default. Without defining those it's hard to know what could happen in an edge case and this could introduce specific bugs that don't present themselves immediately but introduce data weirdness because the cases where something may be ambiguous wasn't defined.
In any case, I think that if a company has a really big, maintained code base in Python 2, it's their fault for supporting an older, in-2020-unsupported version of a programming language and the money/developer time spent supporting the codebase could be spent transitioning it to Python 3. I can understand a little more with an open source project because time is more precious and generally that time is donated, but even then most bigger projects (Numpy, Scipy, Django) have moved to Python 2/3 compatibility so unless the project is gargantuan there's no real excuse besides the project is not maintained.
Except it isn't.
So scoff all you want, but Python 2.x isn't going away that soon.
In the Java world (conservative and slow-moving) JRE 7 (2011) is considered the absolute minimum, and if you're not targeting JRE 8 (2014) you have to have a very good reason.
You mean like, dealing with strings and Unicode? That's usually the case why people have trouble migrating to Python 3.
- There is a community that wants it (largely enterprise).
- The Python team does not want it.
A less controversial word is deprecated - the Python team is discouraging use of Python 2, but not prohibiting it's use or development. That's fair, and if you read this page:
https://wiki.python.org/moin/Python2orPython3
they are not very opinionated about it, largely saying "Use 3, unless you can't, then use 2 and start trying to migrate, unless you can't, then just use 2."
I will say, not to give somebody a bad day but, 2.8 seems like a bad idea. Currently python's development has still largely been a straight line, which is good for transitioning, but 2.8 would cause a fork. It would give a lot of people a short-term win for a long-term lose. Better not to tempt people.
Like you said, Python team sees the 3.x series as the successor AND as a replacement for Python 2.x. They were never meant to exist one beside the other (or, there was no thought put into this before the release).
From my perspective, giving people the choice between 2 or 3 will only give us problems down the road, which is why I vehemently discourage it.
I wonder what about this made this difficult. Was it because it's a language interpreter? Libraries have this problem sometimes, but not as much. (I never hear of issues with gstreamer between 0.10 and 1.0, for example.) Maybe it was just that a binary called python existed? Maybe we should have just said "screw it, python means python2, end of story."
Don't know. What would you have preferred?
In my field what seemed to keep people on python 2 for a long time was numpy or scipy (or both, I do not remember which) which did not get a 3 upgrade for a long time.
Either that, or just call it something different, kind of like perl6. There is no perl6 distribution shipping a perl library or some perl.dll that clashes with perl5.
Regarding adoption, it's 3.0 that's obsolete, and 2.7 that's vibrant. Even for new code (they conveniently only count totally greenfield projects, but most new code is written in fact to work with established 2.x codebases under Python 2, not as a totally greenfield project).
What numbers? All the numbers I've seen -- official numbers from PyPY etc) tell otherwise.
Also, your anecdotal data is arguably biased as well.
My take is that many sources, including the ones linked to in the tweet's replies point to solid growth in Python 3 adoption. Python3 might not have overtaken Python 2 overall, but it's very far from being "dead".
Downvoting because Python 2 is anything but obsolete. People still love using it.
The fact that people love and use something doesn't mean it cannot be obsolete.
At work, I care about more than 35 years old software. It is obsolete (it's written in mainframe SAS with 3270 green screens and some assembly), but people still love using it, mainly because there is no good alternative and it does the job very well.
For reference, Oxford dictionaries define (..."define"? are multiple dictionaries involved here?) "obsolete" as:
1. no longer produced or used; out of date.
Clearly Python 2.7 is in widespread use, and version 2.7.12 came out just a few months ago, so it's neither "no longer produced" nor "no longer used" nor "out of date"...
I mean look at other things. You can still program in C 89 or FORTRAN 77 or COBOL 74 (and no doubt there is somebody still supporting compilers and runtimes for those), but they are all obsolete standards.
Addendum: I think for standards like programming language semantics (which in case of Python is directly embodied in the C implementation), "obsolete" means there is a new standard by some official body (say, the developer of the old standard) that addresses shortcomings of the old standard. So "out of date" is the fitting equivalent of "obsolete" from the Oxford definition.
That's just what the lead project team declared. Not what the user base asked for or wants.
>You can still program in C 89 or FORTRAN 77 or COBOL 74 (and no doubt there is somebody still supporting compilers and runtimes for those), but they are all obsolete standards.
That's because people stopped using them organically. That's not the case with Python 2 -- Python 3 was declared "the new hotness" with a decree from above.
It's like as if the W3C comes out with some incompatible HTML NG on their own and says that HMTL 5 is "end of line", giving billions of webpages the middle finger.
Even worse, it's also as if HTML NG only had some marginal improvements over HTML 5, and was otherwise the same.
The situation with C89 and Fortran 77 is completely different than what you see today with Python 2 vs Python 3. For 99.9999% of C89 and Fortran 77 code you can build the old code with new C and Fortran compilers and use it from today's standards. You can take a piece of code written 30 years ago, recompile it and it usually works.
The word you are looking for is "deprecated", not obsolete.
Really? So even if no one ever uses it, it still renders the old one obsolete?!
feature proposal -> patch -> review -> mergeThe vast majority of current development out there is in python 2.x. A small minority use Python 3.x. How on earth does this make Python 2.x obsolete?
This is like saying Perl 5 is obsolete, just because Perl 6 is out and completely ignore the realities of the real world use of the products.
Here is one recent one: http://www.randalolson.com/2016/09/03/python-2-7-still-reign...
If you have counter statistics showing that Python 3.x is more popular than 2.7 I would very much like to see them.
But to be fair, even with that, I doubt you would get more than 30% of Python 3 users. Which is kinda in line with other surveys, such as the one from JetBrains. (It's probably a good guess that users of Python applications are even more conservative in upgrading than developers of Python applications.)
2.7: 10 million (M)
2.6: 0.5 M
---
3.5: 1 M
3.4: 0.750 M
3.3: 0.05 M
3.2 and 3.1 too low to make a difference.
-------
So totals:
Python 2.x: 10.5 M
Python 3.x: 1.8 M
---
Python 2.x is waaaaay ahead over 3.x
If you're going to create this abomination, at least do us all a favour and DON'T call it Python. Call it Retardython or something. I don't want to imagine people coming into the official support channels and claiming they are using "Python 2.8", then other people lecturing them about what that software really is, etc. Sounds like a horrible waste of time. (Source: I spend many hours a week helping fellow Python users.)
They could license it to Python 2.8 for free if they want to, though it seems unlikely.
Yes, the built in csv module really does that in Python 3.
If you pass it bytes, yes, it does. If you pass it strings, no, it doesn't.
If what you pass to the built-in CSV writer is not a string, the CSV writer will call str() to get a string representation it can write out. The string representation of a bytes object includes the 'b' prefix.
Meanwhile, you discovered your bug: you were treating bytes as text, which is likely to blow up on you sooner or later, and thanks to how Python now handles text, it blew up on you immediately as a way to remind you not to treat bytes as text.
What you probably think you want is for the CSV writer to realize it got a bytes object and, instead of calling str(), call its decode() method to get text it can write. But that is once again a dangerous operation, and sort of the whole point of Python 3's text changes is it won't let you get away with that stuff anymore.
It doesn't "blow up". If it blew up and retired, I would have seen the problem. The problem, like several python 2/3 incompatibilities, is that Python 3 merrily did something different, without telling anyone, until eventually we track down what has changed. I spent quite a while on this very bug myself, and it, along with others, persuaded me to switch to a different language (serious I know, but I was just getting annoyed with python's general loose dynamic nature, in combination with the python 2/3 changes.)
It's not a bug, it's a fix for an architectural error in Python 2, and it was quite well announced at the time: https://docs.python.org/3.0/whatsnew/3.0.html
But the fact that the same code now silently does the always wrong thing in Py3 wrt CSV is clearly a bug.
Actually, the design defect here is calling str() on everything, and assuming that the output is sensible for CSV. It may be a decent rule of thumb, but it clearly does not apply to bytes. Given the likelihood that someone might mistakenly use bytes as a string (for example, because they're porting a legacy Py2 codebase), this should be a hard error, immediately reported as such, and not just a silent behavior change.
Sure, don't change the stuff that works and cause yourself unnecessary pain, but don't blame python 3 for your misencoded data.
If a company tries Python 3 and discovers basic things like CSV produce utter gibberish, they would do well to opt out. And they do--in droves.
My data is not misencoded, you see. It's just misunderstood.
So I won't be able to use NumPy with either of my two native languages. That sounds like a bit of a shortcoming for the majority of the world.
It could be argued that the csv module's behaviour is reasonable, and NumPy's isn't. (I'm not 100% sure about all the details of this issue) Hopefully, NumPy will change it's behaviour to match Python 3, but if not you could still use the NumPy CSV routines like `loadtxt` or `genfromtxt` [0]. So then this becomes a documentation change to add some warnings to both modules.
> they would do well to opt out. And they do--in droves.
This is simply not true. They would do well to handle strings properly and so avoid bugs in future - something Python 3 actively encourages, and Python 2 obscures. And while I can't speak for every company, our metrics show that our Python 3 code has far less customer issues than Python 2, Perl, or Ruby. Now that's business value. (Edit: I mean it's hard to make the comparison - the Perl code is e.g. older - but we're writing code now, and when the interns add new stuff to the Python 3 codebase, it breaks less. All of them are still actively developed, and the Ruby one is about as old as the Python 3 one).
[0] https://docs.scipy.org/doc/numpy/reference/generated/numpy.l...
Text encoding issues are absolute garbage in Python 3.x
I fucking hate the way that csv module works with text encodings.
As soon as I can figure out a reliable way to take latin-1 and save it as UTF-8 without breaking everything, I will try to shoehorn in a PR.
Right now, it's fucking awful. My ETL pipeline hates it, I hate it, my boss hates it, and my internal constituents hate it. Because it sucks.
A file I can read in one encoding and write as another should be readable with the encoding I wrote it in. That is not currently the case with the latest version of Python.
And it makes me hate the world.
with open('some latin-1 file', 'rb) as f:
text = f.read().decode('latin-1')
with open('some utf8 file', 'wb') as f:
f.write(text.encode('utf-8'))
Python 3's string encoding support is super good. I've said it before and I'll say it again: if you use bytes as a string you are Doing It Wrong.If you use bytes as a string you are Doing It Wrong.
If you use bytes as a string you are Doing It Wrong.
in_file = open("in.csv", 'r', encoding='latin-1')
in_csv = csv.reader(in_file)
out_file = open("out.csv", 'w', encoding='utf-8')
out_csv = csv.writer(out_file)
out_csv.writerows(row for row in in_csv)I don't see how silently printing a binary literal, if that is indeed what it does, is reasonable. Simply put, b"foo" is not meaningful CSV.
What it should do is 1) raise an exception by default, informing the user that they need to be supplying strings and not bytes, and 2) provide an explicit switch to treat binary data as pass-thru, which would be useful in scenarios where you're just reading a file and dumping it elsewhere, and don't want to spend time decoding and then encoding everything.
+ if (PyBytes_Check(field)) {
+ append_ok = FALSE;
+ Py_DECREF(field);
+ PyErr_SetString(PyExc_TypeError, "Field is bytes");
+ }
else {
This would then raise a TypeError.I don't think this is the right solution. It seems weird to have a special case because people aren't watching what they're putting in. Garbage in, garbage out, consenting adults and all that.
In fact, the reason I chose Python was because I was able to dive so quickly into real problems like this with no problems whatsoever.
This is a minor issue for someone porting from 2 to 3, it is not a problem with 3
Python's a superlative language but it has a pretty terrible set of included libraries. urllib2 isn't the only library with a superior alternative on pypi. Pretty much all of them do.
I hate Python 3's removal of the (lambda (key, value): blah) tuple unpacking syntax, and the forcing of parentheses for print statements. They might seem minor but they aren't for me. So I'm not at all eager to move to version 3 and don't really see any benefit. Not sure if those who aren't migrating feel the same way, but I wouldn't be surprised if some of them do.
(Edit to address comment below: There are more issues I have with Python 3. It allows more bugs to slip through, for instance. I actually particularly like a comment I just wrote, so I'll link to it here: https://news.ycombinator.com/item?id=13145299 Do note that this was added after the reply below.)
I meant I don't see any benefits for me, not benefits for other people. I assumed that was clear; sorry if it wasn't.
> Unicode.
Yeah, but some people have still been living without the changes, and it's hardly enough of a reason on its own (for me anyway) when there's other things I hate about the language.
> Async.
It's a nice feature, yeah. I can live without it, as people have for many years. Maybe if I was used to having it around I wouldn't want to go back, but I'm not.
> Extended library.
Cool! I'm not sure what exactly falls under this that I'm supposed to be missing, but pip install has sure been taking care of everything in the blink of an eye in version 2.
> Required keywords.
Cool! I need it about as much as I need a donut.
> syntax inprovements (lots of m, many more than just the removal of the print statement).
nonlocal is literally the only positive one I can think of right now that I'd actually care about. But then again, it comes up maybe 50x less often than the parentheses I have to write for print, or the tuple unpacking that I have to do. So yeah, it's hardly a reason to migrate.
> Type hinting.
Nice to have. I'm living just fine without it. Maybe I'd have migrated if it actually optimized things or did something more useful.
You should read some changelogs of past python 3 releases. 3.6, for example, has ordered dicts by default. Which is quite convenient when you need to write test testing a small dict with two items for example.
I like driving an old muscle car, most of m look beautiful and bring me everywhere i want. But damn, those new cars changed a lot and are much more comfortable. (But they do break as much ;))
I think you just don't comprehend the multitude of problems that Python 3 fixes by handling strings correctly... Maybe you've never handled Unicode before.
First of all, the tuple unpacking one is a HUGE readability AND maintainability issue; it's not just syntactic sugar. var[0][1][2] is not only far less readable than the unpacking notation, it doesn't even have the same semantics (doesn't enforce the structure of the tuple).
That means Python 2.7 helped me catch more bugs. Think about that!
Second, no, I just listed the two that irritated me the most every time I tried to switch, because they were the first things that came up by far the earliest and most frequently. Small inconveniences can be amplified through their frequencies. There are lots of things I don't like about it though... division becoming floating point division, having to say list(foo.items()) or list(map(...)) instead of just foo.items(), etc... again, more verbosity and typing for common cases where I really didn't mind the old way. If I wanted imap(), I could've just used imap; they could've just moved that to __builtin__ and made my life easier that way.
By the way -- the lazy nature of map(), etc. also means you catch fewer bugs now. Again, think about that! Just because it looks more efficient, that doesn't mean it's actually better. If there's anything I've learned, it's that even the smallest things tend to come with non-obvious tradeoffs.
Finally, regarding strings: if you look at my earlier comments, yes, I already acknowledged the Unicode changes were for the better. Awesome. I agree. Cool? OK, but there are other things in the language besides Unicode though, and they're not as awesome. I don't spend my entire programming life dealing with Unicode strings, so I care about other things too, and they make my life harder. Simple as that.
I'm a huge lover of Python, but I drop into C# when I need stuff like that. Or Go. Or Rust. Or something.
You're just making bad decisions here.
If you're looking to catch bugs in your code before you test or deploy it, don't use a dynamic language.
Python has never been and probably never will be a language with declarative types and the checking that allows.
Pick a different hammer if that's the nail you need to hit. Don't complain about the hammer you want to use not being a screwdriver.
def foo((birthname,surname)):
...
you can write in Python 3: def foo(name):
birthname, surname = name
...
It's not less readable. I also missed it in the beginning, but it's not really a big deal. (I think they couldn't keep the feature because of how '*' is used, but I am not sure.)Edit: If you have problem with this in lambda expression, just create a named inner function. It's a feature/shortcoming (depends on POV) of Python that you cannot bind variables in an expression. I hope you understand that they couldn't keep the feature in lambdas if they didn't keep it in proper functions.
I am sure if you think about other things, there are good reasons to do it the way Python 3 does it, usually there is a hidden case where things need to be disambiguated (like your list() examples).
In any case, I think I see your problem. You are not the sole user of Python language. There are features that other people like (such as using '*' in unpacking), and so features you like are weighted against their use cases, and a reasonable compromise is made.
And frankly, I think if you like to use lambdas that much, you really want to program in a language where everything is an expression, such as Lisp or Haskell.
list(map(lambda (some, thing): some + thing, everything))
# better
list(some + thing for (some, thing) in everything)
Or, as the parent suggests, just create helper function, preferably one your python environment doesn't need to set up every time your outer function is called # okish, "verbose lambda"
def compute(everything):
def magic(elem):
some, thing = elem
return some + thing
return list(map(magic, everything))
# probably better
def magic(some, thing):
return some + thing
def compute(everything):
return list(magic(some, thing) for (some, thing) in everything)This has been the general trend in mainstream languages lately, not just in Python. E.g. in C#, all LINQ operations are lazy. in Java, the new stream API, to be used with lambdas, is lazy.
You mean print functions. I love the new change because you can pass "print" around like any other function now, letting you write code like:
def my_map(data, func):
for item in data:
func(item)
my_map(dataset, insert_into_database)
# For testing
my_map(dataset, print)If your software is being actively maintained, it's time to move to Python 3.
Maintenance is key. Most people don't stick around for 20 years anymore either. I know I'm going to have an easier time finding a new hire for a Python codebase. And he's going to have a far better chance at understanding said codebase. Code which nobody knows how to maintain will hurt us either with a fiendish bug, or limit out growth. So for me, slowly moving away from legacy stuff is good business value in the long run.
Remember, you can never be sure that Fortran code is 100% bug free. The test of time is as good as any other test, but not perfect.
(FWIW, I'd probably write most new numerical code in Julia rather than Fortran 20xx, and either call into existing Fortran via FFI, or drive it from the command line with some scripting language.)
SciPy has Fortran code under the hood: https://github.com/scipy/scipy/search?l=FORTRAN&utf8=%E2%9C%...
Both SciPy and NumPy use LAPACK - a Fortran library. SciPy also uses BLAS - another Fortran library.
So every time you praise Python for being useful for scientific work, you're actually praising Fortran and C libraries/modules wrapped in Python.
However, I recently had a gig working in py3. Apart from screwing up every single print statement for a long time, it was entirely drama-free, and actually pretty great. There really is a difference between living and dead languages I think. 2.7 is the latin of python.
A lot of people here have strong opinions about the name "Python 2.8". I don't mind changing it, and intend to do so, (https://github.com/naftaliharris/python2.8/issues/47). I picked it initially since when talking with friends about this project it conveyed pretty darn immediately what the project is and does. I'd be very keen to hear people's suggestions for alternate names!
For those of you with 2.7 codebases or projects, I'd be extremely interested in hearing about whether you were able to get this interpreter to run your code. Personally, the biggest challenges I've had so far are with dependencies that check for `sys.version_info[:2] == (2, 7)` as opposed to something like `sys.version_info[0] < 3`. But I'd be very interested in other people's experiences, particularly with larger codebases.
[1] A minor and somewhat pedantic point: The interpreter I've been working on includes PEP 515 (underscores in numeric literals), which is new in 3.6. I didn't think it was right for me to "take credit" for this new feature before it was even out in Python 3.6. Obviously, the real credit for this feature existing (in 3.6 or in any interpreter) goes to the CPython core devs, and especially Georg Brandl.
Keep up the good work, this could help a lot of people!
What if more than one person did that? (paraphrasing Raymond Chen's "What if two programs did this?")
There can be only one official Python 2.8, but since it will never exist (see PEP 404), there can be many unofficial mutually incompatible "Python" 2.8 implementations. Therefore, calling it "Python 2.8" is a bad idea.
So the question is do you want to basically light that money on fire, or just keep your perfectly fine Py2.7 code running and maintained another few years.
More discussions of the monetary value of programming languages please. What is "correct" or "right" isn't all that interesting to many, for good reasons.
Kudos to this project and hope it can set us on a saner migration path to Py3. (Should totally change the name though.)
Honestly, the level of self-entitlement through all 2vs3 threads is staggering.
But I do mind people saying in these discussions "you should all move to Py3 now or you are stupid/evil".
No. There are legitimate reasons for staying with Py2.7 and embracing it.
So I hope "Python 2.8" gets a cool name, perhaps even some funding from a company who wants to keep their Py2.7 code alive and invigorated, and the community part as friends.
Is that self-entitlement in any sense?
All I want is for people who I think have a huge blind spot to stop calling me ignorant for decisons I make about MY code. I totally don't expect core CPython devs help me out though.
It's not just CPython, it's the Python specification of which CPython is the reference implementation. The specification couldn't move forward in significant ways without making some of the changes that came with Python 3.
> All I want is for people who I think have a huge blind spot to stop calling me ignorant
The level of vitriol in 2vs3 threads has been way too high from the start, because people always hate change. You are obviously free to do what you want, you always were. Don't mind the haters, but please don't be one either.
It has already been explained elsewhere in this discussion by other people, I would strongly advise against that. If you for whatever reason have to stay on 2.7, then make sure your new code is 2.7 only (and best if it works with 3 without changes if you decide to change your mind later).
Consider what's more likely in the far future (after PSF will give up on 2.7 support in 2020). That somebody will support 2.7 as it is, or that this guy will support his Python 2.8 hybrid?
Also consider what happens (in the far future) when some library you use will drop Python 2 support. It's not likely it will be easy to run on this Python 2.8 hybrid, either.
And if you for any reason must use Python 3 features in your code base, just bite the bullet and port it.
My py2 code uses unicode properly which may color my view a bit...
But I don't really disagree with what you say.
(Looking at this from the perspective of the "Python Community" or someone who's goal is to adopt Python3) His focus is to back port newer Python 3 features developed since then. Does this help people move to Python 3?
Good on him for digging into cpython. While there's __future__ and the backports module, he seems to have focused on features that aren't just new libraries (which is cool). A few years ago I was trying to backport Python 3's Namespace Packages for my company since our internal import tools effectively do the same thing (except our's had bugs).
[1] https://docs.python.org/dev/whatsnew/2.7.html#the-future-for...
It is very dishonest to call it Python 2.8 then.
Shoot the 2.8 messenger all you want for choosing to call it Python 2.8, but don't dismiss the issue that drives thoughtful people to get value out of this strategy.
I guess the intention was a bit different: he wanted to have stackless features in python. It's not clear to me the reason he decided to back down, whether because of licensing issues or just because the other python developers didn't like it.
[1] http://www.stackless.com/pipermail/stackless/2014-January/00...
I've scheduled time this year for my teams project to update to Python 3. It's expensive in the short term but in the long term we get continued support and new features which is a huge win.
The core driver for the Python 3 break was the fix in text model, this is what allowed literally everything else as it completely broke existing code.
And I, for one, think it's one of the most important improvements of Python 3, the text model of Python 2 is a giant mess and makes it very hard to correctly deal with non-ascii text for any non-trivial software, especially in large teams where not everybody will carefully evaluate the text-ness of their code..
There are counter-arguments to this. Armin Ronacher, author of (among other software) the excellent Flask web framework, thinks that Python 2's system of codecs and byte streams is better in practice [1][2]. Reasons include: You can do byte -> byte conversions with codecs that are no longer possible. You can better handle text encodings besides UTF-8 (and here he describes several embarassing failures of Python 3 to handle OS paths correctly). You can write single APIs that handle byte streams like gzip and text encodings like UTF-8.
[1]: http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/ [2]: http://lucumr.pocoo.org/2014/5/12/everything-about-unicode/
Armin Ronacher works in a very specific context of having to deal with byte/text interfaces in pretty much all his projects, and while I can see where he comes from I work at a different level and at the level at which I work the P2 model is a giant pain in the ass.
> [1] http://lucumr.pocoo.org/2011/12/7/thoughts-on-python3/
http://lucumr.pocoo.org/2016/11/5/be-careful-about-what-you-...
Armin is no foe of Python 3. And as noted in the essaye Python 3 has undergone several improvements or features reintroductions e.g. PEP 461 reintroduced C-style formatting to bytestrings, making generating binary data (especially ascii-based formats) significantly more convenient than it is between 3.0 and 3.4.
Also note that Armin has repeatedly praised Rust's text model, which is much more similar to P3's than P2's (except with static types and no messy legacy).
> and here he describes several embarassing failures of Python 3 to handle OS paths correctly
And (fucking surprise) the issue with that is the text model of FS paths is an embarrassing pile of garbage, Python 2 is convenient because it doesn't try to touch that mess at all and just hands the flaming bag of shit to whoever comes next.
That is incorrect. Rust's text model has (almost) free (and copyless) transmutes from bytes to strings. Python does not. The text model of rust is much closer to Python 2 than 3 in many ways.
Rust's text model strictly separates proper strings and bytestrings, defaults to proper strings and requires that strings be properly formed (so much so that it has additional completely separated platform-dependent types for dealing with OS-originated "stuff").
The one "difference" (which is more in the realm of implementation detail than language text model) is that Rust leverages its ownership system to make UTF8 "encoding" and "decoding" free (literally for the former, essentially for the former). The encoding and decoding are still there and explicit operations though.
> Rust's text model has (almost) free (and copyless) transmutes from bytes to strings.
Only for the specific case of input bytes already in the language's internal encoding (which granted will be common as most inputs would be ascii or utf-8) and with the same ownership constraints as the input, and that's mostly enabled by Rust's ownership model.
> Python does not.
Python doesn't generally do no-alloc/0-copy operations so that's not overly surprising.
Except of course on operating systems where text I/O is done entirely in UTF-16. Say, Windows.
Since Python strings have no fixed encoding, but choose "the most efficient one" (heuristically) when decoding, they can cope better than a fixed UTF-8 encoding in these cases.
>> Python does not.
> Python doesn't generally do no-alloc/0-copy operations so that's not overly surprising.
Indeed. Even when the encoding is not changed, the string will be always copied. One could think of an API that does that, though, to optimize all those cases were memory is already owned by a shim in the runtime.
These are very different and incompatible text models.
At no point is Python's text model fast or overly useful.
As someone who works with Python text processing extensively, I can tell you that the Python 2.7 text model is broken and dangerous, due to the silent bytes-unicode coercion and misguided use of ascii instead of UTF-8 as the default text encoding. Many people don't realize this and will argue that it's not broken, because they have never fed non-ascii text through their app to watch it blow up! And once they realize that they have a problem, they then have to deal with a rat's nest of silent bytes-unicode coercions happening implicitly all over their app, sometimes impossible to deal with due to library code outside their control.
There is a good discussion to be had on whether a language should prioritize bytes or unicode strings as the main data type, but there is no excuse for the "ticking timebomb" string data type design that pre-3 Python has with strings and the default encoding.
For this reason alone I'm very happy that 2.7 is starting to lose its grip. Its continued support is a problem, and I have no love for people who are trying to hold on to it.
There are many other features in 3 that I can no longer live without - most of them now available through backports modules - but types and asyncio can't be easily backported either, and people are starting to use them extensively.
You genuinely are better off using Python 2.7.x and its naive approach to text.
That would mean to me that the API endpoint could be sending me Unicode, in which case Python 3's Unicode-aware CSV is going to work great, and Python 2's csv is fucked. The limitations of Python 2's csv module was one of the key points that moved my company to Python 3.
On Python 3, if you want to be naive about text (not sure why you're celebrating only working in a subset of English, but you have this option), you could open the file as Latin-1 and get the same results as Python 2.
Many CSVs are made with Excel. Excel's only form of Unicode CSV is tab-separated UTF-16. Python 2's csv can't parse those at all, can it?
Nope, not without re-encoding to UTF-8 before parsing (learned that out the hard way and found out it's easier to just take excel files as input).
P2's CSV module works byte-based, and basically only handles ASCII-compatible supersets, assuming your special characters (quote chars, field and record separators) are straight ASCII.
Some points Armin makes are valid and remain valid for Linux-ish systems, but have been shown and refuted countless times for other operating systems; Python is not a Linux-only show. I won't re-iterate all that here.
The fact that there are some parts of "modern" Python which could have been implemented in Python 2.7 backwards-compatibly is irrelevant.
Python 3 is not the language designers worrying about minor subjective issues in python like the print keyword or the design of iterators and deciding that they want to change it all. It is the language designers worrying about major issues like international text, realizing that they regrettably will have to break backwards compatibility to fix those, and then just taking the opportunity to revamp things like printing and iteration since they're breaking backcompat in some pretty major ways anyway.
No, Python had a horrible story for text, and people who worked in limited/sheltered domains didn't realize it. I personally lost all kinds of valuable hours of my life fighting with Python 2's "pretend everything is ASCII until it isn't, then fall over dead" model, because I -- a US citizen, working at US companies, and for quite a while dealing only with English-language content -- still ran into non-ASCII characters with regularity.
And here I'm being charitable; I simply refuse to believe that the overwhelming majority of people who used Python 2 never once had to deal with someone copy/pasting text out of Word or another program that used "smart quotes".
Unicode in Python 2 was fundamentally broken in that whether it had UTF-16 semantics or UTF-32 semantic depended on how the interpreter was compiled. That's a terrible, terrible idea. However, they could have fixed it by sticking to one option: UTF-16 (which provided compatibility with some interesting things that Python interoperated with like Cocoa and, via Jython, Java).
UTF-16 is a sad legacy mistake, but APIs providing Unicode operations of any kind can be build on top. So UTF-16 is a mistake to begin with, but it's not a blocker for supporting all of Unicode and features targeted at the needs of all writing systems and languages. Java, Windows, the Web Platform (including JS) show that proper i18n can be built on top of the bad but backward-compatible 16-bit code unit foundation.
Now, the _even_ sadder part of Python 3 is that if you decide that UTF-16 is a mistake and want to fix it, UTF-32 is the naive and wrong solution. When a Unicode newbie is told about surrogates, they think that UTF-32 is the answer. But then they waste memory and cache line space (and, if dynamically omitting leading zeros on a per-string basis, the compute and copy cost of promoting to different unit width when adding one emoji). And once the damage is done, someone points out that grapheme clusters are a thing, so they still didn't get O(1) indexing to user-perceived units.
The enlightened thing, of course, is to do what Rust does: use UTF-8 and use iterators on top for accessing pieces larger than a code unit (code point, grapheme cluster). (To my taste, Swift strings are too magic and DWIM-y. At least back when I read the Swift book, it didn't even explain the underlying representation. With Rust, the representation is very explicitly known.)
"UTF-16 sucks" is what Python 3 got right. That UTF-32 (with dynamic leading zero omission on a per-string basis) is the answer is what Python 3 got very, very wrong. The correct answers are either UTF-8 (for a new language like Rust) or holding the nose and making stuff work on top UTF-16 (Java, JavaScript) without breaking old programs.
I also think that default Unicode is a major improvement over py2 and is enough to justify breaking the language because modern languages should at least have that.
As someone who writes cross-platform code, Python 3 was a breath of fresh air after fumbling around in the dark with Python 2.
[1] I wouldn't count them out, but they're definitely on the back foot as bash inclusion shows.
Oh, yeah, I agree that Python's unicode model isn't great. I like Ruby's, and Swift has it's own cool thing going where it's very explicit about the uselessness of code points.
However, I think that Python 3 having some form of default unicode support is way better than what Python 2 had. It could be improved (backwards-compatibly too!), but it passes my minimum bar for a "modern" language's text story.
I'm pretty sure the people complaining here about how python2 str only supports ascii and they couldn't paste their smartquotes were bitten either by windows or unnecessarily bad unicode/str interactions due to python not just hardcoding utf-8 auto-conversion. That is the only sane thing to do (Your locale isn't *.UTF-8? Well sucks to be you. By now even the Japanese and Chinese seem to slowly have come around to the utf-8 bandwagon, and they had better reasons then most).
I might be wrong, but I can see basically 3 non-idiotic ways to do text in a programming language:
1. arrays of utf-8 bytes (Rust, Go). Python was close to that already and then messed it up. Indexing indexes into bytes O(1).
Upsides:
- efficient: most text you're going to get is already utf-8 and the rest should be converted on ingress/egress; html/css/most code will be represented fairly efficiently even if the body text is mostly say, Chinese; you can do a lot of text processing by just working on the ascii range (e.g. CSV parsing).
- sane: no BOM, no 32 bit encoding of 21 bit quantities etc; unix-compatible
Downsides: - can't efficiently access individual logical characters or know the fixed-font width of the text
- normalization is kinda nasty (concatenation etc.), in practice people just tend to ignore that
- hard to constrain to only valid utf-8 without significant downsides
- maybe not that beginner friendly
2. use some non-array type that doesn't allow for indexing (e.g. ropes), probably using (mostly) utf-8 for internal encoding.
3. arrays of logical characters. That means you need to make up fake characters to handle graphemes that are not directly representable as a single pre-composed code point in unicode. The upside is that this has beginner friendly semantics in a sense and allows indexing on what's meaningful in the domain (graphemes). The downside is that I can't see how to do this with a lot of complexity and some nasty gotchas. This seems to be what perl6 does https://design.perl6.org/S15.html#NFG
And also that you yourself, and any users of your own project, will need to run it with this guy's own homebrewed version of Python in order to know it will behave correctly.
Useful?
The "horrible hacking" already exists bundled into libraries and tools like Six and Future.
> split it in two seperate packages.
That was tried, and failed every time. Because now you end up with two diverging and hard to reconcile code bases. A single-source cross-version library, while not trivial (and not allowing the user of more advanced P3 features) works way better.
Only for the simplest cases. How is Six going to help you write a regex that matches emoji, for example?
On Python 3 you write a range that includes the emoji you want to match, and you're done. On Python 2, that regex may or may not compile, depending on what sys.maxunicode is. If sys.maxunicode is 65535 you have to fall back on a different complicated regex that has a bunch of cases to find emoji in their UTF-16 representation. If you want to avoid that system-specific behavior, you can I guess encode it to UTF-8 bytes and write an extremely complicated regex that finds emoji in UTF-8.
This is my prime example about how something that's easy on Python 3 can require horrible hacking on Python 2. A wrapper library doesn't fix that problem.
Python 2 and 3 have different semantics, and the only way a correct, automated translation between them would be possible would be to emulate one inside the other.
So?
> How is it that you claim it supercedes six?
Its `future` library goes further than Six and it provides CLI tools to convert both Python 2 and Python 3 code to 2/3 code.
Examples of A in this discussion:
- Disagreements about the degree of library support for 3.x
- '' the ease/value of porting from 2 to 3
- '' the rate of industry adoption of 3 for new projects
- '' the degree to which people are driven away from Python entirely because of the version situation
Examples of B in this discussion:
- Claims of paternalism on the part of GVR/ the PSF in pushing 3
- Claims of unreasonable/emotional attachment to 2 by partisan devs
- Claims of willful distortion of facts by both sides (see A)
- Claims that the writing is on the wall for 2, because usage of 3 is supposedly accelerating
- Claims that the writing is on the wall for 3, because it's supposedly taken too long to drive not enough adoption
It seems outrageous to suggest that the differences between 2 and 3 are anywhere near as significant as the differences between, say political conservatism and liberalism, and yet the level of partisanship seems nearly the same. How did it come to be like this?
I like how your comment stands out as being such a rational, non-aggressive observation in all this. I wish I had the answer for you, but I can't figure it out either. shrug Maybe some people are using this as an opportunity to vent the frustrations they've had with using the language (despite it being so loved, and ranked 3rd in most-used according to IEEE), and since there are two versions of the language instead of one, they can more readily create an object doomed for the epitome of their hatred, the sacrificial lamb going to the slaughter (sounds like something from a bizarre school of thought in psychology). You'd think it'd just be easier to let people do what works best for them. I'm sure many expect to take collateral damage from the differences in versions (imagine being a Python 3 fan and having to start working for a company that exclusively uses Python 2 or vice versa), but these people also readily accept working with languages with radically different designs and welcome in copy-cats (like Clojure is to other LISP-like languages). The important thing is that the language does what you need it to do. The name on the download link shouldn't make that much of a difference. The world has adapted to accepting Python 2 and Python 3 regardless of the bickering. There's room for "Python 2.8", even if it is poorly-named. I imagine there are probably dozens of Python 2.8s in existence - they just haven't been noticed by people here on HN. Besides, in about 3 months, we probably won't be hearing more about Python 2.8 anyways unless someone else decides to use it.
We have a project at work and the only reason we need to support 2.7 is because we want to run under Jython.
Anyway, I would like to hear other people's use cases for Python 2.7.
AIUI, some of those who were paid to work on it previously are now doing so only in their spare time.
See example in the fine manual: https://docs.python.org/3.7/library/urllib.request.html#exam...
You are only told "we know python.org uses utf-8 so just decode it as utf-8." No further discussion, no pointers are provided how to correctly fetch an URL with text content into a string. Even small convenience function that at least tries to look on Content-Type: header would help here!
I am well aware that "in py2k it just worked" was mostly an illusion. But honestly, is the situation above an improvement?
PHP has been successfully deprecating features for decades now and cleaned up their code base; they have never had a schism in the way Python has so you can argue all you want about legacy, or people preferring 2.7. The reason this happend (and is still happening) are human not technical.
You know, as a strong proponent of Python 3.* (it's a significantly better Python), I started to write a withering critique of your statements here (and I do strongly disagree about naming it Python v2.9). But then I started actually looking at the evidence.
For example, when you go to download Python at the python.org website, you're still presented with a choice between Python 2 and Python 3 side-by-side, looking for all purposes like equivalent choices (Python 2 should be much less prominent and toward the bottom and/or the download button should be much smaller than the Python 3 download button)[0].
And then, when you follow the link for "Wondering which version to use?", you get a full page worth of hemming and hawing, and cost-benefit analyses and so on[1]. In fact, they should be saying something like:
Always use Python 3 unless you have a strong reason to do otherwise.
This typically only applies to people who have large legacy code
bases and are for some reason unable to upgrade, or else need a special
legacy library only available in Python 2 (most libraries are
available in Python 3)."
So wow, having read the official python.org statements on the Python 2 vs 3 issue, I'm pretty appalled. No wonder some people who are relatively new to Python are confused or surprised. I'm on Debian or Ubuntu mostly and so download Python via package manager or PPA, or build from source. So I had no idea.It's also been a huge mistake for GVR to have allowed so many new and attractive features to be backported to Python 2.7. It's taken away a significant amount of incentive for people to make the move. Python 2.* should have been bugfixes only for a long, long time now.
That said, if you're in charge of significant amount of Python code for a company and you've not seen for years that Python 2 is a deadend, you've either been engaging in wishful thinking or oblivious to what's happening in the Python community.
The idea that Python 2 needs to be sabotaged and the community forbidden from improving it so that people makes the move, with the people who are stuck to Python 2 codebases held as hostages, should be an indicator that something was done very, very wrong...
Personally, I'm neutral as to 2 vs. 3, I use both, but the schism is the main drawback of Python for me and it often drives me to just use other languages. Python 3 is great and has very neat improvements, but it should have just been called a different name so that both branches could evolve freely and compete on their own merits, rather than on the PSF mandating to use one over the other.
To be fair to the "Should I use Python 2 or Python 3 for my development activity?" page, it starts with:
> Short version: Python 2.x is legacy, Python 3.x is the present and future of the language
Massive breakage of backwards compatibility in a minor release is bad form. They absolutely did the right thing by naming it Python 3.0.
Rather than have this huge break maybe just take the time to move all of Python forward rather than us still 8 years after Python 3 leave us still talking about it.
Interesting. Now I want to make my own Python 2.8 which is 2.7.latest with the "-3" command line argument hardcoded to always be set.
The PHP Project did manage to allow people to write code that worked on both 4 and 5. At the same time, the project failed to get Unicode in. It's unclear how that will be possible without breaking a lot of programs. If PHP ever switches string handling like Python3 did now, it will be just as painful. Well, that's my prediction anyway, maybe they will figure something out. Just realize that their first try (PHP6) didn't fly at all.
Theres a module called __future__ for exactly this purpose
Ridiculous as it might sound, that implies it would actually take less work to devise a path for Python2 that's backwards compatible but still has a future, and invest the time to backport all 3-only code, than it would to proceed with killing off 2 and porting everything over. It's also nothing more than everyone whose code is 2-only is being asked to do sometime in the future (they chose "the wrong competing standard" and have to pay the cost). Either way someone has to foot the bill for a unified Python but couldn't it just be that the early adopters of 3 could pay that cost if it saves updating (say) twice as much legacy code in favour of backporting the 3-only code?
Of course, the future version of Python should be the best form of the language with the right features, and that's why it was decided to kill off 2, but if after this many years the initiative hasn't completely succeeded, there's always the option to reverse course.
Heck, in the time until Python 2 is officially gone, maybe this new fork will evolve into a better language than 3 and still be backwards-compatible with 2!
For now I'm sticking with 2.7 for as long as I can but will just accept it when the time comes.
[1]: https://mobile.twitter.com/vlasovskikh/status/80172061331236... for example (I don't think I've made a controversial claim here but could be wrong)
This is just a really bad idea IMO.
Nope. Can't be done without getting Python 3 either way, because Python 3's text model is not compatible with Python 2's. That is why the core team allowed the other breaking changes, because software was going to be broken in the first place.
it's not 100% seamless, but its almost there.
Unless you're willing to put in the time and energy to extensively test, this is basically a recipe for disaster. You're basically paying at least 80% of the cost of a full Py2 -> Py3 port (since looking for regressions is a huge part of that cost), for only a fraction of the benefit.
It would become a substantial detour and end up being a kit more total work.
Also, porting can be done in parallel to normal dev. You aim your port at a specific release while you continue fixing bugs. When the port is done, you port over all the patches. Repeat until port and original version converge.
I think leveraging PEP414 and carefully using bytes, unicode or "native" literals is a much more resilient mode of operation.
If by this you mean "give me exactly Python 3's text model without breaking my unmodified Python 2 code", you should be aware that A) this is impossible because B) the Python 3 text model is backwards-incompatible with unmodified Python 2 code, and C) that's kinda why Python 3 existed and was backwards-incompatible.
nope - you can mix them in python 2 code. The way it should have been done in the first place.
- Concurrency from Erlang (http://queue.acm.org/detail.cfm?id=1454463)
- String handling from Go (https://blog.golang.org/strings)
- Lightweight data classes (https://kotlinlang.org/docs/reference/data-classes.html)
- Syntactic macros from Nim (https://hookrace.net/blog/introduction-to-metaprogramming-in...)
- more from Lua, Julia, Scala, and Haskell (within reason)
There are already plenty of examples of Python modules that implement significant modifications (e.g. Cython, Rpy2, Dask), so I don't think it would really feel all that different from typical Python programming.
from __past__ import bytestring_literals, loose_comparison, integer_division, ...
? Being able to upgrade one file at a time to Py3 would have made porting so much easier! I realize some things could be hard to support file-at-a-time, but others (like my examples) would obviously not.I've held off on porting the Py2 code that I'm responsible for because of the lack of total test coverage. I know that something tricky with unicode, None comparison, or the list->generator split will cause a failure in a weird, untested edge case. If I were making a numeric library, that'd be one thing. However, I mostly write complicated business logic API glue code that is hard to test without huge amounts of mocks.
Also, this kind of thing gives me nightmares about porting:
WSGI therefore defines two kinds of "string":
"Native" strings (which are always implemented using the type named str ) that are used for request/response headers and metadata
"Bytestrings" (which are implemented using the bytes type in Python 3, and str elsewhere), that are used for the bodies of requests and responses (e.g. POST/PUT input data and HTML page outputs).
So according to the spec (!) WSGI headers MUST be `str` in Python 2 & 3, but that means that the semantic meaning of the types changes. How on earth am I supposed to write good code for working with headers (let alone mocking and testing) when I'm required to decode them in Python 2 but not in Python 3?Something like this? http://python-future.org/translation.html
Surely that already exists?
The arrogance and stubbornness of the CPython dev team is starting to face the consequences we all knew were coming. If only they'd at least done unicode right. Gone with assume UTF-8 and kept the indistinction between bytes and unicode, more of us would be onboard.
What would happen with my python 2.8 when it needs to be converted to the official python 3?
OK, so it may well be terrific, but it should really have a different name.
1. Calling it Python 2.8 is a really bad idea. If you want to fork Python in this way, great. I'm sure very few people in the community would have a problem with it if you called it Brothon or P8thon or Hackthon (could run into trouble with that one, but who knows?) or IHate3.xThon. You can call it LifeOfBrython or Snake, depending on how much you care about where the name came from.
Guido van Rossum is one of the least litigious people you can find in the open source community. When he wanted people to stop naming packages after PEPs (pep8, pep257), he didn't try to go to court and figure it out later. He talked to the maintainers of the packages and explained why he thought it was problematic [0]. I wouldn't expect a cease and desist showing up anytime soon, but it's in poor taste, and I hope you reconsider it. Python is not yours just because you are free to use it and do as you please with it. It's really stretching the concept of open source for you to unilaterally declare a collection of hacks to be a point release that's been specifically addressed and decided against by the community as a whole.
2. I have so much sympathy for people who must maintain Python 2.x applications and cannot make the business case for putting the time into porting it to Python 3. I'm in that situation myself right now, and I have been for years. I can't find a compelling argument to warrant the time so long as the 2.7.x branch is getting security updates. When the security updates stop, the move will finish.
The reality is that those of us who have been working with Python for a long time now (10+ years) have already figured out reliable workflows to get around the most lacking aspects of the language in our domains of expertise. Or we have swapped out subsystems in other languages for cases where we absolutely cannot find workarounds.
I get it. I'm in that boat. I understand that boat because I live in it.
I can even understand why this project happened. Because hacking things you love to make them better for you as an individual is the heart and soul of what makes software development so great.
What I absolutely cannot understand is why anyone (absent certain libraries needed) would start a new project on a deprecated branch of a language. And I don't understand why this is such a contentious issue specifically with Python. I don't hear any of my friends who work with C# or Go getting into arguments about how the latest version has failed or shouldn't have been rolled out. As much as Swift 3 has caused problems, I don't hear any of my iOS developer friends complaining about how that shouldn't have happened and trying to backport the good stuff into the 2.x branch. I genuinely don't understand this behavior in the Python community.
*
I do, however, suspect that it has something to do with the overall age of the community.
Using Python has always been a matter of taste and style. I get that. Python is a language that makes explicit tradeoffs between performance and style. It is no shock to me that people who learned to love the language in its older form want it to stay that way. And I say this as someone who is closer to 40 than I am 30: we older folks in the industry need to fight against the stereotype that we can't or won't learn new tricks. Yes, there's a place for battle scars and pushing back against every new flavor-of-the-month stack that rehashes old problems in new ways. But obstinately sticking with stuff simply because it's what we know does all of us who are getting on in years a huge disservice.
Getting a good gig as a software engineer in your 20s is not that difficult. Getting the same gig when you're pushing 40 and the hiring managers are still in their 20s is a lot harder. If you're making decisions to start new projects on old versions of any language, please, please, please, do all of us a favor: make sure that you are doing it for solid reasons and not simply because it's what you already know and are already comfortable with.
*
Because this post isn't long enough already, I have a second pet theory about the reluctant adopters.
Python 3 is a _great_ language. It's not the same as Python 2. It makes different tradeoffs about styles and choices of expression from what Python 2 does because these things can and should change over time.
I think that the expressive power of Python, the closeness to human language, and the ease of reading well-formed code have a lot to do with the resistance to Python 3. There are people (I'm one of them) who get really really picky about language in general. I think the Oxford comma is a necessity in almost every case. And I think people who casually omit it are stupid, shallow, non-thinking, drones who don't care about the history of language and don't care about precise meaning of words, and therefore don't care about me because they an't be arsed to toss a comma in a place that greatly clarifies meaning.
I don't actually think all of that, but I'm closer to that than I am to not caring at all.
And there are many like me. When you're dealing with a programming language with intent at its heart, people are going to take that in different ways. When you're dealing with a language that cares about whitespace and eschews braces and wants to limit brackets and just expose the pure logic of the program, people really are going to get fired up about print vs. print().
That's expected. And beautiful that so many people care that much. But as much as I lean towards prescriptivism and that rules are good for a language, even I have to admit that language evolves. But it usually evolves to be more inclusive and more expressive, not less. The really bad fights about languages in general revolve around what to include, not what to exclude. Because languages are exclusive by default.
A counterpoint to my idea above that we are all just old and lazy, is this: that we really do genuinely care, and that we care for good reasons. But something has to give. We cannot refute the evolutionary pressure. Python 3 is as necessary for modern speakers of programs as a recent edition of a dictionary is in your preferred language.
Yes, you can get by with something older, and you can even make yourself understood by most people who speak the language.
But you are limiting your ability to express your intent when you make an intentional choice to refuse to adopt what the rest of the culture around you is doing.
Aaaaaaand I'm done. Sorry for how long that was. I didn't intend that. It just sort of happened. If you read all the way to the end, let me know, and I'll upvote you.
This is a bad idea. My guess is the author will not be able to call this thing "Python", and for good reason.
Should it exist? Sure. Why not? The nature of FOSS is that anyone can tinker with it and make it into whatever they dream. Go for it! Just don't call it Python.
I use both 2.7 and 3.x. No drama. New projects go through a "Do we see any issues with 3.x?" phase where we try to list libraries we'll need and check for support. Not that difficult.
I absolutely understand companies/people holding on to 2.7 for dear life. Converting a non-trivial working code-base would be costly and very difficult to debug if that code base doesn't have extensive test coverage (probably true of most). It would also be utterly irresponsible in most business cases.
Most businesses don't have software developers sitting around doing nothing. Revenue comes from existing products and new features along with bug fixing. No customer of a software product will pay one dime for a company devoting a year to port their entire code-base to 3.x. Can you visualize that announcement?
"We stopped delivering new features and fixing bugs a year ago. Instead we ported our entire product to Python 3.x. Today, a year later, we give you exactly what you were using a year ago. Enjoy!"
Yeah. Exactly. The huge sucking sound you'll hear is that of customers leaving the company throughout an entire year of nothingness. In a dynamic free market competitors would eat you alive as you stop delivering features and fixes while they zoom right past you with a better offering to your customers.
That said, sticking to 2.7 for the long haul --say, ten years from now-- will create the Python equivalent of old COBOL code still running in deep dark places within financial institutions. It will create codebases nobody wants to look at or touch. It will create codebases that will be anywhere from hard to impossible to support as libraries will surely evolve to support the 3.x and, eventually, 4.x branch.
It is perfectly sensible for a company to, given today's realities, stick to 2.7. This is almost exactly the problem described in "The Innovator's Dilemma". Good management means focusing on delivering what your current customers are buying and want to buy. Unless they are clamoring for your product to use 3.x it could actually be really bad management to make the switch.
The only way to do it correctly would be to hire a full parallel team of programmers to port the codebase while mirroring every single new feature and bug fix implemented as customer's needs are met. At some future point the two branches would achieve parity in function and reliability. This parity would allow seamlessly switching to the new codebase without damaging customer relationships. Of course, this would cost a ton of money for a non-trivial product and, at the end of the process, the company would probably have to fire one of the two teams. Pretty messy and costly in more than just financial terms, isn't it?
I am not advocating either approach. Just saying I understand this from both engineering and business perspectives. People pushing others to just switch are doing so from a frame of reference devoid of any understanding of the realities of business. Most businesses are not about the technology, they are about what problem you solve for your customer. They don't care about "the geek stuff" behind the curtains. And rightly so.
Live long and prosper.
What's even ironic here is that the print command he used is Python 2 instead of Python agnostic. I wonder which version Randall Munroe would prefer, if either.