Python 3 makes it easier to develop high-quality software
migrateup.com
migrateup.com
Python 3 adoption takes longer than entire NodeJS lifetime. In the same timeframe PHP introduced and EOLed 5.3 (edit: and there were 4 years between release of 5.0 and EOL of 4.x, which involved backwards-incompatible refactoring of OO semantics).
They are not skeptical. Nobody thinks "yeah if I just wait long enough, Python 3 will go away and 2.8 will come out".
They just don't see an incentive to update. Simple as that. It is basic economics of time and money.
Python 2 wasn't terrible and Python 3 isn't dramatically better. Unicode and other stuff in Python 3 is nice. Is it nice enough to start digging into stable, working, money making code just to say "Yay, Python 3"? I think the development community has spoken by how it was handled. It's gonna happen but it will happen at a very slow pace.
Imagine this alternative scenario -- "Python 3 brings 3x speed improvement + removed GIL". Ok, I bet you, there would have been a much faster adoption because now there is a greater incentive to justify it. People would think "Yeah, I'll destabilize my code a bit but I gain performance, I can do that".
Someone in a sibling comment suggested "They should have just stopped updating security patches for 2.7, that would have forced them to use 3.0". That's ok if Python holds a monopoly on development languages. But it doesn't. As soon as that happens developers will not just look at updating to Python 3, they'll start looking at other ecosystems: Go, Elixir, Java, Clojure, Scala, NodeJS etc.
To add to this; if it included even a 2x speed up / memory reduction, I think 90% of everyone would have switched by now, and you know what else? Go probably wouldn't exist.
If people would rather move to another ecosystem rather than update their code that's fine. The Python community doesn't get paid a nickel every time the interpreter runs, (at least slightly) engaged users are what move Python forward.
There is lots of FUD from early 3.0-3 days. Python 3.5 is solid and well supported by 3rd party libs check this out
Python3 has been adopted, now. It just isn't newsworthy so we don't hear about it.
It is hard trying to introduce Python 3 at a place that has systems that depend on both Cpython and Jython. Jumping between python 2 and 3 between projects is quite taxing mentally. And staying on python 2 only leaves me with a sense of lingering doom to say nothing of having to endure all the idiosyncrasies of python 2.x. I really enjoy the few moments when I can use Python 3.
It is perhaps easy to consider the 2 vs 3 debate over and done with for people who can just disregard Pypy and Jython entirely in their environment. I probably would think so myself had I not been so unfortunate to be the unwilling maintainer of solutions dependent on Jython right now.
Also, PyPy is 2.7 compatible, although there's a 3.x beta.
That doesn't make Python 3 a failure or unadopted, it's just the structure of our industry. By search & job stats [1], the most popular programming languages are Java, C, C++, C#, and yep, Python. Companies tend to pick whatever language is best for early adopters when they first get started and then language usage grows along with the company. So the most popular languages now are the ones that were cutting edge 15 years ago, when the biggest companies of today were tiny startups. Go look at what's most popular among startups today and you'll see what'll be the hot programming languages in 15 years.
[1] http://www.tiobe.com/index.php/content/paperinfo/tpci/index....
I've done this anecdotally (i.e. not with a huge quantitative sample size, but looking at specific "hot" tech companies and the tech stacks they end up using) and it seems to work, but is subject to certain gotchas. Data points:
EBay (1995) - Perl, rewritten in C++ in 1997, rewritten in Java in 2002. Google (1998) - C++. Del.icio.us (2003) - Perl. GMail (2002) - Perl, rewritten in Java circa 2002. LiveJournal (2002) - Perl. Flickr (2004) - PHP. Facebook (2004) - PHP. YouTube (2005) - Python. Reddit (2005) - Lisp, rewritten in Python in 2006. AirBnB (2006) - Ruby on Rails. GitHub (2007) - Ruby on Rails. Dropbox (2007) - Python. Hacker News (2007) - Arc, YC internal tools (c. 2012) are written in Ruby on Rails. Twitter (2007) - Ruby on Rails, rewritten 2010 in Java/Scala. Uber (2010) - Node.js, supposedly this was a rewrite.
There are two big caveats:
Sometimes a big competitor comes along right when a language is undergoing a massive rewrite - for example, it looked like Perl would own the Internet around 2003, but just a couple years later we got Rails and Django and Perl stagnated with Perl 6, and so Perl jobs are nowhere near as hot as they would otherwise be. Similarly, Python was poised to take over much of the non-web server space in 2008, but in 2009 Node.js came out just as Python was losing steam in the Python 3 transition. Python still remains pretty hot (it's been helped by its use in data science), but it may've lost the startup server market to Node.
The other big anomaly is that there was a lot of attention around Haskell and Erlang from 2005 - 2008, and yet these are still niche languages. Erlang had one big success with Whatsapp and some smaller ones with Facebook chat and CouchDB, but didn't really become mainstream. My hypothesis there is that because both of these were old languages that were rediscovered, they were built with some assumptions (eg. strings are lists, or funky POSIX interfaces) that didn't fit modern startup development, and so they couldn't gain critical mass. But then, both Python and Ruby were old languages that were rediscovered; perhaps it was the sudden resurgence of Linux (which both Python and Ruby were steeped in) along with the web that carried them forwards.
New companies, like the one I'm part of, use Python 3, all the time. Django, and the other 15 packages we require, all support Python 3. In fact, I've been using Python 3 since 2008, and I've been using it for work since 2013. I think we're past all of that, at this point.
I'm not sure why migration is always regarded as a metric for succes. Python 3 wasn't supposed to be a "upgrade", per se. It was supposed to allow the core to rid itself of a lot of the language's warts. It has done this quite well without sacrificing much. The result was a slow migration, not by existing code bases, but by means of new companies.
It's true that 3 is happily happening for many people, effectively and without fuss.
But it's incomplete to pretend the schism is not an issue; not to try to understand the pro-2.x points which many reasonable users have.
Its a huge issue, even now, and to pretend it's not there is to do BOTH pythons a disservice because it will simply serve to feed the disagreement, to the detriment of both flavours of the language.
The Python 3 project needs to understand where it went wrong, acknowledge it, 'better language' arguments (true-ish but largely ineffective) notwithstanding, and try to bridge the community.
Yes the 3 segment reasonably wishes all this would just go away so we can all focus on moving python forward, but it's not going away. It must be adressed without yet another post about 3 being better, which simply makes 2.7 people feel like second class citizens. They've heard all the arguments before.
Already happened years ago. Breaking backward-compat is the issue, and is a difficult situation for any popular platform. Guido has mention it won't happen again.
(However, it is clear that going to Python 3 would not work.)
There are asyncio ports of many popular and less popular libraries, and it even has a nicer syntax than NodeJS.
Python 3, on other hand, has tried to fix many design decisions and hence requires some transition effort. The effort is many folds more with so much software already written in Python 2 and large codebases. I am not surprised the adoption is slow.
Java has done a great job of evolving over time -- they made the controversial choice of using type erasure in generics. This makes generics much less useful than they are in C#, but Java kept compatibility between the classic and generic collections and C# didn't. Thus the C# code base has two different syntaxes for collections, particularly in the stuff that comes from Microsoft.
Java and PHP have had both in spades. Back in the Perl 4 days, people really wanted a real object system[1].
Python 2 (aside from a few things that annoy the crap out of me, but apparently not many others) is a pretty complete, fairly decent language. Perl5 is pretty complete and people are pretty happy with it. (Well, the ones who write in it.)
[1] Yes, I know, not here to argue the semantics. It is "sufficiently OO" for what people want to do with it.
Yes it is possibly to laboriously add code to check and encode/decode text, that might sometime be text and sometimes binary pretending to be text pretending to be binary. But why go through that pain when you could just use Python3 where things are clear and generally actually work.
It's also very intuitive, which is great for teaching - I've always had a hard time explaining the difference between "strings" and "unicode" objects in Python 2, but the clear boundary makes it straight-forward in Python 3.
For the same reason if you need to write code that should work on Py2 and Py3, it's much easier to write it on Py3 and then add backward compatibility through "from __future__ import ..." and the six module.
It's much less clunky than six. Instead of manually adding the compatibility layers to your code, it automates the process, which results in cleaner code. It can even translate (some) Python 2 modules on-the-fly.
Many projects switched to it from six.
Okay: You can't. You'll need to convert both strings to the same case, and then compare those. In the majority of cases, you can do:
return string_a.lower() == string_b.lower()
But sometimes that won't work, because converting case isn't always easy. And that's a unicode issue, not a python issue. It's not even completely a unicode issue, just a "writing can be very weird" issue.
For example, "ß" (U+00DF) is a character that doesn't/shouldn't have an uppercase form. Unicode does support it as U+1E9E, but all the examples I know of will convert it to "SS" ([U+0053, U+0053]) instead.
Try this in a javascript console, for example.
> "ß".toUpperCase().toLowerCase() === "ß".toLowerCase()
false
> "ß".toUpperCase() === "SS".toUpperCase()
true
> "ß".toLowerCase() === "SS".toLowerCase()
false
> "ß".length
1
> "ß".toLowerCase().length
1
> "ß".toUpperCase().length
2
> "ß".toUpperCase().toLowerCase()
"ss"
Also, try ctrl+f'ing for "ss" on the Wiki page for ß [0] and see what it matches. Or ctrl+f this page for "ß" and it will match every "ss". I think it's pretty neat.[0] https://en.wikipedia.org/wiki/%C3%9F
(edit: You might already know everything in this comment, but I started writing and it ended up being a cool exercise for myself, so I figured I'd submit).
Absolutely! I think there's a lot of baggage in computing because it happened to emerge in places where language could mostly be written in simple characters with surprisingly simple rules. When exported internationally suddenly you deal with "oddities" that are really rather the norm in other languages.
In French for example is customary NOT to put accents on capital letters, even though it's possible. But then again sometime people do it anyway - it's not a hard rule. So often converting to lowercase may require contextual understanding because some words are written the same but with different accents, and would capitalize to the same string. Which word was meant would only be evident from the context of the sentence.
I suspect that if computing had rather developed somewhere where these things were commonplace we would have much better handling of it today.
But, personally, if you're teaching Python to a class, whether it's 2.x or 3.x. You should probably stick to using some variation of string.upper() == string.upper() to explain simple comparisons, due to simplicity.
Either way, if you're not dealing with differently encoded character data in Python 2.x as noted above, you're causing yourself some unnecessary headaches.
But seriously, all what python (and as a matter of the fact majority of languages) is placing a clear distinction between a byte and a character. This forces developers to keep in mind what they are dealing with and reduces number of bugs.
The truth is that if you treat everything as bytes, chances are that your program has bugs. Actually good test is to convert your existing Py2 code to Py3, if you have encoding/decoding errors then your code was buggy.
No, they make a clear distinction between bytes and code points. This is the wrong approach as code points are not characters. People are going to think that everything is fine now that they're using code points and the unicode type when they should be using a unicode library for text manipulation.
Compared to that, Python3 was like gift from heaven.
Did you know, that there is an operating system, where a simple loop that reads lines from file and prints them to the console will end up with errors? (that being Windows and some European code pages, where console is cp85x, filesystem read and write cp125x).
Well, and that's before the user tries to redirect stdout to file, in which case it will end up with error too, because that is going to be 7-bit ascii.
Anyway, Python3 does the right thing with defining the encoding at the edges and everything works right for the user out of the box.
https://docs.python.org/3.3/whatsnew/3.3.html#pep-393-flexib...
Glancing at Julia, I'm not sure "invalid character index" is really a great thing to have to handle when dealing with text. It does seem that systems for dealing with text will eventually offer various apis beyond just a somewhat disfunctional character api (which I think is a fair description of str in Python 3).
It's non-BMP characters that trigger the conversion to 4 bytes per codepoint. Unfortunately Emoji are non-BMP and are super common.
Non-BMP characters in 2.x-3.2 give you implementation-specific behavior. It can vary from one system to the next. Some string methods fail on "narrow builds" of Python (the ones that store their characters in UCS-2), while others return different results than they would on a "wide build".
I happen to know this fact. Are you really asking that everyone that comments on a language acknowledge all the possible implementation details? I at no point complained about the memory required to hold strings in CPython so I can't really understand why you would bring this up.
And I don't want to be forced to distinguish between strings and bytes. To me they are one and the same as I use UTF-8 for everything. Not once in my code do I need to encode/decode anything.
I've always wanted to know how things are done internally. Makes me a better programmer. What I said about all this extra conversion does to performance is correct, especially when scaling up.
And my choice to use UTF-8 internally is a valid way to do things. I make sure that what comes in from the outside world is UTF-8 (either by rejecting or converting) and then don't have to worry about it from that point on. The bulk of my code deals with outbound data so no need for me to do any encoding as it's already UTF-8.
As for regular expressions, they work perfectly if your delimiters are in the ASCII range. If you're trying to match a single character then they don't work for Unicode even when using Python3's unicode type. The reason is that code points are not characters.
For regexes, \w matches any unicode word character (unless ASCII mode is on), just use it on strings. If you use it on bytes, then yes, you can only match bytes.
And this perfectly illustrates what I posted in another comment elsewhere. Thinking that doing operations on code points (indexing and length) is the proper thing to do. You need to use a unicode library to work on graphemes and take into account grapheme clusters.
> For regexes, \w matches any unicode word character
From what I've read, \w just matches code points. I did find that the https://pypi.python.org/pypi/regex module supports \X which matches graphemes as per the unicode spec http://www.unicode.org/reports/tr29/.
Yes, I could build the packages/wheels myself. Yes, I could install a compiler on the target machine and let the binaries build from scratch. I won't, however, since it's a lot of time for remarkably little benefit. Python 2 is still well supported, and is a great language to program in.
Yes, Python 3 has cleaned up some of the rough edges of Python 2; but the improvements are still not significant enough to offset the cost of porting libraries and building packages.
Which distribution doesn't ship python3.3?
e.g.
python-mysql python-django python-keyczar python-pyhsm
> but for node we expect everyone to install/build from source an not use the supplied packages
This is not necessarily an improvement in the state of packaging. Verifying that unsigned source code has not been changed is a hard problem (which NPM has not yet solved). Installing compilation tools and extensive header files on a production machine opens up a new class of security vulnerabilities on that machine (whether a VM, container, or bare metal).
Many companies use it to run a recent Python stack on any Linux distro they like.
One solution, build on your dev machine and produce distro packages instead.
Are we talking about the same Python?
Both sides are extremely persuasive and thus I felt that no matter what choice I made it was the wrong one and would haunt me down the line.
So I just stuck with a different modern language that has wide acceptance and isn't battling with itself.
It really isn't that big or hard of a decision. For the vast majority of projects, the choice doesn't matter, so I'd default to Python 3 at this point. Really, the only reason to choose Python 2 for a new project is compatibility with libraries you need[0], but the list of Py3-incompatible libraries is constantly getting smaller and smaller.
That said, if you already have a lot of Python 2 code written, there probably isn't any point to upgrading--the benefits are marginal at best, and it can be a lot of work to migrate a large code base.
[0]: Here's a pretty good reference for Python 3 support in popular libraries: http://py3readiness.org/
In contrast, with Python, on a weekly basis I run into 2.x and 3.x problems...last week it was running the Google Cloud SDK, which is 2.x only. I say this as someone who has almost switched exclusively to Python (3.x) and enjoy it...but man I miss the days of Ruby's ease-of-use.
(but yeah, 1.8 to 1.9 was not a lot of fun, but thankfully, many years have passed since then)
This is the same issue facing Perl (although the Perl situation seems more dire) WRT Perl 5 vs Perl 6. For Python maybe the benefits of keeping Python 2 alive for large users offsets the harm of any confusion, I have small projects (all in Python 3) so it's a rare annoyance rather than a deal breaker.
The situation with Perl is neither more dire nor really the same issue as with Python. Perl 5 is not going away. Perl 6 maintainers are not the same people as Perl 5 Porters -- both will be maintained for the foreseeable future.
Second, Perl 6 is a significant change to Perl, with far more features and somewhat different syntax. People will want to change for the features. Running Perl 5 code from inside Perl 6 works by calling out to libperl, which runs a real version of Perl 5. Perl 5 will still be popular due to ubiquity, access to CPAN modules, and things like CPAN Testers.
So it's a choice for the developer, to develop in a new language that just happens to be Perl-flavored, or to go with any other language. Not really a crisis.
Python seems more popular at Academics, and it's essentially the Basics for the generic public, or replacing Java at colleges nowadays.
That been said, I use PHP myself, PHP7 looks promising.
Python 2 vs 3, AnguarJS 1 vs 2, the self-conflicted "upgrade" is double-sided.
Speed improvement is great as well. Seeing 2-5x speed improvements. :)
My problem is with its premise. That premise is that Python 2.7 people need to be somehow educated. To be shown some kind of light. It's the same premise that the project leaders have been going on about for 7 years. "Can't you see how much better this is?". Worse: "are you 2.7 people stupid, or what?".
It hasn't worked.
Why hasn't it worked? Not because 2.x people are slow or uninformed. They're very well informed. They assess the merits and downsides. They've had 7 years!
The reason is a combination of inertia, but also, many people actually think that in many ways, 3.x is a regression. If you don't need Unicode, Unicode is a huge pain in the xxxx. I was just having to mess with "b"s everywhere today on strings over msgpack to a non-pythonic endpoint. Lazy ranges (can also be) a huge pain in the bt. So I must move to 3.5. For what? Type annotations? Asyncio? I've been using (the highly impressive) Tornado for 3 years on 2.7.
Now google brings out the brand new TensorFlow. On 2.7. "But porting to 3 is their #1 priority" I'm told. Okay cool. You wait. I won't.
So the advances of 3 meet the regressions of 3. Even if you believe the net is positive in 3's favour, that net is not nearly enough to persuade people with actual lives outside of computer science, to port years of code, and more importantly, to port their brains to work with the new dispensation. Python is a language for people who want to get things done. 2.7 gets things done. Zero problems. For many people, that's the killer feature.
Seriously though, the point is well taken. The right tool for the right job will always apply, and I like to think most developers and programmers acknowledge that there are very few tools that should be used in every opportunity and by every person.
Just because "Use X, it's better" is a frequently made argument doesn't mean it's wrong, especially when X actually is better!
I could say "the main difference between Ruby 1.8 and 1.9 is that programs run twice as fast in 1.9" (true, 1.8 is dog-slow.) And that's "Y is better", but it's also fairly specific.
"Python 3.0 makes it easier to produce high-quality software" is annoyingly vague. They seem to mean software that behaves more consistently by reducing the amount of default unpredictability in the language spec -- that's fair and specific. It's also not what the article said, sadly.
Ultimately the Python 2 -> 3 transition has been as smooth as developers make it. There are certainly a lot of transition efforts put in place, such as the ability to import from __future__ some of the 3 behaviors and abilities and some automated migration tools. At this point, too, most of the major libraries today support 3 and there's few reasons not to build for 3 and maybe do a few tweaks to also support 2, if you truly have to. (A lot of good 3 code runs on 2 unmodified and just needing the Python equivalent of polyfill libraries and __future__ imports.)
Some of the fracture has been binary ABI compatibility and mini-versions of this struggle can certainly be seen in just about any language with an ecosystem of cross-platform native libraries to support (the brief Node/io.js being an interesting fracture more in that it was resolved so amicably so relatively quickly).
But ultimately it's tough to resolve the psychological and political hurdles: for whatever reason there seem to be "die hard" Python 2 fanatics that hate the direction Python 3 has moved and don't seem interested at all in compromise or transition. It certainly doesn't seem to be as strong as the hate divide between VB6 and VB.NET (which will likely forever be a weird blood feud until VBA and VB6 programmers die of old age), but still haters are gonna hate.
https://www.python.org/download/releases/3.1/
(Even if the above is deemed BS, the release of 3.0 was 7 years ago)
Many of the 3.0 features can be switched on in 2.7 and there were some attempts to automate as much of the code change as possible (but people ended up preferring to write code that runs on both 2.7 and 3.x, but I suppose that then is not really evidence of a fracture).
Google seriously spent tons of effort at the time trying to remove that GIL. They were one of the biggest proponents of Python at the time, and frankly they were part of what made it so successful. With their work on Unladen Swallow not making meaningful progress and the Python core developers concentrating on Python 3000, making it clear that they don't even care, it isn't a wonder that Google was hedging their bets by hiring Rob Pike to work on Go, which many now see as most directly being a competitor to Python, and which has sapped a lot of people who would potentially be Python users.
Meanwhile, the broken release of 3.0 needs to be blamed for part of the problem, not used as a reason to delay the historical date of fracture: a lot of things were broken in 3.0, including some basic things that Python got wrong in Unicode. The performance regressions were not quickly fixed, making people hesitant of the benefits of Python 3, and the Unicode issues (which were not fixed until years later, when they added the round-trip-safe Unicode conversions; but even this is a workaround, as filenames are defined to be strings of bytes by the filesystem) were ironic as the whole point of Python 3 was to somehow be better at dealing with Unicode :/.
And you seem to be forgetting that the real problem was not the features added in 3.0 but the features taken away at the same time that made it impossible to write reasonable code that ran on both 2.7 and 3.0 even for extremely simple cases like "catches an exception". They only just a year or so ago have started to crack their party line of "shut up and upgrade to 3: it is better and you are wrong" and add some backwards compatibility features to 3.x, such as supporting the u prefix (which was the absolute dumbest thing they could have removed).
I think there are however a small number of very vocal people who didn't like the change, and their posts tend to linger...
Perl was the workhorse of the early web, when I was a young programmer interested in the web it seemed like the most worthwhile language I could learn. Perl isn't dead but it's lost most of its appeal, and a large part of that is due to the failed development on Perl 6.
Perl 6 still isn't available.
I guess it still counts as unstable software, but I used Gmail beta for half a decade and that turned out fine. At this point, its a matter of nomenclature, because I've certainly used buggier "final release" software.
[1]: https://github.com/coke/perl6-roast-data/commit/2aeeb87edc
Look, I think Rakudo Perl 6 is great, and its quite usable for many things. OTOH, I think it is quite fair to refer to it as "incomplete", because:
> It is very nearly complete
"very nearly complete" is a kind of "incomplete".
You can't really make a fair comparison to Perl 6 uptake before even the Perl project is calling Perl 6 "ready" with Python 3 uptake 7 years after PSF said Python 3 was ready.
If you're on 2.7.x, you can certainly import in quite of the newer features. There are also some resources for writing 2/3 compatible code, such as the six library and http://python-future.org/. It is no problem for fairly trivial applications, but I don't envy anyone who has to ensure 2/3 compat for a large project.
His example is using it as an angle, where comparing two of them is conceptually pretty sensible.
If you pretend it's a structure, sure, you'll need to extract primitive data types out of it before doing anything with it. But if you want it to behave as a higher-level abstraction, reimplementing less-than (say, for comparing the vector magnitude) at every call site isn't great.
> comparing two of them is conceptually pretty sensible.
Well, there are a bunch of sensible ways to do it, depending what the angle represents. You might not want to privilege any of them over the others directly in the class.
Is it a point on the unit circle? It doesn't make much sense to compare (sqrt(2), sqrt(2)) to (-1, 0).
Is it an angular distance to travel from point A to point B? Then you want 355 and -5 to both be less than 10.
Is it an amount that a screw has to be rotated? Then you can't take mods, so 720 > 360 > 0.
Do you have a robot that can only turn anticlockwise? Then maybe 5 < -10 < 355. This is the comparison implied by the implementation in TFA (self.degrees = degrees % 360), but it's pretty obscure.
2. Would you think it's reasonable for a class to implement equality ==, so it can tell you if it's equal to another instance of the same class? Then why not ordering?
3. Take Python's set type as a useful example. A < B tells you whether A is a subset of B, A | B returns the union, etc. All very handy.
1) some libraries I actually wanted to use were still 2
2) When I tried to google a solution for something, more than once I stumbled across a code which just didn't work. Then it occurred to me, that this is because the example is Python 2, while I am on 3.
To sum it up, I later switched to 2, and everything was fine since then. I also asked my friend, who is a professional python programmer, and to this day his company still uses 2 and don't plan to switch.
For anyone still in this situation, re-evaluate them - that issue is mostly solved.
Python 2 is deprecated. It's no longer supported after 2020. There's no point in using it for new projects, unless there's a very good reason. You're missing out on lots of incredibly cool features like asyncio and static type checking.
There's even an increasing number of Python 3-only libraries.
This is why you frequently hear that large, established codebases that control their environment aren't updating to Python 3. It's pretty easy to move a few thousand lines over but complexity scales non-linearly with the size of the codebase.
Funnily, though, I just thought 'yeah there's that library I use, they never ported that', for one of my codebases. Just looked it up and they now have ported it. So I guess it's time for me to do the same. I wonder what proportion of (actually used) libs are Py 2 only still.
For those who don't know: http://python3wos.appspot.com/
(Edit: Not that I care what they use internally... but sometimes they release a really cool open source library and it is Py2 only... arr! But at least they are sharing, so I can't really be too upset :)
There's really no good reason not to support Python 3, and projects which refuse will at some point be forked.
The entire scipy stack, ntlk, tensorflow has partial 3.0 support and plans to support both, caffe, theano, etc.
There are few to no libraries that don't support, or don't plan to soon support both.
I suspect this is just an issue of they didn't want to block the release on not having Python 3 support ready at launch.
Python 3 insistence on unicode hurts my brain. I had a 'pleasure' of writing Python 3 script manipulating data produced by Python 2 (custom pickle). Ended up reimplementing whole Python 2 pickle library by hand, because apparently no one imagined a use case involving reading raw bytes into a buffer.
Python 3 provides no real benefit to an experienced Python developer. So considering all the things you could possibly do, switching to Python 3 is not switchting to a new and shiny thing, it's just a bad decision.
If you really want to switch to something new and shiny, there are now many more interesting language options. Modern Javascript has copied so much from Python, you can feel very much at home there. If you care about concurrency, you have Go. If you care about concurrency and care about using a language that's well designed or just want performance, there is Rust.
David Beazley recently held a great talk on this very topic, in which he goes into these issues in more detail: https://www.youtube.com/watch?v=lYe8W04ERnY
In the real world, software depends significantly more on the properties of a developer or a team. Building software is about people more than it's about tools.
I'm sure I hit Unicode problems in Python 3 a few years ago as well, but if current versions are better that would be a sufficient reason for switching.
It's really hard to keep up with language changes and implementation changes.
EDIT: I am so surprise at the number of downvotes in this HN thread. I bet there is some haters out there today :-)
https://docs.python.org/3.0/whatsnew/3.0.html
The other items in the article are also mentioned there, but not as headline items. The great majority of behavioral changes from 2.7 to 3.x should be listed in that one document, there was a freeze on language changes for several of the 3.x releases.
A lot of the PEPs are way too verbose for a quick bite. PEP are great for a weekend read, but for a quick summary, OP does a better job.
There are lots of other changes since then, but most of them are additions (that don't directly impact old code) or to libraries or whatever.
http://migrateup.com/whats-really-new-in-python-3/
(Actually, maybe not, because it doesn't have code examples.)
Guido should have used his power to force the move forward down people's throats, because that's the responsibility of a BDFL. The D is there for a reason, and its exactly in this cases that the power of the Dictator role should have been used.
Communities rarely move forward on their own if left to decide democratically: they need o be forced or tricked into moving forward by individuals with vision!
...heck, I'm even starting to like the nodejs folks for their "shove the future down people's throats" attitude :)
Python 2.8: - boring easy to fix changes like division, print function, exception handling syntax, removal of iteritems() / etc
Python 2.9 - library reorganization (urllib2 & urllib, etc)
Python 3.0 - unicode stuff
Most organizations that have been lingering at 2.7 for 5+yrs would have long since migrated to 2.9, and would thus be half way to 3.x in terms of backwards compatibility.
Now instead, everyone has to introduce the better part of a decade's worth of backwards compatibility changes at once, or be left in the dust. Hardly a fair choice.
https://docs.python.org/2/library/__future__.html
(in versions where the new behavior is the default, the import does nothing and is not an error)
The reason for that isn't "Python 2 is intrinsically better than Python 3," though. It's the sheer inertia of a language whose last minor version (2.7) was released five years ago and the persistent meme that Python 3 is buggy and unusable, despite clear advantages as noted in the article and the added compatibility of popular Python 2 packages in recent years.
As with hardware, Python 2 will only die when the top packages completely drop support for it, and it wouldn't surprise me if that happens sooner than later. (Django, for example, will be dropping Python 2 support in 2017: https://www.djangoproject.com/weblog/2015/jun/25/roadmap/ )
That is, unless said software needs to include a popular library that has not been ported.
There's really barely anything that is python2 specific, most of the things that weren't ported is no longer being maintained.
I think the main reason you should start moving to python3 is that it will be discontinued. Python 2.6 was EOL in 2013, 2.7 they've been more generous and you have 4 more years, originally it supposed to be 2015.
There's no good reason to stick with 2.7, if you have some legacy apps there's nothing preventing you from having Py2 and Py3 installed side by side.
If you need to write libraries that share code, it's much easier to write the code in Py3 and then make it backwards compatible with use of from __future__ ... imports and six module.
I feel like people who still complaining about the difficulties did not actually tried Python 3. Py3 is much more enjoyable language to write in.
Regarding GIL and JIT. Those problems are mostly performance related and not easy. As you mentioned GIL is available everywhere so it's not a good argument regarding py2 vs py3.
As for JIT you get it with PyPy which by itself is not 100% compatible with CPython (for example as of now majority of python modules written in C simply won't work).
Although they do have PyPy compatible with Python 3 although it is not as stable, but since Py3 is picking up hopefully it will be becoming more popular.
For me, it's the other way around, I have not used PyPy primarily because it's compatible 2.7 and I don't use Python for performance, but if I could get performant Py3 implementation then why wouldn't I want to use it?
The ecosystem is rotting for 2 reasons.
First, there's no easy way for library developers to maintain a codebase that supports both 2/3 development in parallel. Yes, it's easy to use a 2to3/3to2 tool but unlike transpilers in Javascript the output code may still require fine-tuning after the tranform.
This could be solved by adding preprocessor support to python (ie for conditional loading of code) but GvR is adamantly against this approach use to the way many devs write unmaintainable code in c via the preprocessor.
I actually published an OSS project called pypreprocessor that was supposed to solve this issue. It works to add preprocessor conditionals to code but won't work with both 2/3 code in the same file because there's no hook to inline code before lexer runs and throws syntax errors.
Second, package management in python still sucks. Do you use distutils, distutils2, setup.py, etc? There's no single standardized approach that everybody agrees on and each tool follows a unique approach or needs to be supported because of legacy code.
Unlike the Ruby, Javascript ecosystem where package management and publishing is relatively straightforward/easy, it's an absolute PITA to publish packages to PYPI.
Back in 2011, the answer was. Just wait 5 years for everybody to adapt their libraries to support Python 3. 5 years later, nothing has changed. The existing libraries don't want to piss off users by dropping Python2 support and maintaining two separate branches of the code in parallel sucks. So, since most everybody that uses python in production still relies on Python2, Python2 is still the default.
The library ecosystem makes the platform. The tooling sucks so library devs don't want to make the switch.
Until there is a major cultural shift python devs have 2 options. Use a better version of the language without the support of a mature library ecosystem, or use an outdated version of the language with a rich ecosystem.
Sometimes, though, the 2 version had bugs because of this behavior. So, the bugs were both subtle and bi-directional. I definitely fault Python 2 for allowing that practice, and am glad we're now using Python 3.
For example, I just discovered a few days ago while porting a Python2 library that byte concatenation is an O(N^2) procedure in Python3 while it is an O(N) procedure in Python2.
So you have to use bytearrays and really rewrite all the code that deals with bytes. Basically, even if something works, you still have to performance profile.
In python2 strings/bytes are also immutable but concatenation is O(1) w.r.t. original string. Not sure about internal implementation differences.
so basically whenever you see code like:
some_bstr = b''
for something in get_something():
some_bstr += something
And if you have to do it this way, make sure that some_bstr is a bytearray or make it a list that you later join.1. Calling the next( ) builtin instead of .next( ) method on an interator or generator.
2. .items( ) on a map in 3 is like .iteritems( ) in 2. Which means that nasty short cuts like map.keys( )[0] no longer work.
Oh yeah, and print( ) is a function :)
d = {chr(i):i for i in range(65,91)}
print(d)
Do it with python2 and python3. You'll see that the output in python3 changes every time.If someone was relying on consistent ordering, they're going to have a bug.
Python2's ordering is deterministic[0]
[0]https://docs.python.org/2/library/stdtypes.html#dict.items
> If items(), keys(), values(), iteritems(), iterkeys(), and itervalues() are called with no intervening modifications to the dictionary, the lists will directly correspond.
If you restart Python, you have broken the correspondence.
You'll note that either the documentation is incomplete or your interpretation is incorrect, as the most recent versions of 2.6 and 2.7 will use a randomized hash table when the -R flag is enabled:
% ~/Python-2.7.10/python.exe -R x.py
{'M': 77, 'L': 76, 'O': 79, 'N': 78, 'I': 73, 'H': 72,
'K': 75, 'J': 74, 'E': 69, 'D': 68, 'G': 71, 'F': 70,
'A': 65, 'C': 67, 'B': 66, 'Y': 89, 'X': 88, 'Z': 90,
'U': 85, 'T': 84, 'W': 87, 'V': 86, 'Q': 81, 'P': 80,
'S': 83, 'R': 82}
% ~/Python-2.7.10/python.exe -R x.py
{'Z': 90, 'Y': 89, 'X': 88, 'W': 87, 'V': 86, 'U': 85,
'T': 84, 'S': 83, 'R': 82, 'Q': 81, 'P': 80, 'O': 79,
'N': 78, 'M': 77, 'L': 76, 'K': 75, 'J': 74, 'I': 73,
'H': 72, 'G': 71, 'F': 70, 'E': 69, 'D': 68, 'C': 67,
'B': 66, 'A': 65}
I can totally understand how people expect an invariant order. As I pointed out, our regression code broke in the 2.x series because we relied on consistent ordering, and CPython never made that promise. But what I quoted above is the only guarantee about dictionary order. Everything else is an implementation accident.Nor is it the only such implementation-specific behavior that people sometimes depend on.
>>> for c in "This is a test":
... if c is "i": print "Got one!"
...
Got one!
Got one!
That's under CPython, where single character strings with chr(c)<256 use an intern table. Pypy doesn't print anything because it doesn't use that mechanism.Note that 'is' testing is also faster:
% python -mtimeit -s 's="testing 1, 2, 3."*1000' 'sum(1 for c in s if c is "t")'
1000 loops, best of 3: 893 usec per loop
% python -mtimeit -s 's="testing 1, 2, 3."*1000' 'sum(1 for c in s if c == "t")'
1000 loops, best of 3: 1.01 msec per loop
This extra 10% is sometimes attractive.We are talking about the default way of doing things in the most commonly by far used implementation.
I'm not saying someone should have relied on the specific ordering or that the code that does rely on it is a great way of doing things.
CPython2 did make that promise -
CPython implementation detail: Keys and values are listed in an arbitrary order which is non-random, varies across Python implementations, and depends on the dictionary’s history of insertions and deletions.
In other words, the order should be the same no matter how many times you restart the application.It is a subtle source of bugs since it always worked in each specific version of CPython2 without the -R flag.
Because either the documentation means to include -R in the description, in which case your interpretation of the documentation is incorrect, or the documentation is incomplete because it doesn't describe a valid CPython 2.x run-time. Either way, it indicates that the difference isn't, strictly speaking, a Python2/3 issue.
"an arbitrary order"
Where does it say that the arbitrary order must be consistent across multiple invocations? Quoting from https://docs.python.org/2/using/cmdline.html#cmdoption-R :
> Changing hash values affects the order in which keys are retrieved from a dict. Although Python has never made guarantees about this ordering (and it typically varies between 32-bit and 64-bit builds), enough real-world code implicitly relies on this non-guaranteed behavior that the randomization is disabled by default.
I totally understand your point. I remember the debates about how this would break code. But it's there to mitigate algorithmic complexity attacks against an every increasing attack surface. This was the best solution they come up with, along with a migration path to the new default.
If your unit tests are the most vigilant, you'll test the output of that function for just such a comparison.
For example, xmlrpclib in 2.x will return either str or unicode depending on whether the particular piece of data it's returning happens to contain non-ASCII characters. In 3.x it always returns a unicode string.
I think a lot of it boils down to, sometimes you really really don't want to care about unicode. Linux filenames are one thorny issue - Linux thinks of them as strings of bytes, and a tool is broken if it can't also deal with them that way. Py3 makes it very difficult to do that correctly. There are also issues with stdin/out being opened as unicode vs byte streams.
For web development, the "unicode sandwich" (http://nedbatchelder.com/text/unipain.html) works great in my experience. However, I can see his point that for some kinds of tasks, Py3 is unambiguously a downgrade.
Regarding the huge mess, that exists with or without Python 3 if filenames use unknown, incorrect or inconsistent encodings. To treat a path as as text, e.g. to show it to a human or send to another system, you have to know its encoding. Python 3 and GTK (with its G_FILENAME_ENCODING and G_BROKEN_FILENAMES) make this clear - some other software (like Python 2) just silently uses and propagates the broken data.
Many people don't, though. Many people run their Python scripts in cmd.exe.
Did you know that cmd.exe supports unicode just fine? It's Python that does not (out of the box). The development version of Click for instance supports input/output in cmd.exe with both 2.x and 3.x.
Unicode is not a feature of Python 3.
I doubt that anyone has ever managed to output Unicode into cmd.exe in Python before you made this.
I hope Python 3.6 can incorporate what you did.
Any changes that could be made in some compatible way have largely been backported to the Python 2.x versions (the latest stable version of which is Python 2.7.10) that have been released more-or-less in parallel with the 3.x line (search for 'backport'): https://hg.python.org/cpython/raw-file/15c95b7d81dc/Misc/NEW...
Enough such has been done that you can write Python code that runs in recent versions of both Python 2 and Python 3 (which is very important for 3rd-party library maintainers that want/need to support both it in a single codebase), but to do that you will be forgoing most of the productivity and 'niceness' benefits of the sort that are described in the OP.
Basically, I think a lot of the shims in the six library could have been included in 2.6, 2.7 and 3.x.
Also, I think it would have been wise to spread the breaking changes over many releases, not saving them all for one massive breakage. That way you can deal with only a few subtle bugs and breakages at a time.
Reading "If you're starting a brand-new application today, there are plenty of valid reasons to write it in 2 instead of 3. There will be for a long time." on the same day it was written is not the same as 6 months or a few years down the lines.
I don't actually know when this was written except for the copyright: 2014-2015, so it can't be more than a bit short of 2 years ago.
It's one of the more irrational and idiotic suggestions I've ever come across. One of, if not _the_ most important piece of metadata for anything that has or will ever exist is its creation date. In particular for writing, putting it in a historical context usually makes the piece more powerful and increases your understanding of the subject. If it's bad writing then it's not going to have a chance anyway. Sorry to say but that may be true of the post in question. I was under the impression that due to poor Python 3 adoption, Python 2.7 would continue to be supported.
I second OP, please date your blog posts.
This has been true for the last seven years, and I certainly don't see it changing anytime soon.
It's ok to be pedantic, but only if you're correct about the thing you're being pedantic about.
A Python 2 dev can learn Python 3 in a day.
Deleted comment