Moving Away from Python 2
asmeurer.github.io
asmeurer.github.io
20 years ago there was this great MP3 player, WinAmp 2. And then they released WinAmp 3, which broke compatibility with skins and plugins and which was slow. People didn't upgrade. Finally, they came to their senses and they released WinAmp 5, marketed as 2+3, which was faster and brought back compatibility with older stuff.
In the retail trading Forex world, there is this great trading app called MetaTrader 4. Many years ago they released MetaTrader 5, which broke backward compatibility and removed some popular features. People didn't upgrade. Today, they are finally bringing back the removed features and making MetaTrader 5 able to run MetaTrader 4 code.
Me, I wait for Python 5 (=2+3) which will be able to import Python 2 modules, so that you can gradually convert your code to the newer version.
I have tens of thousands of lines of Python 2 code in a big system. I can't just take 2 months off to move all of it to Python 3. Moving it in pieces is also not really possible, since there are many inter-dependencies.
Uglyfying my code with the "six" module it's also not a solution, since when I'll move, I won't care about Python 2 anymore.
So basically I'm just waiting until a consensus emerges.
It should be noted that the quantity of Python 2 code keeps on growing, I wrote most of my code while Python 3 existed. If Python 3 allowed an easy path forward, we wouldn't be in this situation.
I've upgraded a lot of my code with just future imports and it is just a tiny step from there, unless you're doing a lot of character encoding. Leave those projects behind, but moving most other projects is easier than I expected.
There are few paths for upgrading. You haven't mentioned 2to3, but I will. I've migraded projects of the same size you mention (10s of kloc) with 2to3 and the process took days, not months. About three days with one developer, actually.
There are a number of projects which also have code bases which are compatible with both Python 2 and 3. You say that once you move, you won't care about Python 2 any more, but I that's not a good reason to reject six outright. In my experience, the code isn't really "uglified" but there are just a few minor cases here and there you need to think about.
Anyway. Small-ish projects like the one you are talking about are usually not as hard to migrate as you might think. There are some exceptions, of course.
Those exceptions are often bugs in the logic of handling text. Python 2 allows you to sweep that under the rug.
The Windows API had the same problem, it was ASCII and then they added Unicode. But instead of just forcing your app to exclusively use a single one of them (sort of like use Py2 OR Py3), they allowed you to mix and match your code, and call either the old or the new version of the API. And gradually, people stopped using the older version and started using the Unicode one. Today, new Windows APIs are strictly Unicode.
Of course, the solution was ugly from their side, having 2 duplicate APIs for the same thing, but it didn't brought pain to their users and allowed them to proceed at leisure.
The main breaking change is that you can't read a stream of bytes from a byte-oriented resource (say, a network connection that by definition can only send a stream of octets) and magically treat it like text elsewhere. While that's a tempting thing to want to do, it's been an unending fountain of bugs over the years. For example, in Python 2.7:
>>> a = 'this is a test ಠ_ಠ'
>>> b = bytes(a)
>>> b
'this is a test \xe0\xb2\xa0_\xe0\xb2\xa0'
>>> type(b)
<type 'str'>
In Python 3.5: >>> a = 'this is a test ಠ_ಠ'
>>> print(a)
this is a test ಠ_ಠ
>>> b = bytes(a)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: string argument without an encoding
>>> b = bytes(a.encode('utf-8'))
>>> b
b'this is a test \xe0\xb2\xa0_\xe0\xb2\xa0'
>>> type(b)
<class 'bytes'>
In 2.7, you encode a series of Unicode values back to... a string? This is the source of all kinds of fun as functions expecting a string as input don't have a way to tell whether it's already been encoded or if they need to do so. In 3.5, there's no ambiguity. Functions that write to files or network sockets can either receive strings and know that they have to encode the strings locally, or bytes and know that they don't have to.Right up near the top of Zen we get:
>>> import this
The Zen of Python, by Tim Peters
Beautiful is better than ugly.
Explicit is better than implicit.
The Python 2-style implicit conversions were awfully convenient 15 years ago when most things were ASCII. They're absolutely painful now that we've moved past the expectation that strings are arrays of single bytes.Python3 is the Windows way of handling strings so the change was pure churn. Both are valid, but Linux is "pretty popular" right? So why break everyone's code just to churn the pot.
If it was good for Kernigan and Richie, good for Joe Ossanna, good for Robert Pike well it's good enough for me.
And it's not like Python 3 is forcing you to give up bytes. You can still os.listdir(b'/some/dir') and you'll get a list of bytes objects back. It's just that that's not default, since using Unicode by default simply makes more sense on a cross-platform system.
I'll take Python2's POSIX text model over Microsoft's default unicode strings every time because I prefer Linux. At this point in time see no reason to cater to a shrinking server platform.
Not to mention they didn't even get unicode right with Python3. Google did with Go (assume UTF8).
Most python2 code doesn't. It has a nice property of exploding spectacularly when the string you are writing out to console or file contains Chávez or Çelik.
Citation needed. I can't say I ever had practical issues with not being able to handle unicode on Python if I needed. No matter which library. Please give practical examples of where you cannot deal with unicode on Python 2.
for line in yourfile:
print line
with yourfile containing non-ascii characters and running it in windows console.</s>
Works fine on Linux.
I'm sorry to say, but your code for example. While I love your click library it has many small issues. Which annoy the heck out of me.
I would submit a patch but this is not a simple change and requires a bit of effort. Perhaps I'll find some time to work on it, but I'm afraid it might possibly break compatibility.
class ClickException(Exception):
"""An exception that Click can handle and show to the user."""
#: The exit code for this exception
exit_code = 1
def __init__(self, message):
ctor_msg = message
if PY2:
if ctor_msg is not None:
ctor_msg = ctor_msg.encode('utf-8')
Exception.__init__(self, ctor_msg)
self.message = message
def format_message(self):
return self.message
def __unicode__(self):
return self.message
def __str__(self):
return self.message.encode('utf-8')
def show(self, file=None):
if file is None:
file = get_text_stderr()
echo('Error: %s' % self.format_message(), file=file)
This code is arguably wrong. I think I understand what you were trying to do, but I'm not certain. There is a possibility to use it correctly in python 2 though, but many people might be not aware of it.In python 3 you pass the message as is (it might cause another issue, but about that later).
In python 2 you immediately encode the passed variable using utf-8. This means that you're expecting argument to be of unicode type, but at the same time you're discouraging users to use unicode_literals, and in most situations users will pass a regular string.
In python 3 __str__ is trying to convert text to bytes, while the message would already be text. This most of the times will look correct, but it might spew garbage when there are non ascii characters.
Here's corrected code (did not test it though):
if PY2:
unicode = str
class ClickException(Exception):
"""An exception that Click can handle and show to the user."""
#: The exit code for this exception
exit_code = 1
def __init__(self, message):
if PY2 and isinstance(self, unicode):
message = message.encode()
Exception.__init__(self, message)
self.message = message
def format_message(self):
return self.message
def __unicode__(self):
return str(self).decode()
def __str__(self):
return str(self.message)
def show(self, file=None):
if file is None:
file = get_text_stderr()
echo('Error: %s' % self.format_message(), file=file)
There are few other things that gives headaches (not necessarily python 2 only).In python 3 for example click refuses to run if LANG and LC_CTYPE are not defined. Why doesn't it simply do what all other applications are doing and simply fall back to latin-1 (ISO-8859-1) instead of printing the error.
Also, another issue (and my above code is also is affected by this). Click is checking for environmental variables before continuing, yet encode and decode have hardcoded utf-8. It probably should use whatever locale.getdefaultlocale() returns.
For some locales Windows has additional bonus, the Ansi (GUI) encoding is different than OEM (console) and it was heroic undertaking to make your program work correctly for both.
Python3 works out of the box.
For true insanity try adding IDLE to the mix. IDLE under Windows in Python 2 will let you output UTF-8, but won't accept it as input. Python 3 is more consistent.
And yes, it's not really a Python issue, it's a more generic issue with treating bytearrays as strings (what you refer to as POSIX text model). Until UTF-8 became the standard encoding everywhere, there were many Linux apps that couldn't handle non-Latin1 encodings properly, either. It was also the case on Windows in 9x days, for all the same reasons.
Dunno about you, but the amount of encoding/decoding errors in Py2 broke my little code multiple times in the past. "oh yeah I'll just fetch this html page"... all good, until that page contains non-ASCII and then BOOM.
With Py3 you're forced to think about this sort of issues right away and with minimal fuss you're sorted for life.
Using byte strings also doesn't 'just work' if the original format is a Unicode encoding other than UTF-8; e.g. Windows filenames are UTF-16 and can't be converted to UTF-8 without risking errors caused by unpaired surrogates. But that's what WTF-8 is for (or rather, what it should be for, disclaimers in the spec notwithstanding)...
I'm tired of people spreading FUD about it. they either don't understand what they are doing (and do it incorrectly) or are repeating what other people (who did things incorrectly) said:
» echo $LANG
en_US.UTF-8
» echo $LC_CTYPE
en_US.UTF-8
» echo 'works' > $'\377'$'\377'$'\377'
» ls
???
» cat $'\377'$'\377'$'\377'
works
» python3
Python 3.4.1 (default, May 23 2014, 17:48:28) [GCC] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> a = b'\377\377\377'
>>> a
b'\xff\xff\xff'
>>> a.decode()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
>>> b = open(a)
>>> b.read()
'works\n'
>>>
Accessing files with invalid utf-8 characters works just fine. Being explicit and clearly distinguishing between a string and bytes allows python to know when the filename is encoded using utf-8 (or in fact any other encoding you have defined in your environment) or bytes. Actually python2 struggles here, because it treats everything as bytes and can't tell what is used. Yes, legacy python probably won't crash, but it will spew garbage.Python3 on the other side is strict about to point errors in your code. If python3 code will crash it's either bug in python or (most likely) bug in your code and you should fix it.
I (and I'm sure a lot of people) prefer program to crash, pointing where the error is, than silently corrupt data once in a while, and when you notice the issue spending days or months figuring out where the corruption is happening.
On the other hand, if you read a list of filenames from a file, you'd better remember to either use binary or explicitly choose surrogateescape decoding. Which is a special case of "file in a text-based format" from my original post.
I don't know what you mean by "corruption".
As for your other point. You're right, but there's no easy way resolving this. Either treat everything as unicode and remember it's binary on linux, or have linux working and everything else broken.
In that scenario you probably should use .encode('utf-8', 'surrogateescape') as soon as you can and work on bytes, so no need of doing the above.
Also, if you for example call os.listdir() with argument in bytes it'll return entries in bytes:
Python 3.4.1 (default, May 23 2014, 17:48:28) [GCC] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import os
>>> os.listdir(b'.')
[b'\xff\xff\xff', b'test.py']
But IMO, things still could be better. I would love if there was matching sys.setfiliesystemencoding() to sys.getfilesystemencoding(). In scenarios when you are not sure of encoding you could fall back to iso-8859-1 (latin-1) which has 1:1 mapping for bytes.Currently looks like you would have to use:
filename.encode(sys.getfilesystemencoding(), 'surrogateescape').decode('iso-8859-1')Let me guess: you don't work on anything that has any internationalization ever.
Even if we accept your myopic list of features that work and believe your statement that "most operations can be performed on bytes", that doesn't make the language usable for internationalized content. Trading having to call encode and decode, which fail with clear, easy-to-fix errors, for a myriad of subtle, silent encoding failures is not a good trade. Which you would be abundantly clear to you if you had internationalized anything ever.
After a while I started creatively expressing my dislike of phone screens by showing them an implementation that hints at some of the complexity lurking in this allegedly-simple task (i.e., you need to understand Unicode classes and normalization in order to write a palindrome tester).
I think if you showed me an implementation that hinted at some of the complexity lurking in the task, that would be a big point in your favor.
I did bust out a version on one call which did detection of combining classes, normalized into composed form before checking the reverse, etc., and the interviewer seemed kinda lost as I explained it.
Ended up not working for them, but for different reasons.
I question some details in your lists but imagine you'd agree that the Works vs Doesn't split approximates to whole string level vs sub-string/character level operations.
This is incorrect for anything other than an ASCII-only world. In general, "just use UTF-8" is a shorthand for "just let me pretend non-ASCII doesn't exist", because it leads to people writing code that assumes one byte == one character.
By rubbing your nose in the difference between str and bytes, Python 3 complicates your life.
Having had a lot of experience with it, I would instead say Python 3 makes you think about things and find bugs at the appropriate time -- up-front during the development process -- rather than at an inappropriate time, like 3AM on a Sunday when your pager goes off because you trusted the "helpful" but wrong way of doing things.
Only if people writing code are entirely oblivious to what UTF-8 is. The most basic feature of UTF-8 is that it's a variable-length encoding scheme. If the main feature is it's variable-length, why would any programmer assume the length doesn't vary?
Because UTF-8 is the Unicode encoding that looks just like ASCII, so long as the characters stay within the set supported by ASCII, so it lets them say they're doing "Unicode" while really just writing code to handle ASCII.
This could have been done differently. I would have been tempted to store strings as UTF-8 in Python 3, but hide this from the user. Most of the string functions, including regular expressions, don't really random access a string; they're sequential and forward only. They could work on UTF-8. The ones that return a string index should return an opaque type (say, "StringPos") which has a reference to the string and a position in it. The string functions should accept that type as an index.
If someone applies "int()" to a string index, or performs arithmetic on it, or indexes a string with an actual integer, then it's necessary to scan the entire UTF-8 string and create an index of where each Unicode character starts. This should be a rare event. Especially if there's support for adding or subtracting small integers from a StringPos by walking the string forwards or backwards, rather than building the full index.
Examples:
s = "This is a test" # Unicode string, internal represention UTF-8
i = s.find("is") # returns a StringPos, not an int.
# no need to index the string for any of these
s1 = s[i] # StringPos selects the char found by find
s2 = s[i:] # No need to build an index for this
s3 = s[i+1:] # Adding a small int just walks the string
# these force generation of an index
# the first operation like this on a string is expensive
# comparable to a UTF-8 to rune array conversion
j = int(i) # converts a StringPos to an int
s4 = s[i*2] # again, expensive
In practice, those expensive operations are rare.This would cut Python memory consumption way down when processing mostly-ASCII Unicode data. It would also make it more efficient to input and output UTF-8, which is usually what you want.
(Granted, I'm still learning, so maybe I can be as efficient without relying on unsafe, but still...)
Handling UTF-8 like on Linux is definitly akward in the Win32 API. And actually that only was caused by the fact of backwards compability.
Actually it would be huge if they would force users to unicode more. But I doubt that will ever happen. We maybe see Windows Active Directory on Linux before this will happen.
The problem really only lies in file and console encoding. If they'd simply give you the option between legacy code pages and UTF-8, and made UTF-8 the default, life would be grand.
On the other hand, few years ago, when Delphi the RAD development tool, introduced their new Unicode version, completely breaking backward-compatibility, by changing the literal string type from "Ansi string" to "Unicode String". I know many third party libraries were abandoned and many projects were prevented to upgrade. It's also a disaster to me...
You can also do this piece-by-piece, if for some reason your code doesn't compile with Cython right away.
The compilation process also makes it much easier to identify bytes/unicode bugs, because you can choose to declare types.
Thank you for saying this! For those interested in this direction, there was a discussion about it on HN a couple of months ago:
https://news.ycombinator.com/item?id=10823406
Proof-of-concept examples embedding Python interpreters within themselves:
Once you decide not to support Python2 anymore, just drop the imports. They also have a Py2/3 compat cheet sheet : http://python-future.org/compatible_idioms.pdf
What about us who have and want to use massive libraries of existing Python 2 code?
"Just upgrade your libraries!" is not realistic for everyone, especially those who have to justify a mostly idealogical conversion to a similar language, and spend real time and money, outside of actually developing the end product, to do it.
Four more years of official support is a long time.
According to this article Python2 support is shrinking.
https://blogs.msdn.microsoft.com/pythonengineering/2016/03/0...
If you write `print "hello"`, you just increased the QUANTITY OF PYTHON 2 CODE. If there is one guy left in the world who occasionally writes some Python 2--and even if he's the only one left who does--the quantity of Python 2 code will, in some sense, "keep on growing", regardless of what happens to Python 3.
The article you cite refers to PERCENTAGE OF LIBRARIES that support Python 2 or 3, either one or the other exclusively or both, and it demonstrates a decrease in libraries that only support Python 2 and an increase in libraries that only support Python 3. This is presumably correct, too, but it doesn't suggest that nobody writes any Python 2 code anymore.
> Moving it in pieces is also not really possible, since there are many inter-dependencies.
I guess something is wrong with your architecture in first place, so you can't separate domains or layers easily.
Welcome to the modern COBOL then. Have fun maintaining that and running only old libs cause everyone else in the world has passed you by.
> I'm just waiting until a consensus emerges.
Wake up, it's done emerged. Almost no libs aren't avail on 3 now.
What an awkward way to say "most libs work on Python 3 now"... Why is that? I suspect it's because saying it clearly doesn't deliver the message you want.
If you cant evolve a system incrementally because of excessive coupling, that's a problem, sure, but not a problem of either the platform you are on or the one you might be thinking about moving to, unless the former forced a tightly-coupled architecture on you.
The issue with it is that from unicode perspective your python 2 code is inherently broken. There's no automated tool that can fix it, and there's no way python 5 will be able to run python 2 code correctly unless unicode support would be removed. You have to manually tell python which string is made of characters and which one was made out of bytes.
I'm in a shop with solid microservice underpinnings, so our new project could just as easily have been in Go or something else for all its clients would know or care. Given that all the libraries we wanted to use were already available for Py3, this was a no-brainer. There were plenty of reasons to upgrade and no compelling reasons to stay on Py2. Should you find yourself in such a situation, I highly recommend investigating whether you can make the same move.
Library porting to Python 3 did not go well. Many Python 2.x libraries were replaced by different Python 3 libraries from different developers. Many of the new libraries also work on Python 2. This creates the illusion that libraries are compatible across versions, but in fact, it just means you can now write code that runs on both Python 2 and Python 3. Converting old code can still be tough. (My posting on this from last year, after I ported a medium-size production application.[1] Note the angry, but not useful, replies from Python fanboys there.)
Python 3, at this point, is OK. But it was more incompatible than it needed to be. This created a Perl 5/Perl 6 type situation, where nobody wants to upgrade. The Perl crowd has the sense to not try to kill Perl 5.
Coming up next, Python 4, with optional unchecked type declarations with bad syntax. Some of the type info goes in comments, because it won't fit the syntax.
Stop von Rossum before he kills again.
[1] http://www.gossamer-threads.com/lists/python/python/1187134
Beatings will continue until morale improves...
The reference implementation of most compilers/interpreters is usually the slowest one, because it has to be legible to read the code.
So, in other words, “we don't trust our users to do concurrency correctly, and we need to keep our system safe in spite of that”?
Really, this is more about implementations than languages; a fairly-popular Ruby implementation (JRuby) supports native threading (and has no GIL); IIRC the main Perl 6 implementation also.
Elixir is up and coming in popularity (don't know if you consider it "scripting", which is a relatively fuzzy-bounded category), and definitely supports utilizing multiple processors without starting new OS processes (it uses "processes" as that term is used in the Erlang ecosystem, which are a different thing, basically M:N green threads.)
Granted, neither implementation has the popularity of CPython.
(By the way, Tcl also has an integrated event loop as well as a complete implementation of coroutines, both of which Python only very recently got.)
Second, parallelism in Python is not hindered beyond threading (e.g. multi-processing works just fine: https://www.youtube.com/watch?v=gVBLF0ohcrE).
Finally, removing the GIL is not trivial and the other changes to the language that mandated breaking backwards compatibility are pretty worth it (IMO).
Back in the days of Python 1.5, Greg Stein actually implemented a
comprehensive patch set (the “free threading” patches) that removed the GIL
and replaced it with fine-grained locking. Unfortunately, even on Windows
(where locks are very efficient) this ran ordinary Python code about twice as
slow as the interpreter using the GIL. On Linux the performance loss was even
worse because pthread locks aren’t as efficient.
Since then, the idea of getting rid of the GIL has occasionally come up but
nobody has found a way to deal with the expected slowdown, and users who
don’t use threads would not be happy if their code ran at half the speed.
Greg’s free threading patch set has not been kept up-to-date for later Python
versions.
> no parallelism for you!This is a bit of a hyperbole; Python supports both hardware threads and processes, both of which can be used to achieve parallelism. Despite the GIL limiting Python code to 1 CPU, many I/O routines will release the GIL while they perform I/O meaning I/O can be done in parallel, and native code can release the GIL to do long computations. Processes can be used to overcome the GIL directly in Python, and the standard library offers support to make this as easy as it is to launch a thread, as well as some higher-level support for parallel-mapping a function across a pool of processes.
[1]: https://docs.python.org/3/faq/library.html#can-t-we-get-rid-...
But what made the community jump, IMO, was the fact that the Rails maintainers announced that they would be upgrading to 1.9 [0]. And since there is a very small subset of Ruby users who don't use Rails, that was the end of discussion.
Is there any library in Python that enjoys as much dominance over the language as Rails does to Ruby? Not from what I can tell...And virtually all of the big mindshare libraries in Python have made the transition (e.g. NumPy, Django)...So I agree with OP that making libraries commit to 3.x-or-else is the way to encourage adoption of 3.x...but I just don't see it working as well as it did for Rails/Ruby. That's not necessarily a bad thing, per se, in the sense that it shows the diversity of Python and its use-cases, versus Ruby and its majority use-case of Rails. But forcing the adoption of a version upgrade is one situation in which a mono-culture has the advantage (also, see iOS vs Android).
[0] http://yehudakatz.com/2009/07/17/what-do-we-need-to-get-on-r...
Python 3 still seems to reward you by running slightly slower. Where's the hook? Marginal improvements to language design are bit of a weak sell for a language that's already pretty nice.
The problem was not the lack of monoculture, but that all the strongest projects took their sweet time. (To be fair, it might be that they were lobbying for the compatibility hacks that did land in 3.2 and 3.3, about 2 years after the .0 release.) There was no trailblazer among the big projects. The only real pioneers were the Arch guys, making py3 the default interpreter quite early on; but the popularity of a minor Linux distribution is nowhere near the main projects'.
Still today, some of the most prominent leaders on such projects (the_mitsuhiko, kennethreitz etc) are very much Py3-skeptics; but at least the big frameworks have moved on, so everything else is starting to catch up. We just lost 3 years.
I love Python, and I think Py3k is great. But I guess I have written too much C/C++ or something because integer division yielding integers (implicitly "with truncation") is what I would expect and not a float. And I'm a big fan of the "Principle of Least Surprise" but in this case I'm surprised to see int/int give float.
But, hey, like I said -- Py3k is a net improvement. And I'm sympathetic to new programmers and I love that Python's so popular for those new to coding.
I don't know of any of them that expect 2 / 3 to be an integer.
And as ufo says, it creates subtle bugs because it works right on most data until you just happen to get a pair of ints. Then boom.
eg in Matlab all numbers are doubles unless you take action to make them otherwise.
> There should be one-- and preferably only one --obvious way to do it.
If I divide 1 in half, I'm left with 0.5.
(/ 2 3) ; => 2/3
(/ 2.0 3) ; => 0.6666667Exactly. / in Python 3 now consistently means "real" division. If you intend integer division, you can explicitly say that with //.
Apart from breaking compatibility with Python 2, I don't see any reason why this isn't 100% better/clearer/more consistent.
I get that it's confusing and a little subtle to new folks that "1 / 2" is not the same as "1.0/2.0".
But if I wanted float division I would've coerced one or both of the operands to floats.
Those are two different functions, and both are "mathematically correct". The notation is the problem. The former operation isn't specifying the division operator, it's specifying an integer division operator.
So Python 3 fixes this by separating the two different functions to use different operators: "/" for Real (float) division, and "//" for integer division. Problem solved.
I'll agree with this. In the context of Python, the behavior is surprising.
> but in Python it is implemented in a different way that is mathematically (but not pragmatically) incorrect.
I disagree with this statement, for pedantic reasons. I wouldn't expect someone to use "/" in the context of standard arithmetic to mean integer division without explicitly redefining it, but redefining it would certainly be "mathematically correct". Look at the notation for Galois fields. That redefines all of the standard operators to work within a field of limited elements, and I'd consider the use-case to be comparable.
Finite fields define standard operators because using mathematical operations two sets of fields is ambiguous. So I don't consider that "redefining", in the sense that they are being used unexpectedly, I consider that just defining how to use them in this special use case.
Whereas Python is defining / to do two operations: divide and truncate. I call that redefining because there's a certain expected outcome if you've had elementary math, and that's not what Python delivers.
* "real" as in real numbers instead of integers.
* to acknowledge the fact that computer division is only ever an approximation of real numbers.
* real as in "what division actually means". Only in computers would you expect 1/2 to be 0. If you're doing math, you express that with something like floor(1/2) (don't have time to copy/paste floor bracket symbols).
In the context of a programming language, I have to agree with wyldfire that int-on-int operations should yield integers.
Now if SymPy wanted to override the base behavior, I think they absolutely should! I could even understand that the context of symbolic math makes this the least surprising outcome.
But Python is a programming language, not a Domain Specific Language.
x = 1/3 3 * x == 1
2. I would really like somebody to remove the preparse code from Sage and make it a standalone library available on pypi. Here's the code (happy to relicense it under BSD): https://github.com/sagemath/sage/blob/master/src/sage/repl/p...
3. Permalink to your example -- http://sagecell.sagemath.org/?q=nwzdzn
* (/ 1 2)
1/2I'm building a brand new company and I'm being forced to use Python 2.7 because I'm using Lambda. This was my choice, but the point is I can't use 3 even if I want to.
After "forcing" you to build your company on Python 2, they'll build theirs on Python 3 and too bad for you.
There are alternative python implementations like pypy, and they haven't expressed an intent to drop python 2.7 support. Only CPython 2.7 will then become an insecure interpreter.
1) porting millions (possibily billions) of lines of working Python 2.x code to Python 3
2) keeping the Python 2.7.x interpreter in deep maintenance mode?
This idea that you can force people to move to Python 3 keeps coming up and is misguided, IMHO. There are two proper solutions to this issue:
1) make it easier to port from 2.x to 3.x. We are still making progress. Python 3.5 includes %-style formatting for byte strings. Allowing 'u' as a prefix for strings was another example. Practicality beats purity.
2) make 3.x a more compelling platform for new development. The amount of goodies in 3.0 was pretty underwhelming. I certainly wasn't very excited about moving to it. Async IO is a neat feature. Keep them coming.
> Frankly, if these "carrots" haven't convinced you yet, then I'll wager you're not really the sort of person who is persuaded by carrots.
As someone that uses Python as a C replacement for one-off-scripts that work with binary objects and pure ASCII, Unicode is not a carrot. Unicode will never be a carrot. If anything, Unicode keeps me away.
If I was developing web apps, GUI programs, or data processing libraries, sure. These would all be carrots. But I'm not. The simplicity of Python 2 is my carrot.
> Python 3 does have carrots, and I want them.
You have them. But you seem to be more interested in taking away my carrots, so you don't have to worry about them any more. That's arrogance.
Without that, there may be other much more subtle bugs I'd have missed and spent time troubleshooting.
Most likely, if CPython & crowd give up on 2.7, PyPy will carry the flag of "Stable Python" forward and that'll be that.
But if Python dies, I won't really be sad. Better languages are out there. :)
Like what? I'm curious, because I've tried a lot of the newer languages out there, and I haven't found a general purpose one that I like nearly as much as Python. Elixir is really cool, but it has a pretty narrow focus, which makes it difficult for me to make the investment in learning it.
http://tednaleid.github.io/intro-to-elixir/?full#60
I think use-cases like that are included when you say general purpose.
Other languages have eaten some of Python's lunch over the 3 debacle, but Elixir is the answer to any question about "what's better".
For shops with a functional language horror, Java, particularly Java 8, is better.
My opinion on Python is that it's stunningly bad. I'll spare you the litany.
Everyone is going to use what they want. People have been whining about this for almost a decade now, give it up. This community isn't going to be on one version again, and that is NOT the community's fault. It's the core dev team and Guido's for poor decision making, technical churn and refusal to go back and fix their mistakes. There's no technical reason why CPython can't run Python2 and 3 code, JVMs and the CLR have proven this sort of thing is possible.
If there's not strong enough support of Python3, let it die. Technical churn is not supposed to survive, not live on propaganda and coercion. I don't see the problem. If he bought into 3 before it succeeded, he created his own problem. I plan on porting what I have over when 3 actually gets over the hump but not before. That's what's in my best interests, which is what everyone should do (and it's exactly what the core dev team and GVR did).
Besides, if you're not into numerical Python and mainly use it for web, PyPy makes more sense to migrate to than Python3.
If he's so committed to Python3 he can sink on his own. Guys like this, GvR, core dev team ALL think that this is a top-down affair rather than users-up. Wrong. That's the truth regarding the 8-ball he's behind and why his harebrained scheme was hatched. Python3 landed in 2008, people are just getting delusional at this point.
Don't you think that mentality is part of the problem?
Not a big deal if something like screen dies or fades. But the same is not true of something as big as Python.
Open-source is democratic in that you have a voice, but it doesn't mean your opinion matters much in certain cases.
Python2 hardliners have been warned for years. By the time it is sunset it will have been 12 years since it's release. Why should the Python devs stop moving their project forward because you're not interested refactoring and deprecating old code?
Where I work we move forward and adapt our code/patterns over time. If an old thing breaks (due to updating to Py3 compatible features), we deprecate/replace or refactor it to work, whichever the team decides holds more value for us as a group. There's probably very little code in there that was written 8 years ago when Python3 dropped. We iterate forward.
I get pretty frustrated by this "but this stuff I wrote a decade ago works and I don't wanna..." attitude in the tech world. Not saying people should burn the candle at both ends to keep up, but come on. 12 years and you can't refactor forward? Then you find yourself in a position where, oh shit we really have to burn the candle at both ends.
This isn't about the actual work involved then. Because you're going to write code anyway. It's about not wanting to change your headspace (not all that much, esp for experience programmers) to write Python3 compatible code. The burden then is on Python hardliners to keep up or fade away, much in the same way you suggest.
Old is replaced with new. It's the way the world works. Get over it.
Bottom line is that you believe newer is always better. That's not true at all, especially in programming languages.
The popularity of major open source projects in Python have been ported or are in progress of porting over to Python 3 (mostly 3.3+).
I remember my boss telling me a few days ago in his previous job the codebase was still running Java 1.3 or 1.4. There was no incentive to upgrade because people were afraid of breaking changes, even from to the next major release. So when are they going to port over to 1.4? to 1.7 and then finally 1.8? God knows and trust me that'd be the day IBM AS400 support finally RIP. You can choose to live with that, but the pain of actually maintaining a codebase running on obsolete platform is pain. But ask Ansible maintainers, writing code that has to work with Python 2.3 was brittle. Ask anyone who has done RedHat administration they know too well how brittle working with a system totally 10 years ago (even though RH does support these old version for years). Pain in the fucking ass.
> If there's not strong enough support of Python3, let it die.
Where did you or the author get that? There is a strong motivation to go to Python 3. It just so happens people are still running on Linux that ships with Python 2 by default, or on MacOSX which is still on Python 2 by default.
If you maintain Django and you don't want to drop Python 2 support, you make the call. You don't have to join. He doesn't force people to join his collation. You are making a big deal out of his post/tweet.
Relying on the Mac default Python keeps my binary sizes small, and Python 2.7 is good enough. While I have at least done the __future__ imports to adopt features of 3.x such as print(), a full switch to 3.x will be more trouble than it’s worth until Apple decides to install it.
The bigger reason is to make the application as useful as possible to Python extensions. If you don’t force scripts to use a particular interpreter then the entire application can be imported into almost any script, alongside any number of other modules to do interesting things. The only module that has to be compatible is the application; if by chance the application module doesn’t work with one Python version, another copy of it can exist for two Python versions. Whereas, if your application only works inside a special interpreter binary, extensions have to reinstall special versions of every module that they want to use, and possibly port an entire stack of code to support that Python version. Both approaches are workable but an unspecified interpreter is by far the more convenient way.
Debian Stable and pyenv it is. That way, I have a rock solid operating system and my application are completely decoupled.
Using distribution packages for application deployments is insane. It makes sense for workstation or systems software where the application you develop is part of the operating system, but for a server application, it makes no sense at all.
What if you are happy with the current version and don't plan on changing it before changing the OS? What if you don't install more than one app on a system, or are using containers? Then you might not need the extra complexity.
No virtualenv's, no tarballs, nothing special beyond an RPM. I won't deploy software any other way (and if something doesn't already have an RPM I will make one for it).
I don't, I just use the standard one that comes with the operating system instead of insisting on re-inventing the wheel because "packaging is hard". I could have just had tito push directly to a yum repository if I so desired, but it's much nicer to use Koji (which I'm already familiar with).
...which is "ignoring all the tooling people have invented to make this stuff easy". virtualenv and pip were explicitly designed to work around two damn nigh insurmountable problems:
1) OS vendors move slowly, and 2) Sometimes two projects on the same machine need two different libraries
It's trivial to use virtualenv and pip to make Project A and Project B run side by side with different Python versions and completely different dependency manifests. You can come up with your own ad-hoc system to deal with this or you can adapt the best practice of every other Python shop. Make no mistake though: your current way is deliberately choosing the hard road.
Virtualenv doesn't make my life any easier, it makes deploying applications harder since I have to write a custom deployment pipeline - whereas I already maintain RPM's for other non-python software we use internally, and I can use the same infrastructure to handle them both.
Edit: For most of our applications the latest Fedora Server release is also our target host, so "distributions moving slowly" is rarely an excuse as typically the latest packages are available (and if not we make them ourselves, like we do for selenium).
Eggs are older than virtualenv, from the times of easy_install. Now there are wheels that are nicely integrated with pip.
Also virtualenv became integrated into python (it's called pyvenv now).
Among many things you have all versions of python available that don't conflict with system packages or themselves.
What's awesome is that you can install multiple versions and all of them just work.
I'm a casual user of Python, and a heavy user of OS X command line, and it will be really nice if/when Apple finally changes this.
[1] https://github.com/twisted/twisted/blob/twisted-16.2.0/twist...
They've prudently made their last Py2 release a long-term support release, but if you want new Django features after December 2017, you're going to be using Python 3.
Django works great on Python 3, BTW, and if you've got plugins that haven't updated, I would recommend being a bit suspicious of their code quality.
[1] https://www.djangoproject.com/weblog/2015/jun/25/roadmap/
The way I see it, Python 3.4+ is a lot like Windows 8.1. It's actually really not bad, but the narrative around its predecessors will never let it truly succeed. The realistic way for python to move forward is to offer a brand new version (5 or 4, doesn't really matter), give some concessions to the python 2 people that will make them feel heard and feel reinvested (bring back the print statement, give us some long wanted things like not-stupid lambdas), and guarantee python 3 compatibility so they're not forking the damn project again.
I think at this point so many people have said bad things about Python 3 that the brand is toxic. A new version number might not really mean much technically, but symbolically as a sign to the python community that the devs have finally stopped being stubborn and arrogant and acknowledged they fucked up, it would be huge.
I can't say I've come across an issue caused by an incompatibility between the two since the first couple months after switching...
You don't have this problem with PHP. "Hey, Boss! We need to rewrite our code for PHP 7." – "Will it be faster? How much?" – "It will be faster. By factor X." – "Go on, make it so."
Some bosses care about syntax, at least if you phrase it as maintainability.
That are doing this as a hobby. And don't have extensive time.
Why guilt us?
We are neither companies nor fundation's bitch.
Packaging in python is a mess, and every one wants to contribute to the shiny new features, and very few have to deal with the toilet cleaning that packaging (pypi, deb, rpm..), maintaining, dependency management, bug report platform consistency involves.
I am a bored as a maintainer to be expected to deal without a thanks for all the shit coming from the social pressure of "you should do this or blah" to comply with esoteric unproven needs that only results in more work and just more mess and layers of bureaucratic ideas disguised as "best practices".
I am no one's bitch. I am an open source coder. I code what I want, when I want, at the speed I want, and I am no slave that will do what he is being ordered.
But yes, the "now there are two ways to do it" brought on by 2 and 3 really sux.
I should also say I always hated the "small" syntax changes even between the point releases. This was also something that kept me away (ruby has this problem too). Perl never had this problem, stuff written in 2001 can still run on perl of 2016. But I see now that this is a double edge sword. I have worked for several companies that run modern perl versions but refuse to touch their code base that was written like it was 1998 because it just "still works". That was another reason I started looking into python. I very much welcome complete breakage every few years at this point. It will at least force people into staying modern with the stack they use instead of becoming a 15 year old fossil with a disgusting resume. (I mean, what would you do if someone sent a resume to your dev shop that said they have only done CGI perl for the last 15 years and have no idea what javascript is or why they should learn it?)
"Ubuntu 15.10" - Python 2.7.10
"Ubuntu 14.04.4 LTS" - Python 2.7.6
Not sure where you're getting the "by default 3.5" from.
> python3 --version
Or: http://packages.ubuntu.com/xenial/python3But I do feel like tides are shifting and Python 3 is slowly "winning". It's just taking, what, almost a decade longer than most would have hoped?
I find it very easy to write code which is simultaneously Python 2 and 3. A few ``from __future__ import`` statements and you'll be fine. The only reason some people are complaining is that they were ignoring bugs in their code regarding text data. It's not a Python problem, it's a problem with computers -- bytes are not text, but many systems ignore the difference.
Many large enterprises are starting new projects in Python 3 exclusively. It's fine if you want to use some Python 3-only syntax.
No it's not. The syntax is what defines the language:
In linguistics, syntax is the set of rules,
principles, and processes that govern the structure
of sentences in a given language.
It doesn't have the same syntax. It's not the same language, by definition.What impact does one or other way of thinking (absurd vs has merit) about this (a potential break in backward compatibility between language versions due to any addition, removal or alteration of one or more language features) have on the language and its ecosystem and culture?
Imo:
* such potential breakage is relatively benign iff there's an easy-to-follow technique (eg using a linting tool or a program editor's search feature) that always ignores/accepts code compatible with the older language version and always catches/rejects code compatible with the newer language version.
* a further improvement on this is to have the inter-version checking done by a language's interpreters/compilers.
* best of all is when the language is at least partly compiled and compilation using the older interpreter/compiler does the rejecting at compile-time.
When we say it's the same language, most of those aspects are the same from Python 2 to Python 3.
The biggest issue people are having is to convert their python 2 code to work on python 3. There's virtually no downside to use python 3. All maintained libraries work with python 3.
I plan to keep using Python 2.* until the sun grows cold or I die (whichever comes first.) So this entire effort seems like a busybody with nothing better to do working hard to screw me over.
Knock it off!
Roughly 12% of PyPi has been ported to Python3[0] and about 10% of pip downloads are Python3[1].
Most everything you would want already exists on Python2. A lot of those uploads are forks or new projects that already have solutions in 2. I have no problem with Python3 and glad to see people putting their money where their mouth is and 'get to porting'. Rather than blogs/complaining about the userbase/backroom deals and conspiracies to kill Python2.
When there is a competitive advantage to adopt the latest best practices, they will. Maybe not this year, but the day will come when even the legacy python developers will jump.
https://news.ycombinator.com/item?id=8730156
https://news.ycombinator.com/item?id=7005711
It's not a very productive use of time.
Personally I tried upgrading a couple of times, but partial or no support of Python 3 from libraries I depended on, stopped me from doing it.
http://python-future.org/imports.html
Basically don't worry about the MOOCs being in Python 2, the differences are trivial and you don't really need to understand them to be going down the right path.
The older styles of these languages are still relevant to those who will need them for professional use on older code bases, but for a new user, they can be treated as more advanced material to be learned primarily for recognition but not for writing.
Learn Python 3 now; add some Python 2 details later, if needed. And if you really love your Python 2 MOOC, go ahead and use it as an introduction. Almost everything they are showing you is both Python 2 AND Python 3, so you're learning Python 3. Then move on to intermediate-level material that is Python 3 and goes into a lot more detail, and you'll easily switch over.
I'm not a happy camper about the whole thing. But again, I'm not someone who even understands all the intricacies of programming languages, I just want to use the tool to make my life easier. Having two different versions of what is supposed to be the same programming language doesn't do that.
And at least in the case for something like Python (I'd argue it is very much an "every man's" programming language), this is a failure for these non-programmer users. IMO.
When talking to management or client it is useless to name any design improvements, but as soon as you mention speed they will say "yes, let's do it".
The issue with Python 3 is that the performance is not that much better.
For me though, what's the point? I like that lru_cache is a convenient decorator rather than having to roll my own memoisation and that some relatively solid steps have been made towards async programming but that's about it. Its not like the changes were things that couldn't be achieved before (Twisted is solid, from what I have seen so far). I still have to factor in the GIL if I want performance to a certain point (dual core processors have been around for almost two decades now - I am starting to question why cPython is even considered a serious choice for use cases where prod has more than one core).
More crucially though? The APIs. camelCase features heavily (in spite of PEP8), the principle of most surprise is rampant (I recall last week discovering a function signature with infinite arity rather than passing a collection, like, you know, every other Python method) and SO many more warts. When interacting with a file, try and guess what you want: "readlines? Doesn't load each line into collection, which is what you'd expect. writeline? Be sure to manually interpolate a \n because in contrast to the name, it's actually going to insert everything on the same line because... well, who knows. writelines? Well, you know how in Python you usually iterate over over thing and handle it? Well here, you pass a collection. There's a corresponding read meth... No there isn't". If we're introducing breaking changes, how weren't these things fixed? You either commit to breaking changes or you don't make any - py3 trod the line somewhere in the middle and is paying for it with its adoption numbers. The broader point is that this mess is exactly why we all love libraries like requests. If Python were a Fortune 500 company, GvR would have been ousted as CEO long ago, and rightly so.
As a contractor, I have to absorb this new way of doing things in case I find a job that actually requires Py3 knowledge (in the UK at least, these are less frequent than HN prefers to make out) because my marketability depends on it. That said, although I'm currently contributing to Pypy, I realistically see my future in Clojure, Ruby, Elixir or some other language that I'm prepared to put the time into picking up, in contrast to Py3, where I see missteps and hamstrung APIs as active disincentivisation to my learning. Any half decent dev will just move on, because why would they put up with this? Did we not learn anything from PHP?
Yes. Imagine the bewildering inconsistency when a newcomer who tried any of pascal, c or a myriad of other languages try their hand in python.