A simple example is the print statement, in Python 2 you could do:
print "Hello World!"
Whereas in Python 3 you have to do: print("Hello World!")
Now obviously, you can easily change this programmatically. Things get harder when you use more advanced print statements, for example adding a "," at the end to prevent a newline, printing to different targets, new ways of using placeholders, et cetera.Then there is the fact that strings are no more in Python 3, but they are Unicode. Again for average strings that's not really a problem, but when you start using bytes, special characters, et cetera.
There is a tool to translate programs from Python 2 to Python 3, but it does not catch everything, and you still have to fix stuff manually.
In my view, those that made the decision took to seriously "there's only one way to do it." Zen of Python actually says: "There should be one-- and preferably only one --obvious way to do it." Even if the obvious way for 3.x can be different from 2.x, I can't imagine any real-life effects from allowing the alternative syntax for formatting and printing except for a few special-case branches somewhere in the source. Yes, then the people would continue to write it with the "wrong syntax" in the new programs too, so what?
You could try to make the interpreter guess when you want to use print as a function and when as a keyword. But that would be horribly complicated, and is bound to go wrong.
No it wouldn't. There would be some corner cases, but mostly it would just work for the existing code whereas, from the perspective of the 2.x users, now it just doesn't. I know, I write compilers for living. The maintenance cost from the point of view of the compiler maintainer would almost invisibly increase (there are much less trivial things to worry about) the benefit for the current users would be significant. The reason it wasn't done is much more "political" ("just one way to do it") than technical.
Specifically: once it's declared that print is a function, you don't need to treat the string print as keyword. Then you can notice comparing print expression and print( expression ) that if you know that print is a function the braces aren't giving you any new information. So the difference is do you want to encode the knowledge "print is a function" in the compiler or not. That encoding is trivial, and even if it can be called "a special case" isn't anything that anybody would spend any significant energy maintaining. It obviously appears to be "less elegant" to describe your compiler having "a special knowledge that print is a function" but there are even ways out of that: you can generalize such constructs (function calls without using the return value). But then "there would be more than one way to do it."
But orders of magnitude more 2.x Python programs would "just work" when started under a such 3.x. Of course, once you accept that the transition should be less painful, you'd need provide the way for libraries to also have the "newer" and the "older" ways to do it. "More than one way" is potentially contagious. But, sometimes "worse is better."
print
As Python 3, this program does nothing. As Python 2 it prints a newline.I do agree that Python could have adopted the ML/Haskell syntax for calling functions that does away with most parens. But I don't think anyone in Python land would have swallowed that.
Lesson for language designers: never, ever mess with your debugging/printout facilities once you've hit the mainstream.
print "%s remember to %s in %s if %s but never %s" % (some_random_tuple)
vs
print("{when} remember {what} in {room} if {condition} but never {dont}".format(* * dict_with_explicit_keys))
A bit verbose but much clearer and future-proof.
print "%(when)s remember to %(what)s in %(room)s but never %(dont)s" % dict_with_explicit_keys
was supported back in at least 2.4.The only actual difference in that example is that print became a function in Python 3, requiring some additional parentheses. (Something that can be mechanically translated without much hassle.)
New Python fixed that.
Problems: 1) new Python wasn't very tolerant of real-world conditions where incorrect text encoding happens once in a while.
1b) there was great resistance to fixing that because "Python now did it correctly!" and the rest of the world was just assumed to do it correctly, always.
2) not everybody who already programmed in Python really understood the need for the new (and mostly correct) way of handling strings. Also not everybody in the US really understood Unicode and encodings.
3) there was no proper update path from Old Python to New Python.
4) it was not possible (or very difficult) to write code that was both valid Old Python and valid New Python and which did the right thing in both cases.
5) at the same time the interface for libraries written in C was changed.
5a) the new way was better.
5b) a change was needed for New Pythons string handling anyway.
5c) it could still have been done in a backwards-compatible way...
5d) ... but it wasn't, since the new way was Better.
6) lots of important Python libraries are partially written in C because pure Python is so slow
7) ... so porting all the important libraries was necessary for New Python to take off while at the same time being rather annoying and difficult work.
There were other changes at the same time that were improvements but which made upgrading hard. Print was no longer a keyword with lots of special handling and corner cases but an ordinary function. You could switch newer versions of Python 2.x to the same behaviour but not the older ones. Many libraries needed to work with many different Python 2.x versions, which made it hard to support Python 3.x at the same time. You could convert all your print statements to using a function that worked like the new print function and then either use that natively (3.x and newer 2.x) or provide a compatibility function that wrapped the print statement (older 2.x). You would put this code in a module you would selectively import depending on the Python version -- or some variant of that. Not actually hard but really annoying and something that required changing many lines of some libraries.
Ad 1, things have improved a lot and there are now known ways of fixing the remaining problems (with newer 3.x versions).
Ad 2, Unicode is better understood these days.
Ad 3 and 4, things have again improved. If you skip the first few 3.x versions and the early 2.x versions then it is not too hard to write code that works on both Old and New Python.
Ad 5, I don't know if is possible to support both ABIs in a single shared library, but at least the core code that actually does the work can be shared.
PyPy might have helped as well by making ordinary Python -- the kind you would otherwise write in C -- run fast.
Edit: formatting is hard.