Porting a historic Python2 module to Python3
lucasg.github.io
lucasg.github.io
I'm not blaming the person who builds the machine in the end, but the person that sets up the specification – they are doing a bad job.
All the projects I've worked on moved to Python 3 and never looked back.
VTK itself recently does support Python 3, but it's not in Debian/Ubuntu yet. And we'd have to upgrade our code to VTK 7. (Now 8, I believe, seems we've skipped a version.) So... lots of inertia all around, basically.
The result is that in the meantime (meantime being like, 2 years), we compile Python 2 and Python 3 versions of everything.
Not to mention that it's hard to justify breaking user's code that we aren't in control of just to tell them to fix their print statements -- though obviously we'll get there eventually.
When all dependencies were there, we made the switch.
The only times it doesn't work is with python that is bad to an unusual degree. For example, if you make a variable named "list", overriding the built-in list type. 2to3 freaks right out. The other type of "bad" python is stuff that abuses the "open kimono" nature of the language and heavily monkeypatches python internals. Or things that depend on version specific binary data, like pickle files.
This is absolutely not the case in my experience.
2to3 tends to produce something that appears to work at first glance, but it actually doesn't.
Second, those only do the fairly straightforward changes, fixers are simple AST transformations they don't do complex type analysis or anything, you should use them as a starting point but don't expect the work to be over by then: the 2->3 transitions has a number of API and semantics changes which these tools can't handle, text model changes are the biggest one[0] but they're not the only one by far.
If the project is non-trivial, having a good test coverage is absolutely crucial.
Basically, 2to3 handles the first 90% of the work which are mostly drudgery and syntactic fixes, you're still on the hook for the second 90% which is more subtle.
Source: finalising the conversion of a ~200k SLOC codebase to cross-version compatibility.
[0] not just the strict separation of text and bytes, but APIs being set to one or the other so you have places where you want to ensure bytes, others where you want to ensure text and yet others where you need "native strings", some APIs (csv) are also incompatibly altered.
This works in python 2. It silently fails in python 3 because b"cookie" and "cookie" are different keys. Not really related to whether the original code supported Unicode or not.
What bites me all the time is the division operator. The 2to3 code will run, but produce incorrect results. String / byte handling is an obvious other one.
Is there a more polite way to go about this? I submitted some "Py3 support"-style pull requests to a repo out of the blue a while back. In that case it was just a few minor changes to some fairly small scripts and I did those by hand, preserving Py2 compatibility. But I did wonder at the time whether it wasn't a bit jarring for the owner to see a stranger pop up with a bunch of pull requests like that.
This is the usual process for submitting a large change to a project:
1. Check the issue tracker or mailing list for an existing discussion about Python 3 support. If there is one, read the discussion. If there isn't one, start the discussion. Mention that you're willing to make the changes if there is interest. Remember that supporting both Python 2 and Python 3 places a burden on all future contributors.
2. Keep the changes minimal and easy to review.
3. Review the changes yourself before sending a patch/pull request.
4. Run the test suite.
5. Test your changes in all versions of Python the software supports.
Forgetting #1 is usually the cause of tension when people submit large patches. The person who submits a patch wants to see their patch included. The authors of the project want to make sure it aligns with the goals of their project and doesn't break their software.
There's little reason to though, generally you want 2.7 and >=3.x (the simplest is 3.5, 3.4 is fairly easy, 3.3 is possible but starts being more annoying)
> As far as I know, 2to3 doesn't produce py3-only syntax.
It does: print function, metaclass kwarg. It will also generate syntactically valid but P2-incompatible code (with_traceback calls) as well as possibly semantically break P2 code (rename __nonzero__ to __bool__, basestring and unicode to str, stdlib changes, raw_input to input, …)
print("a", end=" ") is not, neither is class Foo(metaclass=Bar)
And...that's exactly why you want static typing. Code review to catch stuff like this is poorly and primitively doing a job a computer can do for you perfectly. Dynamic typing is fine for small projects but when you're wanting to port over all Python 2 code that's ever been written to Python 3 types would have made the task much less daunting.
Look at c. It has static typing, but fairly weak types. Strings are represented as a series of bytes (char *). The exact same work would be required to port c code to semantic c++.
The "bytes" that was put into 2.6 was just an alias to str as hinting for the 2to3 utility.
And as mentioned, the byte "backported (sort of)" to 2.x acts and works differently than the byte in 3.x. So there actually 3 types in play, plus bytearray too.
I just think it's a good example of how strong static typing makes refactoring large projects significantly easier compared to strong dynamic typing. Getting everyone to move from Python 2 to 3 is basically a massive refactoring task.
The solution -- introducing a new type which explicitly is not to be used in string-like fashion and does not expose the same interface as 'str' -- was implemented in Python 3, and does not require static typing.
Also, I've been involved in porting very large codebases and have never felt a need for static typing to assist. The "dynamic typing is only for tiny puny baby child's toy programs" meme does not reveal your sagacity; it reveals your ignorance.
You'd have to change type definitions then...if you define the type to allow those operations then obviously a static type checker isn't going to help. From the article they have this Python 2 -> 3 example which seems like static typing would catch:
self.streams = zip(sizes, page_lists)
and # zip return a generator on Py3
self.streams = list(zip(sizes, page_lists))
> Also, I've been involved in porting very large codebases and have never felt a need for static typing to assist. The "dynamic typing is only for tiny puny baby child's toy programs" meme does not reveal your sagacity; it reveals your ignorance.Why would you not want assistance from the computer to automatically check things for you given the choice? Why would you not want ways to automatically refactor your code? Those are good reasons, not memes or ignorance... Basic refactorings like changing an ID field between being a number and a string type are just horrible with languages that have dynamic types and implicit conversions for example.
This is why if you're going to do it, a dynamically-typed language that allows adding hints/annotations later on is an ideal choice.
And this is not a purely hypothetical objection: once you put types in a program, you're constraining the future of what you can do unless/until you go back and change them. Those constraints are a very real and, in my experience, very high cost.
What high cost do you mean? The extra time you spend writing type annotations and getting your program to compile is time you save puzzling through runtime errors. Python, JavaScript and even PHP all have ways to get static checking from the use of types. Why would you not want to take advantage of this and rely on less reliable dynamic checks instead?
> My IDE shows me right away, whenever I have the types wrong and with all the annotations visible I find it pretty easy to reason about types.
Do you find it makes you more productive? Honestly not trying to troll but my feeling is people that defend strong dynamic typing just haven't used strong statically type languages enough because the benefits are painfully obvious to me when you switch between them.
If you don't like to have it in your code base, then you could use it for the migration phase and remove it afterwards, or just leave it in the parts of the project, where it was helpful. It's totally optional, meaning, you can use it just on some parts of your code without affecting the rest.
> Do you find it makes you more productive?
It makes me a lot more productive. Especially when I touch code that I haven't seen in a while. I use it everywhere. Actually my IDE highlights items that I forgot to annotate.
> Honestly not trying to troll but my feeling is people that defend strong dynamic typing just haven't used strong statically type languages enough because the benefits are painfully obvious to me when you switch between them.
I'm not defending strong dynamic typing. I just think that the typing module makes Python a better language. I use the typing module for the same reason I switched from JavaScript to TypeScript and the experience I have working in both, Python+typing and TypeScript, is very similar, even though TypeScript is strong statically typed.
I think you've misread me, I had the exact same experience with both languages and agree with you.
And yes, 2to3 fixes 90% of porting problems. But in my experience, it was especially bad at handling, of all things, import statements. Definitely not a silver bullet.
I've worked in academia (as a research assistant) before so I'm familiar with shitty grad student software. But going to the effort to document and present on the software without making sure it even worked? That seemed excessively stupid to me.
> document and present on the software without making sure it even worked? That seemed excessively stupid to me
It sounds like that individual is responding rationally to the incentive structure he's embedded in, and succeeding as a result.
HA ha ha ha ha. No. Not even a little.
There are a bunch of reasons for this, but the big ones are: 1. Scientists aren't developers (some of us are exceptions, but we prove the rule). 2. The objective function of a grad student is "publish and graduate" not "write good code and use it forever".
The bytes/string changeover is probably the biggest pain point for RE tools written in Python2. Indeed, despite lucasg's excellent work porting things over, there was still at least one byte/string issue that needed to be fixed:
https://github.com/moyix/pdbparse/issues/39
Edit: Also, I have one small correction – pdbparse is actually 10 years old, not 5! I started writing it my first year out of college, which also explains (in part) why the code is a bit crap. As evidence I offer the first article I wrote on the PDB format:
http://moyix.blogspot.com/2007/08/pdb-stream-decomposition.h...
Can't comment on the site itself, so: PEP 484 - Type hints ( https://www.python.org/dev/peps/pep-0484/#suggested-syntax-f... ) does suggest Python 2.7-compatible type hints, and PyCharm, and probably other IDEs, do support them (e.g., https://www.jetbrains.com/help/pycharm/type-hinting-in-pycha... ).
Do: introduce and require the compile step for awhile, flesh out the language and build it up, get by-in and an ecosystem while also leveraging compatibility with python2/3 ecosystems -- then eventually release a python4 that runs this new typed python language directly!
Bring us the beauty!
But some other projects are so hardcore about CLA enforcement, they won't even accept one-letter typo-fixes without CLA signed.
+1 for articulating this bit of funny business in the Python Windows ecosystem... "Moreover on Windows the building process was historically so bad that you usually end up downloading binaries from some random guy on the Internet."
Having said that, I know the decisions were not made lightly, and were backed by logical analysis. If nothing else, it's a good case study for the evolution of future languages (looking at you golang 2.0).
I haven't seen them used anywhere, but that's an okay solution I guess.
NB : by the way, love your blog :p
> NB : by the way, love your blog :p
Thanks :)
Also, that doesn't help those of us stuck on Python 3.4, the last version that supports Windows XP (don't ask).
While it was unrealistic for Python to be that aggressive given its larger community, not forcing the issue created a lot of unnecessary work for the community. A benevolent dictator should have moved developers along faster. Having the language fragmented for this long is extremely unproductive.
There was no backwards compatibility with Python 2 and it had relatively little adoption.
They added Python 3 features to Python 2 over the years to encourage adoption of Python 3 as the standard. It worked.
Now we just need Apple to ship Mac OS with Python 3 instead of 2 as the default and that'll push adoption for developers.
Did it really? My impression was the opposite: the fact that so much of Python 3 became available in Python 2 made people feel like there was no point in moving to Python 3.
On macOS 10.12.5, released 2017-05-15:
$ /usr/bin/python -V
Python 2.7.10
That's a release from 2015-05-23. Latest is 2.7.13 released 2016-12-17.And no system python3. Only through Homebrew/Macports.
Do you think that means they don't care about Python? Perhaps it's not a strategic language for them?
No one said they the provide a full Unix system or keep everything else up to date. Apple has a very narrow focus on what's important to them.
Django's next version will be the first that doesn't support Python 2 anymore (because that Django version will be supported longer than Python 2), I hope that that will start the mass updating at my employer's.
https://en.oxforddictionaries.com/usage/a-historic-event-or-...
Which we do (in Britain and the rest of the English speaking world outside North America), so "an historic" it is.
e.g. from BBC:
"An hour" on the other hand makes more sense, as that is truly a soft H.
$ ./anorack test_file
$ out:1: an historic -> a historic /hIst'0rIk/
Interesting. What if...
$ ...hacks phonetics.py to use "en-us" voice...
$ ./anorack test_file
$ out:1: an historic -> a historic /hIst'0rIk/
Hum... same result. What about:
$ ...hacks phonetics.py to use "en-wm" voice...
$ ./anorack test_file
$
Ah, finally someone gets it right :)
(just having a bit of fun, this is a cool work!)
Edit: formatting...
No need to feel ignorant; it's a nonstandard, eSpeak-specific notation.
$ espeak --voices | grep ' en-'
2 en-gb M english en (en-uk 2)(en 2)
5 en-sc M en-scottish other/en-sc (en 4)
5 en-uk-north M english-north other/en-n (en-uk 3)(en 5)
5 en-uk-rp M english_rp other/en-rp (en-uk 4)(en 5)
5 en-uk-wmids M english_wmids other/en-wm (en-uk 9)(en 9)
2 en-us M english-us en-us (en-r 5)(en 3)
5 en-wi M en-westindies other/en-wi (en-uk 4)(en 10)History: /ˈhist(ə)rē/
Historic: /hiˈstôrik/
The 'h' is stressed in 'history', not in 'historic'. Try pronouncing 'historic' with a stress on the first syllable, it sounds wrong. Since it's softer in historic it's more natural (to me anyway) to use 'an'.
(Also, I'm french so I barely pronounce h's to begin with - so in my speech the stressed 'h' in history is kinda soft, and the unstressed h in historic is barely there.)
(edited)