Let's Write an LLVM Specializer for Python
dev.stephendiehl.com
dev.stephendiehl.com
To me, the killer feature is better lazy evaluation (generators). In particular, important builtins like map, filter, zip, enumerate, etc are generators, instead of returning lists. This makes it feasible to write things like
(process(line) for line in map(str.upper, open('giantfile.txt')) if line.lstrip()[0] != '#')
Some of the above can also be done with itertools package in Python 2, but not everything.Python 3.4 changelog is here, it contains e.g. asynchronous io facilities (asyncio module): https://www.python.org/downloads/release/python-342/
edit: added enumerate() in the example above, for line in open(filename) returns a generator in Python 2.x too.
edit2: enumerate is lazy in python2, I replaced it with map(str.upper)
But the point should be obvious, without generators that would potentially consume a lot of memory.
Generators provide a convenient syntax to implement that sort of object.
(process(line) for line in open('giantfile.txt') if line.lstrip()[0] != '#')
Is that line really using any new features in Python 3? The lazy evaluation there is in the file object and the generator expression, both of which have long been present in python2. (process(line.upper()) for line in open('giantfile.txt') if line.lstrip()[0] != '#')Here's another one you can't change that easily:
tests_pass = all(process(input) == output for (input, output) in zip(open('inputs.txt'), open('outputs.txt'))The fact that a bunch of builtins and the values/items methods of dictionaries have become iterators is not very siginificant IMHO. Python 2 code could already be written to use iterators or generator expressions, so in the parts where it was crucial it was already done. In this regard Python 3 has not added new functionality but only changed defaults.
The unicode change is the big one.
Changing that to line.lstrip().startswith('#') would be an alternate approach.
The key point is that 2.7 is a language frozen in time, while 3.4+ is continuing to develop and improve. And most of the hand-wringing was before the critical mass of third-party modules was ported to 3.x.
https://docs.python.org/3/whatsnew/3.4.html
https://docs.python.org/3/whatsnew/3.3.html
https://docs.python.org/3/whatsnew/3.2.html
- function annotations (allows runtime type checking via third party modules)
- asyncio (not as easy to use as Go's goroutines but still vastly superior to the multiprocessing module)
Storing types via traces could be another step for gathering types. As well as using the more advanced static type checking code that is around for python.
Now I have something to work through on the weekend. Looking forward to part 2!
The author is definitely helping people learn about LLVM and how it can be used with Python --- which is great, because this is exactly what Numba is: http://numba.pydata.org. But, please don't start another "Numba". Just come help us improve the current one.
[1] http://www.wedesoft.de/hornetseye-api/ [2] http://www.wedesoft.de/downloads/thesis_wedekind.pdf
EDIT: In my approach I didn't go through the Ruby AST though. Rather I used the approach of injecting "GCCVariables" which emit C code instead of doing the actual computation.
I'd love to see a new language exactly like Python but compiled and statically typed. Something similar to Cython, but rather than generating a bunch of C code it would target LLVM. Additionally it would be able to generate pure Python code simply by removing any typing syntax.