PyPy 4.0.0 Released – A Jit with SIMD Vectorization and More
morepypy.blogspot.com
morepypy.blogspot.com
For me, targeting development on PyPy has taken the place of migrating to Python3. Which just moves too fast in effort to pull people over, Python is too big for that strategy. At this point, people want speed with a great language, they want PyPy! The PyPy team has pretty much enabled the perfect development stack. I use PyPy+Django+gunicorn+Nginx and no problems.
Keep up the great work.
If prestige and money are desirable things, absolute no brainer if they have a business bone in their bodies. It would also secure the future of the stack for everyone concerned with Python3 threatening to cut off Python2. It could also enable a bigger return on those of us who have invested financially into PyPy development. It would certainly snowball even more financial investment.
It would then be the CPython3 project's move. Abandon and integrate all the non-breaking Python3 features into 2.7? Or just move forward with a significantly weaker and less relevant project, working against their past selves? I don't think most would really care and it's best they just move away with their own project.
Could be on the way judging from the favorable naming convention, PyPy4? Has a nice ring to it.
"the Python2.7 compatible release — PyPy 4.0.0 — (`what's new in PyPy4.0.0?`_)"
"the Python3.2.5 compatible release — PyPy3 2.4.0 — (what's new in PyPy3 2.4.0?)."
The PyPy gods are in an enviable position for sure. If they seize the opportunity it's the chance of a lifetime.
I don't get this hate for Python 3. Yeah some things changed, yeah having a big codebase in Py2 sucks if you don't have the time/resources to switch, but starting new projects in Py3 by default is awesome. Asyncio and async/await are fantastic and worth the switch alone.
Basically 3 gives us lots of headaches, very marginal improvements, an ivory-tower sense of "we know best" (message to Guido - no you don't), and a real concern that the powers that be have no clue what's going on at Golang Towers, Fort Javascript, not to mention Clojure, Julia, or even Elixir.
As annoying as the transition to Python 3 is, it's the future of the language. I was actually one of the curmudgeons who stuck with 2.7 and grumbled about not wanting to switch until a few months ago. I've accepted that I need to, at the very least, start new projects in Python 3 wherever possible. The vast majority of popular third party libraries have been ported over at this point. Python 2's death is inevitable, even if it will take many more years.
Let me warn you. If you persist with your propaganda on Python3, then GOLANG will be the future of Python. Does that register with you guys?
Nobody who uses Python 2.7 all day every day thinks it has any of these imaginary "worts" you 3.x guys like to pitch incessantly.
Maybe people who use 2.7 all day don't see the warts as much as someone who has been rid of them?
I got 20 000 inserts per second from a 360 gig JSON file into Cassandra using 30 goroutines. I was maxing out at 900 in Python 3. Golang was parsing the JSON at a frankly RIDICULOUS 2 million lines per second. Python 3 was around 50 thousand.
Python 3 is adding features that other languages do better. Python smashes other languages in numerics, and approachability. Nobody needs 3 for numerics, and 3 goes backwards in approachability. If you're using 3 for its back-end tools, you're using the wrong tool.
I would never move to golang for numerics because pandas/numpy/scipy/theano stomp all over it, but its only for numerics that I'll stay with Python. Everything else I'm moving away (mostly Golang, looking at Elixir) and talking to the numeric stack over zeroMQ/msgpack.
Newsflash, Go is fast. I'm sure those numbers will be worse in Py2 than Py3 as well, but aside from that what's your point? Python is slow so... don't switch to Python 3?
> Python 3 is adding features that other languages do better. Python smashes other languages in numerics, and approachability. Nobody needs 3 for numerics, and 3 goes backwards in approachability. If you're using 3 for its back-end tools, you're using the wrong tool.
Sure, people using it for numerical stuff might like Py2, but is that the only use case of Python? Should I not switch because you do numerical stuff? Django is great for backend stuff, as is the whole ecosystem (twisted, flask, DRF, asyncio).
Python 3 is also far more approachable than Python 2. Try explaining to someone why their program exploded after typing an umlaut into it, the magical "from __future__ import division" incantation or even the silly super() syntax.
So what's left that Python does better than any other language? Numerics. That's the only place where you still have a sticky user base that is not likely to dump Python. So don't be so quick do dismiss it. It's the only leg Python is left standing on that other languages cannot easily devour.
Approachability. Not much in it. For every 3/2 = 1 you cite I give you unintuitive lazy ranges, the complete idiocy that 01 != 1, and the unnecessary print functionization which removes one of Python's greatest first-line simplicity selling points. But I will give you that Python 3 remains approachable as a first language. However on that front, inroads every day from JS, where a new coder can get "wow" results in the browser much more quickly than the boring ol' VT100 terminal.
This abstract future of Python has already been ceded to something that is not Python, just as once the abstract future of Perl was Python. Ruby once tried to take the mantle but failed. Go is now taking a lot of mindshare (from 2 or 3 more I wonder?) and I agree the Python 3 crowd just makes Go more likely to succeed, which is terrifying since it's probably the worst possible choice. My message is simply despair and accept the inevitable: Python will fade -- by all means keep using it, people still use Perl, scientist will be using it for a long time -- but the rest of the software world will pass you by. If you care about what the rest of the world is doing at all, you can help make sure the future passes to better general purpose high level languages like Nim or Clojure or OCaml, and for specific niches like scientific computing goes to the likes of Julia eventually. Anything but Go.
Yes I find Golang to be dry and uninspiring, but as per one of my other comments, I get 20x the performance, and the concurrency model beats Asyncio hands down. Honestly I needed to get 360 gigabytes of JSON into Cassandra and Python was going to take 6 days (there is some light conditionality on each datapoint preventing a raw dump). I took 1.5 of those days to learn how to do it in Golang, started it at midday today, and estimate it'll be done when I get into work tomorrow morning. Sometimes I just need to get stuff done.
EDIT: Just re-read your post and cannot help but agree that Python's demise is already telegraphed. History. Perl->Python->(?) is a nice analogy. Python won't die for sure, but it will not flourish like it did in the past decade. And yes I think it's good advice to look to more ambitious languages than Go, even if Go solves many "today" problems quite well. Personally wish Clojure would come off the JVM but am looking into Ocaml, Elixir, and when I need imperative, good ol' C wrapped up in Nim. All good candidates, though I really wish something would come up and go "fully vectorized" for the brave new world of GPUs. Personally need a REPL which is why Ocaml wins for now though Spark looks to me to be the "new R" for data science so I shouldn't exclude Scala (JVM notwithstanding). Decisions decisions...
People use Python because the language gets out of the way fairly quickly to let you get things done fairly quickly. Sure, it isn't a rigorous engineering language, but the ratio of small projects that just need something good enough and easy to work with to large/complex projects requiring careful engineering is staggeringly large. While Python is definitely cooling off; people have been working with it for long enough to want something better - unfortunately there just isn't anything on the radar that is significantly enough better to overcome Python's third party library momentum at the moment.
Believe me, I'd love a language with a really flexible/optional functional type system, better metaprogramming facilities, a better concurrency design, slightly more consistent syntax, etc - just as long as it is still easy to think in and convenient to use like Python.
As someone who, among dynamic languages, dismissed Perl quite a while back as not as useful as Python and Ruby, I'm starting to think Perl 6 may be the language that hits that spot.
I'm not a fan of asyncio, but the other stuff is quite awesome. It's not compelling enough to port some of my older Python 2.7 projects, and isn't a revolution of the language or anything, but they're good improvements. And even if they weren't good improvements, Python 3 is the future, and Python 2.7 simply won't be a viable option for a lot longer.
Your comments have been crossing into incivility. Please don't do that; it breaks the site rules and leads to tedious flamewars. I'm sure you can make your argument civilly and substantively if you try.
Now that I have made all my points, it's bye-bye.py. I won't comment on Python again.
That makes it sound like you had no control.
Being rude is not "speaking up against the received wisdom", it's just being rude, and it's against the rules here. In the future, please post civilly and substantively or not at all.
The most popular package(s) that don't work with it have been numpy, scipy, etc. When I first encountered pypy, there was a numpypy and I wouldn't be surprised if that's different now. But bottom line, tons of python code out there needs more performance and has no dependency on numpy or any C extensions.
What they need is nice JIT engines like PyPy and not to switch languages.
If no one invested in improving implementations for modern languages and switched to something lower level all the time, we would still be using Fortran for business applications.
In terms of performance naysayers only get convinced when someone proves them wrong, not because of what we think might be possible.
So it is important that we have research in GC systems, JIT compilers for dynamic languages, optimizers for FP and LP languages and so on.
Otherwise we might as well keep on using just Assembly.
PyPy also has a slower startup time/overhead so very small amounts of work may end up slower than CPython but in general larger chunks of work will be faster if not its a bug and should be reported.
There is currently some work going on to improve the C extension issue. It's currently in its beginning stages and its goals are to provide both a significant amount of compatibility with existing extensions and a significant reduction in the overhead to call out to C extensions over what is currently implemented in PyPy. On the performance side the new approach was said to remove about 40% of the existing overhead. There are tons of corner cases that still need to be worked on for the new approach to become production worthy and there is no guarantee that there will be no show stoppers that come up or that the tail of corner cases that need to be dealt with becomes a nightmare and the approach is abandoned. If this works succeeds, it will be a game changer as suddenly the majority of libraries would then become compatible with PyPy.
Personally I LOVE Python but I always feel it is the 2nd best choice. I do use Python but it rarely is the best tool to use for the problem you are solving. Maybe Pypy will turn this around, and I hope so.
For me business applications are what you would use Cobol, Clipper, Java, Delphi, C#, Eifel, ....
Of course Fortran makes sense for number crushing, but that is the language domain, not doing CRUD, ETL or distributed computing stuff (not counting MPI here).
Do you intend to refer to a product done in those areas Fortran in modern days, maybe with Fortran 2008?
Outside scientific computing and heavy number crushing algorithms and libraries, I don't see a use for it.
I'm not saying anyone should use it today lol. Julia and R are the best for that sort of stuff. Way the hell ahead of Fortran. Although I thought about making a nice web application in Fortran just to screw with people who eventually try to extend it and gasp in horror. ;)
Actually, the time-critical loops are usually implemented in C or Fortran. Much of SciPy is a thin layer of code on top of the scientific routines that are freely available at http://www.netlib.org/. Netlib is a huge repository of incredibly valuable and robust scientific algorithms written in C and Fortran. It would be silly to rewrite these algorithms and would take years to debug them. SciPy uses a variety of methods to generate “wrappers” around these algorithms so that they can be used in Python. Some wrappers were generated by hand coding them in C. The rest were generated using either SWIG or f2py. Some of the newer contributions to SciPy are either written entirely or wrapped with Cython.
A second answer is that for difficult problems, a better algorithm can make a tremendous difference in the time it takes to solve a problem. So using scipy’s built-in algorithms may be much faster than a simple algorithm coded in C.
http://www.scipy.org/scipylib/faq.html#how-can-scipy-be-fast...
Too many people just have the wrong impression, thinking that PyPy plans on re-implementing all the libraries that Numpy and scipy use but that's just completely false.
They have re-implemented the Numpy array so that it can take advantage of the JIT and so that parts of an algorithm implemented In Python that uses Numpy can also be optimized. Unlike what occurs when using Numpy under CPython where the Python code does not get optimized unless it is converted to Cython, C Code, or some alternative to Python to have it be optimized.
The thing is I've never found another developer that new about it, something is very wrong...
>>> π = 3.14
In Python 2: >>> π = 3.14
File "<stdin>", line 1
π = 3.14
^
SyntaxError: invalid syntax Python 2.7.10 (850edf14b2c7, Oct 29 2015, 17:32:05)
[PyPy 4.0.0 with GCC 5.2.0] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>>> π = 3.14
File "<stdin>", line 1
π = 3.14
^
SyntaxError: Unknown character
Are you sure you weren't trying it in Python 3 (or PyPy3) rather than PyPy? Python 2.7.8 (default, Sep 30 2014, 15:34:38) [GCC] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> π = 3.14
File "<stdin>", line 1
π = 3.14
^
SyntaxError: invalid syntax
>>>
also: Python 3.4.1 (default, May 23 2014, 17:48:28) [GCC] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> π = 3.14
>>> print(π)
3.14This doesn't match my experience of the PyPy project. I found a tiny bug in the stdlib matching against CPython, went into IRC to ask a question about test running to be sure I got it right and was quickly engaged in conversation about why I was running core tests. Next thing I know my small bug has been fixed by a core developer and my chance to contribute is gone.
I'll stick to being a user.
Interestingly, I've had the same experience when reporting and asking how to contribute to Python 3, not by a core dev but somebody "stealing" my bug and quickly submitting a patch!
The other side of the coin is that without a patch, the developers have no idea what your motives or capabilities are. They can't read minds, so if you pique their interest sufficiently, they're very likely to take it upon themselves to do the legwork. No malice intended. I should hope you'll change your mind about contributing to FOSS projects after reading this thread, because I think it's counterproductive to take this sort of thing personally (I don't know if you are, but it appears to me that you're unhappy about this at the very least).
That said, there aren't many projects that post credits or thanks to people who have discovered bugs or whose line of questioning has lead to fixing application behavior. While it's extra work and isn't always feasible, it can encourage users to pitch in if they know their name might appear in some credits (even if it's only per-release). Then there's the question of how big the problem needs to be to credit someone...
What you say about hackers getting nerd-sniped is true, but "that's just how coders are" certainly isn't. If you're a core developer for a project and you want the community to grow rather than diminish it's your responsibility to make it socially welcoming as well as having a low technical barrier to entry. I personally have spent many hours responding to emails about problems I could have fixed trivially, but if I'd done that the potential new contributor may not have grown into someone that's regularly fixing bugs I can't fix trivially. I'm not taking the unwelcoming attitude personally, I'm acting accordingly. If I can't get the information from a community I need to participate but I can get them to do the work by just throwing it over the fence I'd be crazy to continue to try and help them unless I have a very strong reason to outside of requiring fixes.
As to your point about reading minds, this is specifically mentioned in my initial comment. I asked about a problem running tests, and when I provided the context that I was asked for (I'm trying to put together a patch for this bug) there should have been no question as to motivation.
I don't think its reasonable to expect a projects core developers, when they become aware of a bug through mechanisms other than formal bug reports, to deliberately stall confirming and resolving the problem because the person raising the issue is entitled to a crack at fixing it.
One experience, with one developer, and without going through official channels to asks for participation etc.
Here's Neovim's "entry-level" label, for example:
https://github.com/neovim/neovim/issues?q=is%3Aopen+is%3Aiss...
https://github.com/rust-lang/rust/issues?q=is%3Aopen+is%3Ais...
Here are the public IRC logs for what I assume you're referring to (which is from July 2014 by the way):
https://botbot.me/freenode/pypy/2014-07-02/?msg=17362202&pag...
You found a bug in an old version of PyPy. A core developer told you you might want to try a new one. No one stole your shot out from under you.
You would undoubtedly be welcomed to contribute. Don't post FUD.
Disclaimer: I am not a PyPy core dev but I do work closely with them.
> I'll push this change quickly
> Alex_Gaynor - uhh, wait the check seems to exist?
I dunno, I kinda see both sides here. First, I don't think the devs were off-base at all here. They were friendly and looking to solve problems. However, they did seem to want to simply solve the problem, while Matthew did state that he was looking at this as a way to contribute.
The whole interaction could have been improved with a little encouragement at the end, and possibly a recommendation of an 'easy' bug that was currently in need of a fix. It was fairly clear that there was an enthusiastic new contributor, so a little effort in this direction may have been warranted.
I certainly would consider this an example of being unfriendly or unwelcoming.
The fact that I wanted to run tests and couldn't, and the team went on to why I wanted to run tests and looked to solve my problem is my complaint.
Also, to be clear, I'm not trying to say the PyPy channel is unfriendly, it certainly isn't. It's just not set up to help potential new contributors get started.
Fixed a bug the same day someone asked about it and some potential contribute blame you for not letting them fix entry level bug.
Answering people with "patch welcome" and people blame you for not fixing it.
https://github.com/kenrobbins/python-rapidjson
>>> data2 = 1.23456789e-13
>>> rapidjson.loads(rapidjson.dumps(data2))
1.23456789e-13
>>> ujson.loads(ujson.dumps(data2))
0.0Our vectorizing JIT can use SIMD semantics on all numpy looping calls. For instance, non-matrix multiply A*B or for ndarray + scalar calls
While many numpy users are in the habit of manipulating large square matrices, there is a significant number of users who use small arrays, or process RGB pixels
Seems like a huge improvement, don't know if it's the recent additions or in general pypy vs cpython but it's enough of a speed bump to sit up and take note.
import numpy as np
%timeit sum(np.sum(np.random.randint(0, 10000000, 5000)) for i in range(5000))
1 loops, best of 3: 586 ms per loopA relevant comparison is counting words or something like that. Things people don't have an easy way to do much faster.
That's very good news, since monetary contributions don't seem to be abunding (none of the goals have been met according to pypy.org). Maybe it's a marketing issue, like said before?
Thanks for all the effort made.
> The End Of Life date (EOL, sunset date) for Python 2.7 has been moved five years into the future, to 2020. This decision was made to clarify the status of Python 2.7 and relieve worries for those users who cannot yet migrate to Python 3. See also PEP 466.
> This declaration does not guarantee that bugfix releases will be made on a regular basis, but it should enable volunteers who want to contribute bugfixes for Python 2.7 and it should satisfy vendors who still have to support Python 2 for years to come.
> There will be no Python 2.8 (see PEP 404).
Take note that the initial date was set to 2015 [2] and it was delayed only last year.
[PEP 373]: http://legacy.python.org/dev/peps/pep-0373/
[2]: https://hg.python.org/peps/rev/76d43e52d978In any case the 3.3 branch seemed good when I compiled it last, it'd be nice to see a partial release.
If I were the PyPy guys I'd just announce that PyPy was now officially a fork of 2.7. This would bring all the numerical guys with them, all the corpos with big investments in 2 would fund them, and we could then let the 3.x zealots go off on their own tangent.
Anybody who likes what 3.x brings to the table should really spend a weekend learning Go, which does it all better, and faster.
I've been saying in these threads for ages that vectorization is the big win in Python (thank you PyPy), and not hobby projects like asyncio (miles better elsewhere), or (eyes rolling) type annotations.
Where are those blog posts?
According to wikipedia: https://en.wikipedia.org/wiki/Tracing_just-in-time_compilati...
Tracing just-in-time compilation is a technique used by virtual machines to optimize the execution of a program at runtime. This is done by recording a linear sequence of frequently executed operations, compiling them to native machine code and executing them. This is opposed to traditional just-in-time (JIT) compilers that work on a per-method basis.
[1]:https://en.wikipedia.org/wiki/Tracing_just-in-time_compilati...
It's a JIT that traces runtime execution, rather than looking at whole method or function bodies. In particular it eliminates control flow, which simplifies optimisation.
Basically you observe loops and loops that satisfy certain criteria are traced, that is all operations performed within one iteration are stored. Such a trace is then optimized, compiled to machine code and then executed for every iteration when the loop is encountered again.
I've read that a few times and I always come away confused -- it sounds like a huge fundamental type change needs to be made by someone well versed in the PyPy internals. i.e. not something you'd typically defer to an external contributor.
I'd like to experiment with PyPy and PyParallel, but I'm basically exclusively 64-bit Windows, so it sounds like a non-starter.