A New Chapter for PyPy
morepypy.blogspot.com
morepypy.blogspot.com
> PyPy core dev here. PyPy will always remain free and open source. No worries there. The question we and many other open source projects are trying to deal with is how to fund our developers under that constraint. We felt that it was time to explore other alternatives.
One issue that happens is when fiscal sponsors take a cut on all revenue/donations raised by the project in exchange for their services, which is 10% for SFC [1]. Fiscal sponsors also have policies around what they will reimburse, which can become as stringent and as corporate travel policies [2]. I can totally understand a project getting frustrated at their fiscal sponsor and wanting to either start their own and do it all themselves or find another sponsor.
[1] https://sfconservancy.org/projects/apply/
[2] https://sfconservancy.org/projects/policies/conservancy-trav...
When projects run these grants through parent institutions like universities, the typical admin fee is >40%. In some cases, it can be as high as 60%. Many projects are eager to enter fiscal sponsorship agreements with organisations like NumFOCUS, because so much more of their grant money goes to funding the work.
In the case of NumFOCUS, admin fees do not adequately cover the staff requirements to manage these grants. NumFOCUS "loses money" when servicing the administrative needs of these grants & it takes this responsibility onto itself solely for the betterment of the projects.
Rather than assess administrative fees similar to universities, NumFOCUS uses its other fundraising—corporate donations, event (i.e., PyData) sponsorship, individual giving—to finance its operations.
source:
- I serve on the NumFOCUS board of directors as its co-chair.
- I presented on this topic last year at the NumFOCUS Annual Summit to an audience of core developers from projects like Julia, Jupyter, Pandas, NumPy, AstroPy, &c.
- NumFOCUS budgets are public, and all of the above information can be corroborated from materials published on https://numfocus.org/
I realized a rewrite was the best course of action, but in the meanwhile the old thing had to stay up and running, and as the volume of data increased, it started to run in to HTTP timeouts as often requests took longer than 2 minutes.
I moved the thing to PyPy, and got about a 30% speedup from that. Only one lib had to be replaced with a pure python alternative, as it was using a C extension.
It bought me enough time to finish the new implementation (duplicate the data in Elasticsearch, hey presto from over a minute to about a second to get results).
For some workloads PyPy's JIT can do wonders.
I have document parsing and SPARQL queries that can take a few minutes that I'd like to run frequently so I can keep all parts of the system up to date.
I've only benchmarked it a bit, but I found I got approximately the five times speed-up that PyPy promised. This is with PyPy based on Python 3.6. I think PyPy is switching to cffi as the way to connect to C code so most native code "just works" now.
I had to backport my code from Python 3.8; Python 3.6 lacks contextvars, but there is a polyfill for that, otherwise there was no problem.
I stayed away from PyPy for a long time because it was tied to Python 3.5 which was busted in various ways. One of those was that the filesystem path objects were half-implemented, you should have been able to pass them into anything from the stdlib that expected a string path and at that time you couldn't. Little accidents like that can slow down a technology like PyPy from being adopted.
As far as I know extensions need to be written for cffi specifically.
cffi is a newer way of writing C extensions, developed by the PyPy project. It was designed to have a smaller&cleaner interface to let you call C code from Python. Here's Armin Rigo talking about it at EuroPython: https://www.youtube.com/watch?v=ejUzVcvTLgI
The CPython way of writing extensions is documented here: https://docs.python.org/3/extending/extending.html It seems to require you to deal with the internals of the CPython interpreter (deal with PyObject structs, reference counting, etc).
I know PyPy has some support for CPython extensions, but it has to emulate some internals and it's slower as a result.
blist was the one C dependency which I replaced with a pure python alternative http://www.grantjenks.com/docs/sortedcontainers/
The control plane was implemented in Python and Twisted (event driven I/O framework for the unfamiliar), which was fit for purpose at the original scale running CPython (few thousand compute nodes).
As the number of compute nodes scaled up, we developed hotspots in ser/des of control messages, which ultimately started to affect overall cluster efficiency.
Switching to PyPy gave us an immediate substantial performance boost without really having to redo any code at all (just some FFI stuff that was probably wrongly implemented in the first place).
Eventually we realized we were going to out-scale even that (at the hundreds-of-thousands of compute node level) and ended up with a Scala/Akka reimplementation, but moving to PyPy from CPython got us a lot of free breathing room.
I think he paid them a visit, seen on local tv.
LOL source in spanish: https://www.abc.es/sociedad/20130810/abci-jeff-bezos-villafr...
One of the downsides of an open society is the possibility unwanted publicity.
They issued a statement to “control the narrative”.
[1]: https://paritylicense.com/ (not affiliated, I just think it's a good and fair license)
GPL requires that any project that uses a GPL project as a part of it must also be free and open source.
It looks like Parity requires that you pay a licensing fee to the original creator of the Parity licensed project if you're using it for non-open-source reasons. So, it can be used for private/for-profit projects, but you have to pay for it in that case, whereas open source projects can use the code for free.
It might be free software or open source (i.e. in spirit), but I don't have the patience to read potentially-dubious licenses (that's what the FSF and OSI are for!)
In other words, volunteers can contribute but others get to monetize?
People like you helping people like us help ourselves - Processed World.
I wish there was more corporate giving to foundations that could handle this sort of thing but we never built that culture in software unfortunately.
I don’t think it’s fair to frame this negatively at all, really misses the nuance of these situations
What if you work on something but just before you finish someone else submits the same code.
What if their code is bug ridden but gets the bounty, that would be frustrating.
There are a few people who basically run PyPy development. They can do as they please. It's open source, so if you're so against it, you can make a "nobody profits" fork. Most outside contributions to open source projects are made by people who wanted to scratch an itch & then let the existing maintainers maintain that improvement. Their reward is the great software. This is still there so long as PyPy commits to remaining freely available
I never said nobody can make money.
The reality is that the PyPy folk (all single-digit-number of them) have fought tooth and nail to keep the project going for well over a decade. I can't begin to imagine how much highly skilled labour has been poured in by such a small concentration, all for little more than praise and repute on a handful of IT forums.
Give them a break
The tragedy is that the Python ecosystem broadly doesn't use PyPy and doesn't contribute much to it, neither code nor cash. Our compiler engineers are just as good as the folks working on CPython (and there's some overlap), but don't enjoy the powerful deep-pocketed corporate support.
For example, TensorFlow with numba preprocessing is easy, just install both packages and it'll work. TensorFlow with pypy requires a 5 hour compile and 40 GB of temporary storage. Plus some source code fiddling inside TF, if I remember correctly.
Even as an open source project, pypy should honestly consider who they're competing with for users and funding.
Besides, I don't see why they should mention any other project in a post announcing their departure from the Conservancy. The only surprising thing is no mention of the funding model they're moving to, other than a rather vague hint, "exploring options outside of the charitable realm".
If someone explores other options then announces changes they are told “should have said they were exploring options”
And I stand by my opinion that that is something that the pypy developers should consider: is this actually usable as a solution to practical problems? Or is there something else that people use instead? If so, why? Analyzing your competition is usually a good way to learn about your own strengths and weaknesses.
Well, you guessed wrong.
> it can also do general python and loop optimizations.
Yes, it can be used in general purpose workloads, with varying degrees of success. But its main purpose is made abundantly clear:
Accelerate Python Functions
Numba translates Python functions to optimized machine code at runtime using the industry-standard LLVM compiler library. Numba-compiled __numerical algorithms__ in Python can approach the speeds of C or FORTRAN.
Built for Scientific Computing
...
> ... Analyzing your competition is usually a good way to learn about your own strengths and weaknesses.
Except this is an announcement on their funding situation, so strengths and weaknesses are completely irrelevant, unless Numba has a particularly interesting funding model. (The funding model is government grants and corporate sponsorship, so, not particularly interesting.)
I mean the thing's called 'numba' lol.
I always liken Pypy to HotSpot in that to this day the numerical performance of the latter isn't spectacular and nobody really cares - it's built to handle the harder job of making vast tangled codebases of non-numerical application code run fast, not just tight math loops which are already handled perfectly well by other more specialized tools.
"All good things must come to an end". No, they don't.
> @intgr the wind-down with the SFC hasn't been smooth and this is the politically-neutral, agreed-by-both-parties post. PyPy remains the same free and open-source project. Essentially we just switched to a different money-handler. We're announcing it in the next blog post.