It's always frustrated me that the computing world isn't as it should be. Python as a language seems to have found a sweet spot for accessibility to a lot of people, but CPython the interpreter steals all the air from PyPy, which is a vastly better implementation.
https://doc.pypy.org/en/latest/faq.html#does-pypy-have-a-gil...
>I recently tried to migrate my R code to Julia. Even though I already knew R data.table is faster than DataFrames.jl, I was totally blown away by how slow Julia is. So I quickly gave up. I think I will have to write unavoidable hard loop in cpp, which I really don't want to do...
This even goes as far as choosing not to make improvements that improve performance but make the implementation more complex (for example adding a JIT), which allows Python to continue to add language features.
And yes, I already said that packages like blas have had massive amounts of work put into them and have lasting power.
It seems we agree, after all.
It's not just being tied to BLAS and friends. Fortran lives on in academia, for example, because academics don't necessarily want to put a lot of effort into chasing wild pointers or grokking the standard template library.
For my part, I'm watching LFortran with interest because, when I run up against limitations on what I can do with numpy, I'd much rather turn to Fortran than C, C++, or Rust if at all possible, because Fortran would let me get the job done in less time, and the code would likely be quite a bit more readable.
Almost. If you're using C extensions for performance, those extensions can release the GIL, and then you get parallelism in CPython. Although if you've got a bunch of small steps, Amdahl's law is going to bite you as each step needs to re-acquire the the GIL to proceed.
Note that for Numpy, I don't know which functions/methods actually bother to release the GIL, but it's possible in theory.
It's not the performance per se, that doesn't really matter for Python's use cases, it's that the performance charateristics of your program are so erratic it doesn't feel like the pieces are fitting together. Denotationally equivalent operations having drastically diverging operational interpretations means that changing anything is a nightmare.
If you have an API call that does what you want, Python is basically C++. If you don't, you're paying a 100x penalty.
Most of the times performance doesn't matter, and when it doesn't, life is great. When it does matter, the situation feels worse than it actually is because optimizing Python is such a fight.
The frustrating part is that the language doesn't seem to "compose" any longer, it doesn't feel uniform: you can't really do things the simple way without paying a hefty peformance penalty.
With that said, the #1 reason why any code in any language is slow is "you're using the wrong algorithm", either explicitly or implicitly (wrong data structure, wrong SQL query, wrong caching policy, etc.), and in that sense Python definitely helps you because it makes it easier (or at least less tedious) to use the correct algorithm.
It's when you pull in fast libraries that composition breaks down e.g. the many-fold difference between numpy's sum method and summing a numpy array with a python for-loop and an accumulator.
Again, I'm not really a fan of Python as a general purpose language for large projects, but it really, really excels at being a kind of "universal glue", and it's why I use it again and again.
In that role, "uniformly slow" is better than "performance rollercoaster".
But it's usually being driven from a Python script that's in charge of loading the data and feeding it into spaCy's processing pipeline. That code is subject to the GIL, and, since it cannot be effectively multithreaded, it tends to become the program's bottleneck.
So what you typically end up doing instead is using multiprocessing to parallelize. But that comes with its own costs. You end up wasting a lot of computrons on inter-process communication.
Yeah, I should've said it more carefully. That's what I was hinting at with the mention of Amdahl's law. The separate bits of I/O are done in C extensions which can release the GIL, and those can be done in parallel (with each other and your computation), but if each bit of work is small, the synchronization cost (reacquiring the GIL for each packet) becomes relatively large.
Could you share a little more about your experiences with mp? I'm in the other camp - I use multiprocess in several python applications and am more or less happy with it. My real curiosity I suppose is what data you are trying to pass around that isn't a good candidate for mp.
My own use cases normally resemble something like multiple worker-style processes consuming their workloads from Queue objects, and emitting their results to another Queue object. Usually my initial process is responsible for consuming and aggregating those results, but in some cases the initial process does nothing but coordinate the worker and aggregator processes.
Regarding the data itself in this case, there are some objects that are references to in-memory structures of a C extension. Last time I checked I didn't see simple ways to share those structures via pickle. Shared memory was an option, but it required me to implement and change quite a lot of things just to get things working, not to mention that whenever the C extension gets new features I need to invest extra time in making those compatible with pickle. It wasn't worth it, specially knowing that I would still run into problems every now and then due to pickling the context.
If your heavy lifting is done in C extensions, and those extensions release the GIL before doing that work, you can still get the benefit of parallelism within CPython.
When a thread has been waiting for IO, you often want to switch to it as soon as the IO is finished in order to keep the system responsive. So the biggest issue with the GIL seems to be more about scheduling and less about actually having true concurrent threads in Python.
1. pypy has a GIL
2. not every python program is compatible with pypy.
3. not everyone can program in C.
4. it would be a free speed up for troves of multithreaded code.
that being said, the GIL is not going anywhere, there are enough workarounds and the task is monumental.