Python Concurrency: Threads, Processes, and Asyncio Explained
newvick.com
newvick.com
I reality no one writes from scratch with threads, processes or asyncio unless you are a library author.
Is that really true in the case of asyncio? From my experience, it isn't.
I am not familiar with the other three libraries you mentioned, but gevent predates asyncio and is separate from it. It does not build on top of asyncio, and is outdated because of that. It doesn't even have support for WebSockets.
Not necessarily. It means multiple computations could happen at the same time. Wikipedia has a broader definition [1]: "Concurrency refers to the ability of a system to execute multiple tasks through simultaneous execution or time-sharing (context switching), sharing resources and managing interactions."
In other words it could be at the same time or it could be context switching (quickly changing from one to another). Parallel [2] means explicitly at the same time: "Parallel computing is a type of computation in which many calculations or processes are carried out simultaneously."
0: https://wiki.python.org/moin/Concurrency
1: https://en.wikipedia.org/wiki/Concurrency_(computer_science)
While I am no fan of Python's backwards approach to multi-core programming, asyncio achieves concurrency when a coroutine reaches out to the operating system for network calls. So while two tasks in asyncio are awaiting a return from a long running network call, the operating system can be running the network calls concurrently.
Surprising to some, this is the literal definition. The word "concurrent" is a portmanteau from Latin, "con" translates as same and "current" translates as time. Concurrent literally means "same time". Comp Sci really needs to use a different word.
I don't think it's surprising to anyone that speaks English what the definition of concurrency in everyday language is. It is not the same as the computer science definition and that's unlikely to change anytime soon. Can anyone with permission fix the Python wiki?
if cpu_intensive:
'processes'
else:
if suited_for_threads:
'threads'
elif suited_for_asyncio:
'asyncio'
Interesting takeaway! For web services mine would be:1. always use asyncio
2. use threads with asyncio.to_thread[0] to convert blocking calls to asyncio with ease
3. use aiomultiprocess[1] for any cpu intensive tasks
[0]: https://docs.python.org/3/library/asyncio-task.html#asyncio....
It's unreadable in dark mode.
I’m also confused about the perf2 performance. For the threads example it starts around 70_000 reqs/sec, while the processes example runs at 3_500 reqs/sec. That’s a 20 times difference that isn’t mentioned in the text.
* Task Parallelism and Multi-threading are good for computational tasks spread across the ESP32's dual cores.
* Asynchronous Programming shines in scenarios where I/O operations are predominant.
* Hardware Parallelism via RMT can offload tasks from the CPU, enhancing overall efficiency for specific types of applications.
This important especially in non-english world. As forbidden words have different significance in other cultures.
One thing I am struggling with right now is how do I handle a function that its both I/O intensive and CPU-bound? To give more context, I am processing data which on paper is easy to parallelise. Say for 1000 lines of data, I have to execute my function f for each line, in any order. However f using the cpu a lot, but also doing up to 4 network requests.
My current approach is to divide 1000/n_cores, then launch n_cores processes and on each of them run the function f asynchronoulsy on all inputs of that process, async to handle switching on I/O. I wonder if my approach could be improved.
Where does your implementation bottleneck?
Python concurrency does suffer from being relatively new and being bolted on to a decades old language. I'd expect the state of the art of python to be much cleaner once no-Gil is hammered on for a few release cycles.
As always I suggest Core.py podcast as it has a bunch of background details[1]. There are no-Gil updates throughout the series.
[1] https://podcasts.apple.com/us/podcast/core-py/id1712665877
It will need a cycle or two to mature.
Yes. When you use N batches by the number of cores the total time is defined by the slowest batch. At the end it will be just one job running. If you make batches smaller, like 1000/n_cores/k then you may get better CPU utilization and start-to-end total time. Making k too big will add overhead. Assuming n_cores==10 then k==5 may be a good compromise. Depends on start/stop time per job.