Overhead of Python asyncio tasks
textual.textualize.io
textual.textualize.io
Any good resources to clear things up?
It doesn't necessarily improve performance. It's just a much easier way to do non-blocking I/O (note that blocking and threaded I/O is easier to do but much much heavier).
Very few locking and care is needed with asyncio, as opposed to using threads. Race conditions are basically not a thing if you write reasonably idiomatic code.
It might be (or not) faster than using threads, but that's not the main benefit in my view - this easiness of use is.
Basically, your program will normally halt when doing I/O, and won't proceed until that I/O is done. During that halt, your program is doing nothing (no cpu being used). With asyncio, you can schedule multiple tasks to run, so if one is halted doing I/O, another can run.
Edit: And AFAIK, the GIL does not come into play at all with async. Only when multithreading.
In the sense that the GIL is still held and you can have at most one path of Python code executing at a time regardless of whether you use async or threads, sure. Most blocking I/O was already releasing the GIL so the difference is purely in how you can design your modules; for any reasoning about performance the GIL behaves the same way whether you use asyncio or not.
The best mental model for me always was to think: Here's an await, that means "interpreter, go and do something else that's currently waiting while this is not done yet."
And that's all about IO, because what you can wait on is essentially IO.
By the way, I really wish there was a better story to executors in async python. To think we still have the same queue/pickle-based multiprocessing to send async tasks to another core is kind of sad. Hoping for 3.12 and beyond there.
[edit] one really neat example that helped me get asyncio was Guido van Rossum's crawler in 500 lines of python [1]. A lot of the syntax is deprecated now, but it's still a great walk-through
[1] http://aosabook.org/en/500L/a-web-crawler-with-asyncio-corou...
If you ever coded in sockets in C (and its a good exercise to do so), you probably have at some point ran across `select` which is essentially a non blocking way to check which sockets have data available to read, and then sequentially read the data. This gives the ability for a program to appear parallelized in the sense that it can handle multiple client connections, but its not truly parallel. Different clients can be handled at different time depending on which order they connect, which is asynchronous in nature (versus processing each client in sequence and waiting on each one to connect and disconnect before moving on to the next one)
Async in Python is basically this concept, with a core fundamental feature of time limited execution. Functions can say that they are pausing for x seconds, allowing other functions to run, or functions can say that they give a certain function x seconds to run before resuming execution. If you async code (along with any library you may use) doesn't contain any sleeps or timeouts, its exactly equivalent to synchronous code (since the event loop never really recieves a message that it can suspend a routine or cancel it). With sleeps and timeouts, you gain control over things that can potentially block, both from a caller perspective of not having a function call block your own, and from a callee perspective of not making your function blocking.
The use case is for it is that it is good for I/O bound operations like Threading is, but with the addition that you don't have to worry about synchronization or race conditions, since by design your code will have predictable access patterns. The downside is that your code and any libraries that you use within your code has to be implemented as async libraries, and any library that is async has to have async wrappers around the calls to its methods, which in turn means that your entire code has to be async.
Threading with Python is generally not useful, as its not true parallelism because of GIL. GIL allows only one thread in Python to run. Threading is safer in Python because of this, however it obviously has drawbacks. In general its best used if you want asyncio like performance with a library that is not written with async, since GIL is smart enough to detect when a thread is waiting for input and switch context.
True parallelism in Python is achieved with multiprocessing, however the use case is a little different. Rather than spinning off processes, you generally launch a bunch of worker processes up front (to avoid the larger overhead), then use smart scheduling to distribute work between these processes. Here though you do have to worry about race conditions and synchronization, and use things like locks and mutexes.
It's needed when you're spending a lot of time waiting for an I/O request to complete (network/HTTP requests, disk reads/writes, database reads/writes)
> what happens under the hood,
https://tenthousandmeters.com/blog/python-behind-the-scenes-...
(Please read the entire blog post - It goes through the necessary concepts like generators, event loops, & coroutines)
> does it actually help with performance,
Refer to the first answer: You'll see improved performances if your workloads are mainly comprised of waiting for other stuff to complete. If you're compute-heavy, it'll be better to use the 'multiprocessing' library instead.
> sometimes the GIL comes into play and sometimes it doesn't,
The GIL comes into play when you have a lot of compute-heavy tasks: Otherwise, you'll rarely encounter it.
It's only when you have that many compute-heavy tasks that you start to use the 'multiprocessing' library.
> why do we ever use threads at all if there's a GIL,
Threads exist because it was there before asyncio & event loops came into Python.
> why is it called asyncio if we can use it for anything.
Its name came from PEP 3156, proposing the asyncio library back in 2012.
https://peps.python.org/pep-3156/
As for why, asynchronous I/O stands in contrast to synchronous I/O, where the program/thread had to wait for the I/O request to complete before it can do anything else. Making tasks asynchronous allows it to do other stuff while it waits for a task's request to complete, increasing CPU & I/O utilization.
> My mind is kind of scattered and people seems to be using a lot of async Python for some reason. Any good resources to clear things up?
Highly recommend this video from mcoding: It's fairly simple & goes through a sample implementation.
https://www.youtube.com/watch?v=ftmdDlwMwwQ
Also, this article:
Takeaways:
1. Creating async tasks is cheap. 2. It is important to confirm intuitions, before acting on them.
Be careful to hold your references, because async tasks without active references will be garbage collected. I've been bitten by that in the past.
Long discussion here: https://bugs.python.org/issue21163
Docs: https://docs.python.org/3/library/asyncio-task.html#asyncio....
"Important
Save a reference to the result of this function, to avoid a task disappearing mid-execution. The event loop only keeps weak references to tasks. A task that isn’t referenced elsewhere may get garbage collected at any time, even before it’s done."
This prevents tasks from being garbage collected, but also prevents situations where components can create tasks that outlive their own lifetime. Plus, it's a saner approach when dealing with exception handling and cancellation.
The whole beauty of async tasks is that you can spawn, retry, and consume them lazily. When you create a task group, you again end up waiting on a single long-running last task, desperately trying to fix individual failures and retries that hold up the entire group.
I found this article: https://blog.dalibo.com/2022/09/12/monitoring-python-subproc... and while async/await syntax is the same, it's not entirely clear for me, why there's some event loop and what exactly happens, when I pass function to asyncio.run(), like here: https://github.com/pallets/click/issues/85#issuecomment-5034...
So, you can use it and it's not that hard, but there are some parts that are vague for me, no matter which language implements async support.
Edit: Updating per request below.
Tested on a TR2950x. I wonder if NUMA issues in my case or bad python version (3.9.4).
Python
100,000 tasks 77,108 tasks per/s
200,000 tasks 69,945 tasks per/s
300,000 tasks 72,453 tasks per/s
400,000 tasks 74,636 tasks per/s
500,000 tasks 66,253 tasks per/s
600,000 tasks 77,576 tasks per/s
700,000 tasks 69,673 tasks per/s
800,000 tasks 68,176 tasks per/s
900,000 tasks 73,846 tasks per/s
1,000,000 tasks 68,013 tasks per/s
.NET 6 100000 Tasks 523000 Tasks/s
200000 Tasks 550000 Tasks/s
300000 Tasks 550000 Tasks/s
400000 Tasks 559000 Tasks/s
500000 Tasks 547000 Tasks/s
600000 Tasks 539000 Tasks/s
700000 Tasks 547000 Tasks/s
800000 Tasks 540000 Tasks/s
900000 Tasks 560000 Tasks/s
1000000 Tasks 542000 Tasks/sEDIT. My results on a 5950x (undervolted)
python3.8.exe test.py
100,000 tasks 139,130 tasks per/s
200,000 tasks 121,905 tasks per/s
300,000 tasks 120,000 tasks per/s
400,000 tasks 114,286 tasks per/s
500,000 tasks 119,403 tasks per/s
600,000 tasks 117,073 tasks per/s
700,000 tasks 130,612 tasks per/s
800,000 tasks 122,488 tasks per/s
900,000 tasks 120,000 tasks per/s
1,000,000 tasks 110,155 tasks per/s
python3.11.exe .\test.py
100,000 tasks 206,452 tasks per/s
200,000 tasks 185,507 tasks per/s
300,000 tasks 186,408 tasks per/s
400,000 tasks 179,021 tasks per/s
500,000 tasks 167,539 tasks per/s
600,000 tasks 177,778 tasks per/s
700,000 tasks 188,235 tasks per/s
800,000 tasks 180,919 tasks per/s
900,000 tasks 168,421 tasks per/s
.\test.exe (go 1.20 compiled)
100000 tasks 2710563.336378 tasks per/s
200000 tasks 3076885.207567 tasks per/s
300000 tasks 3332292.917434 tasks per/s
400000 tasks 3040479.422795 tasks per/s
500000 tasks 2810232.844653 tasks per/s
600000 tasks 3004138.200371 tasks per/s
700000 tasks 2738877.029117 tasks per/s
800000 tasks 2893730.985022 tasks per/s
900000 tasks 3043877.494077 tasks per/s
1000000 tasks 2857992.089078 tasks per/s
async function time_tasks(count=100) {
async function nop_task() {
return performance.now();
}
const start = performance.now()
let tasks = Array(count).map(nop_task)
await Promise.all(tasks)
const elapsed = performance.now() - start
return elapsed / 1e3
}
for (let count = 100000; count < 1000000 + 1; count += 100000) {
const ct = await time_tasks(count)
console.log(`${count}: ${1 / (ct / count)} tasks/sec`)
}
Outputs (Python 3.11, Bun 0.5.1): % bun textual.ts
100000: 3767797.000743159 tasks/sec
200000: 9001406.4697609 tasks/sec
300000: 8281002.001242148 tasks/sec
400000: 10038491.340232708 tasks/sec
500000: 8976653.913474608 tasks/sec
600000: 10437550.828698047 tasks/sec
700000: 9443895.154523576 tasks/sec
800000: 11021991.118011119 tasks/sec
900000: 9790550.215324111 tasks/sec
1000000: 10263937.143648934 tasks/sec
% python3 textual.py
100,000 tasks 303,063 tasks per/s
200,000 tasks 270,058 tasks per/s
300,000 tasks 271,621 tasks per/s
400,000 tasks 261,945 tasks per/s
500,000 tasks 251,070 tasks per/s
600,000 tasks 272,520 tasks per/s
700,000 tasks 250,977 tasks per/s
800,000 tasks 253,131 tasks per/s
900,000 tasks 244,696 tasks per/s
1,000,000 tasks 266,061 tasks per/s $ deno run tasks.js
100000: 2777777.777777778 tasks/sec
200000: 3225806.4516129033 tasks/sec
...
800000: 2395209.580838323 tasks/sec
900000: 1679104.4776119404 tasks/sec
1000000: 1851851.8518518517 tasks/secSee: https://gist.github.com/jimmy-lt/4a3c6ad9cab1545692e5a3fe971...
$ python3.11
Synchronous
100,000 tasks 22,716,947 tasks per/s
200,000 tasks 22,706,630 tasks per/s
300,000 tasks 22,742,779 tasks per/s
400,000 tasks 22,614,202 tasks per/s
500,000 tasks 22,760,379 tasks per/s
600,000 tasks 22,799,818 tasks per/s
700,000 tasks 22,842,971 tasks per/s
800,000 tasks 22,778,395 tasks per/s
900,000 tasks 22,854,241 tasks per/s
1,000,000 tasks 22,470,395 tasks per/s
await
100,000 tasks 10,336,986 tasks per/s
200,000 tasks 10,405,286 tasks per/s
300,000 tasks 10,451,505 tasks per/s
400,000 tasks 10,482,455 tasks per/s
500,000 tasks 10,451,287 tasks per/s
600,000 tasks 10,485,478 tasks per/s
700,000 tasks 10,508,302 tasks per/s
800,000 tasks 10,505,167 tasks per/s
900,000 tasks 10,492,568 tasks per/s
1,000,000 tasks 10,457,516 tasks per/s
asyncio.create_task()
100,000 tasks 219,858 tasks per/s
200,000 tasks 196,281 tasks per/s
300,000 tasks 201,530 tasks per/s
400,000 tasks 193,674 tasks per/s
500,000 tasks 187,611 tasks per/s
600,000 tasks 201,972 tasks per/s
700,000 tasks 187,505 tasks per/s
800,000 tasks 191,531 tasks per/s
900,000 tasks 198,127 tasks per/s
1,000,000 tasks 173,259 tasks per/s
asyncio.gather()
100,000 tasks 291,095 tasks per/s
200,000 tasks 193,324 tasks per/s
300,000 tasks 129,177 tasks per/s
400,000 tasks 107,024 tasks per/s
500,000 tasks 123,023 tasks per/s
600,000 tasks 122,304 tasks per/s
700,000 tasks 121,674 tasks per/s
800,000 tasks 106,530 tasks per/s
900,000 tasks 135,841 tasks per/s
1,000,000 tasks 106,153 tasks per/s
asyncio.TaskGroup.create_task()
100,000 tasks 319,629 tasks per/s
200,000 tasks 283,560 tasks per/s
300,000 tasks 204,328 tasks per/s
400,000 tasks 203,584 tasks per/s
500,000 tasks 200,968 tasks per/s
600,000 tasks 214,506 tasks per/s
700,000 tasks 206,512 tasks per/s
800,000 tasks 204,556 tasks per/s
900,000 tasks 210,298 tasks per/s
1,000,000 tasks 202,523 tasks per/sI only ever write an asyncio daemon when I want to launch and manage the results of a bunch of Celery tasks.
Not speaking from some superior position of research here, I'm just saying what I prefer to use.
Each of those task_create calls is roughly 10,000,000,000 / 250,000 = 40,000 instructions.
Thats 40'000 instructions of pure overhead as it does not contribute to the task at hand (accidental complexity).
edit: I'm not actually sure you are counting the context switch, but I still don't think estimating instruction count that way is particularly useful.
Having an operation that can be executing 250,000 times per second on a modern processor is extremely slow... not fast.
You wouldn't be able to write your comment if the browser were written in extremely efficient assembly code because it wouldn't exist.
NASA and every big organization also has a lot of waste, but only that way you can get to the moon.
If you didn't prevent preemptive context switches during your benchmarking, it's entirely possible the only thing you measured was the context switch time.
This is a fun experiment, but to get a rigorous idea of the overhead involved takes more work than what anyone in the post or comments has done.
Reasonableness is relative and use case dependent. The post itself illustrates how the cost is insignificant compared to other "wasteful" operations related to CSS handling.
If this is too much overhead for your use case, there are plenty of other approaches and languages to choose from.
If you're not any better, just replace your whole hosted concurrency system with a statement that triggers sched_yield.
/ Recent Python convert, in spite of the horrible general performance of the official implementation of the language. That sweet, sweet module library. Also, with Docker containers the deployment issues have been solved. It might be slow to execute but it's really efficient to develop with.
100,000 tasks 177,778 tasks per/s
200,000 tasks 150,588 tasks per/s
300,000 tasks 152,381 tasks per/s
400,000 tasks 134,031 tasks per/s
500,000 tasks 160,804 tasks per/s
600,000 tasks 129,293 tasks per/s 100,000 tasks 155,257 tasks per/s
200,000 tasks 138,569 tasks per/s
300,000 tasks 134,779 tasks per/s
400,000 tasks 144,371 tasks per/s
500,000 tasks 135,672 tasks per/s
600,000 tasks 135,299 tasks per/s
700,000 tasks 146,456 tasks per/s
800,000 tasks 139,192 tasks per/s100,000 tasks 184,167 tasks per/s
200,000 tasks 160,964 tasks per/s
300,000 tasks 165,278 tasks per/s
400,000 tasks 149,577 tasks per/s
500,000 tasks 160,593 tasks per/s
600,000 tasks 168,098 tasks per/s
700,000 tasks 161,837 tasks per/s
800,000 tasks 160,364 tasks per/s
900,000 tasks 149,479 tasks per/s
1,000,000 tasks 155,919 tasks per/s
>> python3.10 create_task_overhead.py
100,000 tasks 185,694 tasks per/s
200,000 tasks 165,581 tasks per/s
300,000 tasks 170,857 tasks per/s
400,000 tasks 159,081 tasks per/s
500,000 tasks 162,640 tasks per/s
600,000 tasks 158,779 tasks per/s
700,000 tasks 161,779 tasks per/s
800,000 tasks 179,965 tasks per/s
900,000 tasks 160,913 tasks per/s
1,000,000 tasks 162,767 tasks per/s
>> python3.11 create_task_overhead.py
100,000 tasks 289,318 tasks per/s
200,000 tasks 265,293 tasks per/s
300,000 tasks 266,011 tasks per/s
400,000 tasks 259,821 tasks per/s
500,000 tasks 251,819 tasks per/s
600,000 tasks 267,441 tasks per/s
700,000 tasks 251,789 tasks per/s
800,000 tasks 254,303 tasks per/s
900,000 tasks 249,894 tasks per/s
1,000,000 tasks 266,581 tasks per/s Python 3.11.0 (heads/3.11-dirty:8d3dd5b9647, Dec 7 2022, 08:17:48) [Clang 14.0.0 (clang-1400.0.29.202)]
on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>>
[~/Documents]$ python test.py
100,000 tasks 127,992 tasks per/s
200,000 tasks 115,960 tasks per/s
300,000 tasks 117,205 tasks per/s
400,000 tasks 113,131 tasks per/s
500,000 tasks 109,609 tasks per/s
600,000 tasks 116,649 tasks per/s
700,000 tasks 110,743 tasks per/s
800,000 tasks 111,361 tasks per/s
900,000 tasks 109,688 tasks per/s
1,000,000 tasks 117,064 tasks per/sIf your Linux machine is working an order of magnitude slower than you'd expect from the hardware, that's the first thing I'd check.
How slow do you think python is?
Obviously the latter is up to the developer(s) and whatever standards they set for themselves but any software project tends towards paths of least resistance inherent in the language and frameworks being used over time as maintainers change and PRs fixes from contributers are merged. This is pedantic of course, but that doesn't change the fact that Python is not the optimal solution for this problem.
200k ops per second not fast enough for you? How many words per second do you type?
So why wouldn't you use it? I've have since the late 90s and even then it was faster than I could keep up with. Unoptimized Java Swing on SGI was the only thing that wasn't, from memory. More recently Windows Terminal had a very bad implementation for a while, but fixed it after their ass was handed to them over it here at HN.
> but it is slow to start
This is not really true either. Sure not as fast as C, but imperceptible for the most part. And they have improved it in recent versions.
Yes, you can cause it be an issue with a poorly written or though out system, this happened at one job I had. But that wasn't Python's fault they decided to pull in thousands of files each invocation.
My Python scripts respond instantly, even big ones. I have a CLI photo editor and implementation lang is not an issue that I even contemplated until now.
> and projects using it are prone towards difficult to read/messy code.
Primarily large ones with a long history of alternating developers. There are great tools to improve its scalability; use them. The simple pyflakes will eliminate most issues. Type checking gets the long tail for the mission-important+.
> It may be IO that gives AsyncIO its name, but Textual doesn't do any IO of its own.
So why on Earth are you using AsyncIO? You don't need it, if that's true...
> Those tasks are used to power message queues
How are your message queues not doing I/O? What on Earth are they doing then?
Needless to say that the whole benchmark is worthless because it never even initiates anything that would be involved when creating actual asynchronous I/O tasks...
----
I mean, I know, in Pythonland this is just your average Wednesday, but dear lord, if you don't visit that land all that often it shocks you more every time you do.
It's just using asyncio as a task scheduler, nothing more, nothing less. Maybe when you visit Pythonland you might instead be shocked in a more positive way :-)
>Needless to say that the whole benchmark is worthless
Overhead startup isn't worthless.
'I mean, I know, in Pythonland this is just your average Wednesday'
and here we go. The whole post was really just you trying to make out that you're superior to everyone else. In this case looking down on an entire ecosystem. Why? Python is one of the most popular languages. It has an elegant syntax and it's capable of solving most problems. Go jerk yourself off in private.
Dude, where's everyone, where?
I'm smarter than the guy who posted this nonsense about asyncio, that's for sure, but that's a very low bar...
Now, when it comes to Python, then, it's an environment that is flooded with the programmers with the lowest skill level imaginable. This is where all those month-long bootcamps pump their graduates, this is where all the people who took a month-long intro to CS with Python class go, this is where a lot of people who only use Python accidentally, to compliment some other programming activity (or just general computer-related activity) go. It's a swamp, and there's no reason to pretend it isn't.
On the other hand, anyone who had any serious aspirations for Python left the scene ten or so years ago. Today, Python is the worst parts of Java and PHP combined. So, again, it's a very low bar to be better than most Python programmers. Without even trying, when I searched for a job and had to do a bunch of automated assessment tests, I was consistently placed in the "best" decile, often in the "best 5%" of all applicants, and I hate the language. I'm honestly not good at it and don't want to be good at it. You don't need to try hard to be "as good" as I am, because I'm not good... but, yeah, in this particular area everyone is hands down awful.
I also had to interview about two dozens of applicants in my last job. I had people who couldn't explain the difference between expression and statement. I had people who couldn't tell what __str__() method is for. I had extremely low expectations, and yet I was consistently disappointed by the knowledge level of applicants. It's surreal what's going on there.
---
> Overhead startup isn't worthless.
Yeah, maybe... OP never measured it anyways. OP never created any asynchronous I/O tasks. OP measured how long does it take to run couple functions in Python. Admittedly, it takes a long time, but it's irrelevant to async I/O. Especially in the context of comparing whatever OP's doing to threads. Which was their goal stated in the opening statement.
Of course you were. I'd expect that for a large fraction of gainfully employed devs.
People who can't code their way out of a paper bag are way over-represented in tech screens. Because a highly competent dev will apply at 3 companies they choose and get a job. A terrible dev will do 100 applications and see what sticks.
We opened a ML intern position recently and got over 200 applications and 95% of them were terrible. Should I conclude most ML grads are incompetent? I don't think so... probably 190 of them are the bottom 5% of the local market, and these same 190 CVs are on the desk - well, in the rubbish bin - of everyone with an open position right now.
>How are your message queues not doing I/O?
We generally wouldn't consider in-process moving data around to be "I/O", now if you started interacting with an external database/file/pipe/MMAP-ed-file/etc than that would be "I/O".
>What on Earth are they doing then?
Tasks. Kind of like processes but lighter weight. Running an event loop, dispatching signals, that sort of thing. You can do all that by manually writing your own event loop but python's asyncio (IMO) makes it easier to reason about exactly when your yielding the event loop to some other task, and makes it easier to write code that can yield control of the event loop at arbitrary places. So like if you want to update a widget every 10 seconds you can write something like
while True:
await asyncio.sleep(10) #Other tasks can run during this 10 seconds sleep
#Update widget contents
Useful for stuff that needs to periodically poll data, like a process monitor, or even for just simple clock widgets.---
As an aside that cooperative multitasking can be really nice in micropython, where you can write tight loops in straight assembly if you need to and still get a pleasant task interface for managing higher level tasks/threads. Combined with some interrupt handlers it makes a pretty elegant real-time-ish operating system. (You probably need to manually deal with garbage collections though)
But, what do you think happens when you write (input) data to memory and then read (output) it from memory? I'll help you: it starts with "I" and ends with "O"!
Game over. You have no clue what you're talking about.
That's.. that's not what people mean when they talk about I/O in this context. But I think you know that, you're just grasping at straws to win an argument.
Textual is a framework for building desktop apps.
I assume the message queues are in-memory structures used to pass messages between tasks, hence no I/O.
This is really fairly standard stuff.
I understand you may not be familiar with GUI software and/or the Python ecosystem but jumping straight to condescension when you don't understand something is not really a good attitude.
Must be so frustrating, knowing that this silly language is one of the most populair programming languages in the world, despite its short comings, one of which is slower performance.
This isn't about python or any language. The extreme negative tone and sense of superiority seems to point to something else.
What is it?
But, yes, I'm very frustrated this trash is so popular. I never made this a secret :|
Python deserves negative tone, and if you think that popularity somehow makes it exempt from criticism, then you deserve the same.
You don't really criticise. You just exaggerate and emo-dump all over python and the people who use it. Your stance seems to be that python needs to be destroyed and dumped and replaced by some other language. Your stance is that people who use python are idiots.
That kind of 'criticism' is so totally useless it's kind of seriously pathological. Your 'critisism' says more about you (due to the extremely antagonising tone) and not much about Python really.
Quite sad when you think about it.