Python Asyncio
superfastpython.com
superfastpython.com
In fairness they were only included in asyncio as of Python 3.11, which was released a couple of weeks ago.
These were an idea originally from Trio [2] where they're called "nurseries" instead of "task groups". My view is that you're better off using Trio, or at least anyio [3] which gives a Trio-like interface to asyncio. One particularly nice thing about Trio (and anyio) is that there's no way to spawn background tasks except to use task groups i.e. there's no analogue of asyncio's create_task() function. That is good because it guarantees that no task is ever left accidentally running in the background and no exception left silently uncaught.
[1] https://docs.python.org/3/library/asyncio-task.html#task-gro...
Request passed to a thread pool. Each thread is juggling coroutines representing in flight requests
Also nobody mentions this for some reason, but asyncio doesn't make your programs faster, in fact it makes everything 100x SLOWER (we measured it multiple times, compared the same thing to the sync version), but makes your program concurrent (NOT parallel), where the tradeoff is you use more CPU, maybe less memory, and can saturate your bandwidth better. Never do anything CPU-bound in an async loop!
Really? I definitely had a huge struggle reading async python, because generators and yields break the concept of "what is a function" and struggled writing async python because of function coloring.
To me this violates a core principle: the principle of least surprise. It completely changes how I have to reason about function execution, especially if there are conditionals and different yields.
You can write web servers in a sync manner; Flask, Django is way better if you only need an API, choosing FastAPI for a simple REST API is a huge mistake.
FastAPI can be sync or async, although sync will have "a small performance penalty"
https://github.com/tiangolo/fastapi/issues/260#issuecomment-...
The only thing async io gives you is scalability of many MANY concurrent io operations. You have to either be pushing seriously large traffic or doing a lot of long running websocket/sse/long polling type requests.
The only other use case is lots of truly concurrent io within one request/response cycle. But again that is unusual, most apis have low single digit db queries that are usually dependent on one another removing any advantage of async.
Why just "within one cycle"? What about situation where you spend 1s in single db query and are 100ms cpu bound during request handling?
The point is asyncio is to have find grade control of yealding.
https://christophergs.com/python/2021/06/16/python-flask-fas...
With Python and asyncio, neither assumption holds, but the idea of shaving off some latency persists: if we rewrite everything using asyncio, our DB accesses can be parallelized, yay! This may or may not be a net win in latency; reworking your DB access patterns may gain you more.
Nope: https://techspot.zzzeek.org/2015/02/15/asynchronous-python-a...
But often the lower-hanging fruit is doing more joins in the DB, fetching fewer columns, and writing your query more thoughtfully. Async only helps if you can do something while waiting for the DB to complete your query. A common anti-pattern, exacerbated by ORMs, is doing a bunch of small queries while passing bits of data between them in Python; moving that to asyncio won't help much, if any.
Reluctant to add a "me-too" comment but I think you have it there. For whatever reason people are much keener to investigate different ways of scheduling IO than they are to investigate what the query plan looks like and why.
In Python, thread pools and async work mostly the same in practice, except that threads have a bigger overhead. So they can be used to solve the same problems.
maybe the right aproach would have been a Threadpool? Plus you don't have to refractor the task. Just let it run sync, but you can make it also async
So I've also used coroutines in Kotlin for CPU bound tasks with a nice speedup, where multiple sub-tasks get executed in parallel and then gathered when needed to produce new tasks etc., in a granular way that would be very intrusive doing with threadpools. Here I basically took the existing coroutine code and slapped some more threads on it.
With that said, I'm not entirely sold on Kotlin's way either. Both Kotlin and Python have the "what color is your function" problem.
I'm not very experienced with Kotlin, but that sounds similar to python's run_in_executor [1]
[1] https://docs.python.org/3/library/asyncio-eventloop.html#asy...
In Python there's no benefit in doing this due to the GIL, unless you're using a module which implements its own multithreading (for example in C). Python is not alone with this issue.
In other languages you're basically stepping out of the async paradigm in order to use threading in parallel. You can wait for result of the thread in the async loop without blocking.
I really enjoy using asyncio in Python for things where I have to do a lot of stuff in parallel, like executing remote scripts in a dozen of servers in parallel via AsyncSSH or for low workload servers which query databases.
In any case, what keeps me hooked on Python like a junkie is the `reload(module)` which I invoke via an inotify hook every time a file/module changes:
server.py
server_handler.py
reloader.py
server.py loads reloader.py (which sets up inotify) and that one takes care of reloading server_handler.py whenever it changes. server.py defers all the request/response handling to server_handler.py which can be edited on the fly so that the next request executes the new code. Instant hot reloading.
This also works very nice with asyncio (like aiohttp) and one ends up with very readable code.
If you need performance, then use Java, Rust, Go or C/C++, but for prototyping and tooling I absolutely love this approach with Python.
There you also have to act differently based the task
The docs on that page also point you to a different task library which does what I'd expect:
"If the work is appropriate for concurrency and parallelism, also consider using the Task Parallel Library."
I checked those docs and that is a framework that makes sense.
You can amplify this; pure Python is not meant for CPU-bound tasks. If you have a pure Python task, you've CPU optimized it somewhat, and it's still too slow, the answer is, get out of pure Python.
There's half-a-dozen options that still leave you essentially in Python land (just not pure Python anymore) like cython or NumPy, and of course dozens of other languages that leave Python entirely.
To accelerate a pure Python task to the speed that you can get out of a single thread of a more efficient language, you have to perfectly parallelize your task across upwards of 40 CPUs (conservatively!), because that's how slow pure Python is. asyncio isn't a solution to this, but neither is any sort of multiprocessing of the pure Python either.
This is not criticism. This is just engineering reality. Pure Python is a fun language, but a very slow one.
The value is allowing you to do something else while waiting for IO to finish. If you are constantly doing IO, async is a good tool. Otherwise, it won’t help.
I've used curio as an alternative concurrent programming engine for a med-large project (and many small ones). In comparison to asyncio, it's a joy to use.
The overhead to do a similar service using threadpools is much higher and the reason asyncio exists.
You wouldn't do it that way for any reason other than you can get better performance than way than by using threading or other pre-emptive multitasking.
It's not a better developer experience than writing straight forward code where threads are linear and interruptions happen transparently to the flow of the code. There are foot guns everywhere, with long running tasks that don't yield or places where there are hidden blocking actions.
It reminds me a bit of the old system 6 mac cooperative multitasking. It was fine, and significantly faster because your program would only yield when you let it do it, so critical sections coule be guaranteed to not context shift. However, you could bring the entire machine to a halt by holding down the mouse button, as eventually an event handler would get stuck waiting for mouse up.
Pre-emptive multitasking was a huge step forward -- it made things a bit slower on average, but the tail latency was greatly improved, because all the processes were guaranteed at least some slice of the machine.
"Scalability" is a better word to use when talking about asyncio. Along with describing complex concurrent programming such as for a GUI, where the added syntactical complexity if outweighed by the reduced boilerplate of traditional GUI programming.
Just this week, I got to the bottom of a performance issue in django because the developers were using async, and then doing eleventy billion db queries to import a big csv, thereby blocking all other requests.
One of the questions I ask on our programming interviews is "how would you make this (cpu bound) thing go faster" -- Async is definitely a low quality answer to that, when things like indexing or hash lookup vs array scan are still on the table.
I'd need to know the methodology. Asyncio is "gated", meaning that coroutines execute in tranches. When you queue up a bunch of coroutines from another coroutine, they don't execute right away instead they go in a list. Then when the current tranche completes the next one starts. There's some tidying up which occurs with every tranche.
As an (only one) example of "measuring the wrong thing", if you queue up a single task 1000 times which takes as its argument a unique timer instance it matters whether you queue them up one at a time as they finish running or queue them all up at once.
#!/usr/bin/python3
# (c) 2022 Fred Morris, Tacoma WA USA. Apache 2.0 license.
"""Illustrating the tranche effect in asyncio."""
from time import time
import asyncio
N = 1000
class TimerInstance(object):
def __init__(self, accumulator):
self.accumulator = accumulator
self.start = time()
def stop(self):
self.accumulator.cumulative += time() - self.start
return
class Timer(object):
def __init__(self):
self.cumulative = 0.0
return
def timer(self):
return TimerInstance(self)
async def a_task(timer):
timer.stop()
return
def main():
loop = asyncio.get_event_loop()
timing = Timer()
overall = time()
for i in range(N):
loop.run_until_complete( loop.create_task( a_task(timing.timer()) ) )
print('Sequential: {} Overall: {}'.format(timing.cumulative, time() - overall))
timing = Timer()
overall = time()
for i in range(N):
loop.create_task( a_task(timing.timer()) )
loop.stop()
loop.run_forever()
print('Tranche: {} Overall: {}'.format(timing.cumulative, time() - overall))
if __name__ == '__main__':
main()
# ./tranche-demo.py
Sequential: 0.02229022979736328 Overall: 0.04146838188171387
Tranche: 6.084041595458984 Overall: 0.012317180633544922I either use multiprocessing, or celery. using async/await looks ugly, doesn't solve most problems I would want it to, and feels unstable.
Maybe once development around it settles down I'll revisit, but I can't imagine wanting to use it in it's current form.
And yeah, async programming (in Python) isn't really for CPU bound stuff. You might benefit from multiprocessing and just use asyncio to coordinate, which is what it excels at. PyPy can really help with CPU bound stuff too, if the code is mostly pure Python.
So they are operating outside of the request without slowing it down.
We had used aync since 2014 ( twisted, tornado), and officially having async await in python starting 3.7 was the best thing happened to python.
* A Postfix TCP table.
* A milter.
* DNS request forwarding.
* Reading data from a Unix domain socket and firing off dynamic DNS updates.
* A DNS proxy for Redis.
* A netflow agent.
* A stream feeder for Redis.
https://github.com/search?q=user%3Am3047+asyncio&type=Reposi...
By the way you can't use it for disk I/O, but you can try to use it for e.g. STDOUT: https://github.com/m3047/shodohflo/blob/5a04f1df265d84e69f10...
class UniversalWriter(object):
"""Plastering over the differences between file descriptors and network sockets."""The problem that I'm aware of is at a deeper level and has to do with the ability (or lack thereof) to set nonblocking on file descriptors associated with disk files.
Almost all of the time a normal multithreaded Python server is perfectly sufficient and much easer to code.
My recommendation is only use it where it is really REALY required.
I have several python services running as glue between some of our components. They all use threading. I find it alot easier to reason about.
The first purpose is allowing more throughput at the expense of per request latency (Typically each request will take longer than with equivalent sync code).
The main scenario where an async version could potentially complete sooner than a sync version is when the the code is able to start multiple async tasks and then await then as a group. For example if your task needs to make 10 http requests, and make those requests sequentially like one would in sync code, it will be slower. If one starts all ten calls and then awaits the results, then you might be able to a speedup on this overall request.
Other main purpose is when working with a UI framework where there is a main thread, and certain operations can only occur on the main thread. Use of async/await pattern helps avoid accidentally blocking the main thread, which can kill application responsiveness. This is why the pattern is used in javascript, and was one of the headline scenarios when C# first introduced this pattern. (The alternative being other methods of asynchrony which typically include use of callbacks, which can make the code harder to develop or understand).
But basically, unless you have UI blocking problems, or are concerned about the number of requests per second you can handle, async-await patterns may be better avoided. It being even more costly in python than it is in some other languages does not really help.
We use Gevent or Asyncio in a small number of routes that either have long running requests, SSE in our case, or have a very large number of rear facing io requests we can parallelise (a few back office screens and processes).
The complexity that asyncio would add to the the code for the 95% of routs that don't need it would add a lot of unnecessary overhead to development.
My point is, I don't believe it adds any value 95% of the time, but does, sometimes significantly, 5% of the time. I think it's better to only use it where it's really needed and stick to old fashioned, none concurrent, code everywhere else.
My setup is a bunch of microservices that each run an aiohttp web server based api for calls from the browser where communications between services are done async using rabbitmq and a hand rolled pub/sub setup. Almost all calls are non-blocking, except for calls to Neo4j (sadly, they block, but Neo4j is fast, so its not really a problem.)
With an async api I like the fact that I can make very fast https replies to the browser while queing the resulting long running job and then responding back to the Vue based SPA client over a web socket connection. This gives the interface a really snappy feel.
But Complex? Oh yes.
But the upside is that it is also a very flexible architecture, and I like the code isolation that you get with microservices. Nevertheless, more than once I have thought about whether I would choose it all again knowing what I know now. Maybe a monolithic flask app would have been a lot easier if less sexy. But where's the fun in that?
How does this compare to doing the same with eg. Django Channels (or other ASGI-aware frameworks)?
I have yet to find a use case compelling enough to dive into async in Python (doesn't help that I also work in JS and Go so I just turn to them for in cases where I could maybe use asyncio). This is not to say it's useless, just that I'm still searching for a problem this is the best solution for.
I have also never used Go, but I am comfortable saying that Python async is much easier to user than JS async. I find JS to be as frustrating as it is unavoidable.
Using aiohttp as on api is not bad at all. Once you have the event loop up and running, its a lot like writing sync code. Someone else made a comment about the fact that Python has too many ways to do async because everything keeps evolving so fast. I this this is true. The first time I ever looked at async on Python it was so nasty I basically gave up and reconsided Flask but came back around later because I so despised the idea of blocking by server that I was compelled to give it another go. The next time around was a lot easier because the libraries were so much improved.
I think a lot of people think that async Python is harder than it is (now).
pull the rows from postgres that match query <x>. Process the data and push each row as an event into rabbitmq. In <200 lines of code i was easily processing 25k/rows per second and it only took me a few minutes to figure out the script.
> They are suited to non-blocking I/O with subprocesses and sockets, however, blocking I/O and CPU-bound tasks can be used in a simulated non-blocking manner using threads and processes under the covers.
If you're using it for anything besides slow async I/O, you're going to have to do some heavy lifting.
I've also found the actual asyncio implementation in CPython to be slow. Measuring purely event-loop overhead (and doing little/nothing in the async spawned tasks), it's 120x slower than JavaScript on my machine. https://twitter.com/bwasti/status/1572339846122991617
It really seems that if you're doing asyncio, you must do EVERYTHING async, it's like asyncio takes over (infects?) the entire program.
JavaScript has an extremely nice "out" (I'm sure other languages do too):
async function() {
await async_thing()
do_sync_stuff()
}
can be written function() {
async_thing().then(do_sync_stuff)
}
which is quite intuitive and helps glue code together. I think the callback-based thinking really clarifies user intent ("just tell me when it's done") without impacting performance.You could easily have, say, a thread to do all your asyncio stuff, another to do some CPU intenstive stuff (so long as it blocks the GIL) and yet another to do some blocking I/O e.g. interacting with a database with its own blocking APIs. In that case, you could spawn separate threads and exchange messages between them like usual.
Where this falls down a little is if you want asyncio to infer your whole program - or more accurately, if you want to use it in all of the task-management code at the top level of your app. In that case, you would use asyncio's thread API [1] although it's a bit clunky for when you want to maintain state bound to a specific worker thread (like SQLite database objects).
[1] https://docs.python.org/3/library/asyncio-task.html#running-...
1. You jump into a huge codebase and find a very useful function you want to use in your code (it uses asyncio)
2. Your code doesn't use asyncio, so you start converting functions to be async (or else you can't use the `await` keyword)
3. It turns out your function is called by a bunch of different users, some of which do not use the asyncio runner
4. You end up having several meetings with upstream teams asking them to use the asyncio runner
5. A single team says "no"
6. Scrap the project and just rewrite that original function you found synchronously
What you've written is true, but "infects" a smaller proportion of the code than you'd expect, in my experience. In fact, the better organised the program is, the less needs to change when converting to async. For example, reading from a connection your top-level async function would often be something like this:
async def read_and_handle_messages(connection):
while True:
next_message_bytes = await connection.get_next_message()
next_message_parsed = my_parser.parse(next_message_bytes)
my_message_handler(next_message_parsed)
All the application-specific code is in the parser and the message handler but neither of those are async. (If you need to send a response, your message handler could post a message to an async queue that is read by a separate writer task.) In your hypothetical scenario, you'd maybe write two a wrapper functions, one for blocking and one for async, which both read the message, parse it and handle it. But that only needs to be two versions of a three line function while all the subtantial code is completely shared.I find that there's very little actual async code in async programs that I write, even with no blocking IO historical baggage. Admittedly, that's partly because I organise programs to reach that goal, but it actually works out for the best because all the async-task lifecycle management ends up in one place, which makes following overall program flow particularly easy.
The problem is that the event loop is not reentrant [1], so if your sync function is being called from an async function, things do not work. Why would you ever do this you might think? Well, asyncio might be an implementation detail of whatever environment you are using. Jupyiter or ipython for example.
[1] there are... hacks to make this sort of work of course.
The python motto.
This works well if the callback is just there to notify you about a result (which is the most common case in fairness) but doesn't really work if the callback is meant to synchronously determine and return a result that is then used within the function that's calling the callback.
But it feels wrong having to spawn a thread and event loop just for that.
Modern frameworks should be orthogonal and compositional. Even orthogonality (when you use multiple frameworks, each framework solves a problem along a different "axis" and does not interact with any of the other axes) is optional.
Compositionality means I should be able to combine several systems without one affecting the other unnecessarily. async does not fulfill that requirement.
(If you really, really need to do this - triple check - use queues & worker threads to decouple the sync parts and keep them out of your event loop thread, as a sort of sync-async impedence match.)
There are bridges from async to sync and vice versa, but they must be used very sparingly, as they aren't nestable.
I find Python async to be fun and exciting and interesting and powerful.
BUT it is a big power tool and there’s so much in it that it’s hard to work out how to drive it right.
I have pretty good experience with Python and javascript.
I prefer Python to javascript when writing async code.
Specific example I spent hours trying to drive some processes via stdin/stout/stderr with javascript and it kept failing for reasons I couldn’t determine.
Switched to Python async and it just worked.
The most frustrating thing about async Python is that it has been improving greatly. That means that it’s not obvious what “the right way” is, ie using the latest techniques. This is actually a really big problem for async Python. I’m fairly competent with it, but still have to spend ages working out if I’m doing it “the right way/the latest way”.
The Python project really owes it to its users to have a short cookbook that shows the easiest, most modern recommended way to do common tasks. Somehow this cookbook must give the reader instant 100% confidence that they are reading the very latest official recommendations and thinking on simple asyncio techniques.
Without such a “latest and greatest techniques of async Python cookbook” it’s too easy to get lost in years of refinement and improvement and lower and higher level techniques.
The Python project should address this, it’s a major ease of use problem.
Ironically, pythons years of async refinement mean there’s many many many ways to get the same things done, conflicting with pythons “one right way to do it” philosophy.
It can be solved with documentation that drives people to the simplest most modern approaches.
Anything I need to do that doesn’t use this simple IO pattern, like cpu bound workers, I prefer to use processes or multiple copies of the app synchronized with a task queue.
The Python project just don’t make it obvious how to do it easy.
import asyncio
def background(f):
def wrapped(*args, **kwargs):
return asyncio.get_event_loop().run_in_executor(None, f, *args, **kwargs)
return wrapped
@background
def my_io_bound_function():...Things I've found particularly useful:
* Optional / debug logging of coroutine start/exit.
* Stats logging, including count, runtime, request queue (multiple instances of the same function) depth.
* Printing tracebacks within the coroutine context.
Haven't expended any effort on decorators, good for you.
I've found random slowdowns of a factor of 10 for no reason.
A context manager needs to have __enter__ and __exit__ methods. An async context manager needs to have __aenter__ and __aexit__ methods. The error you get if you use a context manager as an async context manager or vice versa is big and scary and offers no clue to most programmers of what the problem is.
Using async with SQLAlchemy it isn't hard to wind up exiting the event loop before all cleanup has happened. They fixed the main cause of this, but it isn't hard to still trigger it.
And so on.
And yeah, Python's documentation is useless. Never how to use stuff, only listing of everything that's possible to do / the API. Unfortunately that style is being mimicked by most other Python projects as well.
https://GitHub.com/samsquire/preemptible-thread
I am deeply interested in parallel and asychronous code. I write about it on my journal (link in my profile)
I am curious if anybody has any ideas on how you would build an interpreter that is multithreaded - with each interpreter running in its own thread and sending objects between threads is done without copying or marshalling. I I think Java does it but I am yet to ask how it does it. Maybe I'll ask Stackoverflow.
I wrote a parallel imaginary assembly interpreter that is backed by an actor framework which can send and receive messages in mailboxes.
Here's some code:
threads 25
<start>
mailbox numbers
mailbox methods
set running 1
set current_thread 0
set received_value 0
set current 1
set increment 1
:while1
while running :end
receive numbers received_value :send
receivecode methods :send :send
:send
add received_value current
addv current_thread 1
modulo current_thread 25
send numbers current_thread increment :while1
sendcode methods current_thread :print
endwhile :while1
jump :end
:print
println current
return
:end
This is 25 threads that each send integers to eachother as fast as they can. The sendcode instruction can cause the other thread to run some code. It can get up to 1.7 million requests per second without the sendcode and receivecode. With method sending it gets ~600,000 requests per secondIt reads:
> The asyncio.to_thread() function creates a ThreadPoolExecutor behind the scenes to execute blocking calls.
> As such, the asyncio.to_thread() function is only appropriate for IO-bound tasks.
It should say it's only appropriate for CPU-bound tasks.
[0]: https://superfastpython.com/python-asyncio/#How_to_Execute_a...
Async and threads only help you with I/O (where the OS affords waiting for multiple operations in parallel).
To execute CPU-bound code in parallel you need multiple processes.
I think this is a major disadvantage of Python, because processes are much costlier to spawn, and if you implement long-running workers to avoid frequent spawning, you have to incur serialization/deserialization costs, because shared memory support is very rudimentary (in essence, just fixed size numeric code).
Though I think the GIL isn't really that big a deal for most CPU-bound code, with plenty of third party libraries and even within the python standard library a lot of the code is not native python, and therefore releases the GIL lock.
I think there is often a lot of confusion around that because a lot of tutorials simply say to use threads for IO-bound tasks and processes for CPU-bound tasks, without going deeper into the differences. Threads can easily share memory which often increases complexity and requires locks to run safely, with processes you avoid locking (including the infamous Python GIL) but also have to pass around data which might hurt performance.
Regardless, it still doesn't make sense to say threads are only appropriate for IO-bound tasks, the official documentation [0] seems to lean towards preferring threads for mixing IO and CPU-bound tasks, and processes for purely CPU-bound tasks.
[0]: https://docs.python.org/3/library/concurrent.futures.html#th...
> How to Execute a Blocking I/O or CPU-bound Function in Asyncio?
But to be precise, we should differentiate between Python’s interpreter threads and OS threads. In general, OS threads are a great way to parallelize CPU-bound tasks (with the issues of locking you mention), but Python’s interpreter threads are not (because of the GIL).
Async Python is slower than "sync" Python under a realistic benchmark. A bigger worry is that async frameworks go a bit wobbly under load. https://calpaterson.com/async-python-is-not-faster.html
If we had those, all this async idiocy would go away, no more code colours, no more single-core, even better isolation and protection against context-switching tangles.
(the implementation not so nice though. Python threads on machine threads ? Whose stupid idea was that ?)
Why is only 1/3 of the screen used for content?
They are not exclusives, the best approach is to be multi-threaded while each of the threads uses concurrency to prevent blocking the cpu.
Even this isn't necessarily true, especially given you can limit a process to a single core (or, if you use a single core system, lest we forget those exist!), or, in Python's case, the GIL! You can still have parallelism with asynchronous programming, it's just not necessarily guaranteed. IIRC tokio lets you spin up runtimes on different threads which (putting the above caveat) _can_ run in parallel.
Python is simply not the right language for asynchronous programming. Things that are easy in many other languages are hard and don't work as you'd expect. If you are in a python-only shop, brush up on your presentation skills and see if you can convince them to branch out.
Or write a shell script. You're better off with "./read_socket.sh &" than you are with any python library.
(I know little about ML.)
Right now the typical paradigm is this really clunky lock-step across nodes waiting for each other (everything is blocking). If your hardware is non-homogenous (both interconnects and accelerators) or your program needs to be run in a pipelined way, you're fighting an extremely uphill battle.
Go, JavaScript etc are the usual languages of choice for this type of stuff because it's easy to express non-blocking code.
That would be closest to true multithreading.
const y = await remote_model_a(x) // different machine
const z = await remote_model_b(y) // different machine
await z.backward()
is trivially pipelined when run in parallel. With multithreading suddenly the backing C++ library has to be aware of this and figure things out for youhttps://tryexceptpass.org/article/asyncio-in-37/
One of my ongoing frustrations is python isn't just the lack of backwards compatibility but that if you inherit a script there's no way to know which version of the language it was written for. If you're lucky you can track down the developers, but IME even they often say "I'm not sure, let me run python --version".
I don't agree with this, but I will concede that async Python is rarely the correct choice.
---
#!/bin/bash
python3.8 read_one_socket.py $1
---
Use the python socket libraries if you must, but don't use python asynchronously.
I was tasked with improving a performance-critical object detection codebase for video that was written in Python. Before attempting a complete rewrite, I identified that outside of FFI calls related to PyTorch/OpenCV, some of the biggest latency overhead was due to the fact that the network I/O for fetching data to process was blocking the actual inference, so I tried to decouple the two so they could happen synchronously.
My attempt at doing this with asyncio went horrendously at best. There were no mature client libraries for the cloud services we used at the time. I was digging through PyTorch forums to figure out the relationship between the GIL and CUDA Streams. Boundless headaches.
Then I realized, I could just separate the code that loaded the data onto the machine from the inference code into two separate code bases and just have them communicate over really simple signals on a UNIX socket. Two different processes, sending a simple fixed-size string back and forth, and then I just used multiprocessing.Pool for parallelizing the downloads.