Gunicorn
gunicorn.org
gunicorn.org
Many people misunderstand asyncio thinking it will make your code faster, when what it does do is make it more scalable (if you have allot of IO for example). But at the expense of developer UX. Mixing async and sync code within Python is a nightmare. I am completely unconvinced by the move to asyncio everywhere in Python.
I have been using Gunicorn with Gevent for nearly 10 years and never had a problem with it.
Languages like JavaScript that have been async from the start work well as you (mostly) only have an async api.
https://docs.gunicorn.org/en/latest/design.html#async-worker...
Uvicorn and ASGI do not. For those you need to handle async explictly.
in a general programming language sense...ur right. But from the POV of what gunicorn/fastapi are actually used for - web frameworks - its actually quite good.
I suppose also nginx docker?
ASGI is similar to WSGI but for async oriented web servers. The person you respond to is saying that they use Uvicorn via the ASGI interface.
I suppose also newlines?
gunicorn bills itself as "Gunicorn 'Green Unicorn' is a Python WSGI HTTP Server for UNIX. It's a pre-fork worker model."
uvicorn's one-liner "Uvicorn is a lightning-fast ASGI server implementation, using uvloop and httptools."
So it seems redundant at first to be running both, but what the gunicorn+uvicorn-worker-process pattern really does is throw away the "WSGI HTTP Server" part of gunicorn and just use the "pre-fork worker" process management part, with the underlying processes actually being ASGI.
If you run in a different kind of process-management environment like k8s, does running gunicorn+uvicorn give you anything else?
You have sync, gthread, gevent, evenlet, tornado (tornado was also a popular server). UvicornWorker is another type of gunicorn worker... incidentally authored by the same guys who built uvicorn.
Nothing strange at all.
But ur question is genuine - but it is orthogonal to the uvicorn part of this answer. Fundamentally, in a k8s model...why bother with gunicorn?
The answer: I don't know. I do love the control that gunicorn gives me...so I'd probably use it even if I was using a single worker.
However there is a complex answer with Threads vs cores vs pods that I'm unqualified to give.
What is your point?
If you are only using async library’s then async/await works perfectly.
I've worked on several projects recently where async was used throughout and haven't had issues. There are many async libraries now, like ones to utilize GCP services, make HTTP calls, connect to Redis, etc.
Personally, I like that. An asynchronous function is different to a synchronous one, and if you call an asynchronous function the caller is asynchronous. The "colouring" just reflects the realities of the problem domain. I'm not keen on libraries that monkeypatch the world to hide that fact, as gevent does.
However, some people prefer those libraries. Which is fine. They both exist, use what you like. There's no need to have wars about it.
This results in ecosystem fragmentation though, which means that more effort is spent developing and maintaining different kind of solutions to the same problem, only because of their different developer experience.
On the contrary, look at Go, where it was decided from the beginning (i.e. stdlib) that we're going with coroutines.
It's a bit like (I'm exaggerating) arguing for what should be the default coding style, where a built-in/blessed formatter would make this a non-issue. gofmt is a nice example of this.
Also, honest question: AFAICT Python was more of the "there's 1 preferred way to do things". Doesn't arguing between async and coroutines goes against this mantra?
That's only true syntactically. If your sync function actually does any IO, it will block your async worker thread. Said differently it just synchronized your async code.
What's wrong with sending the sync code to a thread pool? My CPU-, disk- and database-intensive application is filled with calls like:
result = await asyncio.to_thread(long_function())
...and its resource usage and performance are great.And because C# is statically typed, it can reliably warn on missing awaits (a common issue in the also-task-based javascript, though less problematic than it is in Python).
> I have been using Gunicorn with Gevent for nearly 10 years and never had a problem with it.
The point is that almost all JavaScript library’s have always been async, traditionally with callback based apis (now async/await). You never mixed async and sync code.
The only advantage of green threading is the capability to do problem-specific optimized scheduling.
If you're not doing this then you don't need green threads.
Isn’t the main advantage simply that you avoid making an absurd number of costly system calls?
The million threads risk slowing each other down with context switches, preventing any from finishing in time. An async event loop will focus on one connection at a time until `await`, and at least finish the tasks it attempts.
I haven't been in this situation, but async seems to have real advantages here. With OS threads, you don't control scheduling without costly synchronisation primitives.
You're assuming that an async event context switch is somehow vastly less costly than an OS context switch, which isn't true in the general case. (And unless you really went out of your way to make it happen then yours is the general case.)
Threads have separate stacks. Surely this costs at least a little, and adds up with the number of threads?
Maybe using a stack means using more cache lines than coroutines, but you'd have to be right at the edge of capacity for active requests for it to matter.
For async tasks, it means return, then call. For threads, it means moving to an entirely different stack, somewhere else in memory. “Loading context” is more expensive in the latter case.
> Maybe using a stack means using more cache lines than coroutines, but you'd have to be right at the edge of capacity for active requests for it to matter.
Why? Cache doesn’t only speed things up when capacity is full. Anything we have to reload from RAM will take time to load.
It only matters if you are at capacity because otherwise the active requests will be cached.
But a thread has a separate call stack. Returning and calling another function in the same stack just involves moving the stack pointer by a few bytes. Switching to another thread’s stack will invalidate much more cache, while running an async task until await will make the best use of the existing cache, and switch (much closer in RAM than in another thread) only when it makes sense, i.e. the previous task started waiting on something.
> It only matters if you are at capacity because otherwise the active requests will be cached.
CPU cache is a few megabytes. That definitely won’t hold hundreds of requests if they all transfer substantial data, even if you can handle thousands or more.
They are, and this is the entire premise green threads are built on. An OS context switch has to do a lot more including switching page tables and possibly performing a TLB flush.
(I do agree that the majority of python projects will get by just fine with regular threads though)
The kernel might allocate that much, but it's just virtual memory.
EDIT: Argh, double posting, did I hit "back" by mistake ?
This is why node rose so quickly in popularity as a server. The async IO handling could be leveraged with callbacks to scale low number of processes to handle large number of requests. Plus easy to learn and lots of people knew some JS.
JS always had an event loop (in browser so async events work) and in nodeJS it can be used for async IO. So JS has basically always handled async of sorts (although that has improved significantly).
gevent workers are best suited when your app spends a lot of time waiting on I/O (as you note), and they recommend setting `--worker-connections` to the ratio of {network IO time}/{CPU time}. If you set that too high, then it can result in higher latency for requests.
Me too, now. I've just been bitten by asyncio and it wasn't nice. Running many concurrent tasks a few of which are websockets, the state of Python websocket libraries that work with asyncio is terrible. The sockets would very silently crash until I added asyncio.sleep(0) between the more expensive CPU tasks. It's incredibly hard to debug and exceptions don't get raised properly to outside an event loop, silently making your program run in a zombie state. Even stack traces instantly become a mess when you add async. Bringing something like pandas to an environment like this is probably an immense headache.
It's too bad we got asyncio as the standard library module, though some SC concepts are going to make it into the library in future Python versions.
once properly configured :)
That is exactly the problem with asyncio, the minute you want to do something a little more cpu heavy it blocks. Although admittedly, you would probably see the same problem with gevent, and have to use the same sleep(0) trick.
Point is for good scalability you need both traditional threads and either coroutines/asyncio to work together.
Isn't that by design? Sending CPU-heavy tasks to the default executor with asyncio.to_thread() works perfectly for me.
Or maybe this doesn't scale in your use-case, because of the thread count limit? (Although in that case, I don't think you'd be better off without asyncio.)
If you use pure Python for long tasks, you're right that a ProcessPoolExecutor would be better.
Now, I'm used to this stuff from doing GUI programming on RISC OS back in the day, and got to know instinctively when I should be using WIMP_PollIdle there. For asyncio at work, I've some wrappers around range(), iter(), enumerate(), &c., that will periodically hand off control with a zero length sleep occasionally in CPU-bound code.
A pet project I would like to get started on would utilize gevent and a dynamic pool of processes to recreate an actor model environment on top of Python. That is - I hope to engineer around the GIL by distributing work to distinct processes, but at the end of the day the user just interacts with one "runtime". Trying to improve my academic knowledge of scheduling algorithms and bin-packing techniques.
I have some client projects that would greatly benefit from a move to Erlang/Elixir, but the cost/time/risk of forklifting the entire project to a new runtime is too high. I've used dramatiq.io which is advertised as an actor system, but it really isn't (actors do not have persistent state, no semantics for at-most-one, etc) I think there is a void to be filled here.
Ostensibly the underlying concurrency method is not important here, and could be stubbed out, but I will definitely get it started with gevent.
FreeRTOS is that anyway, except it's longjmping between heap- or statically-allocated stacks, and you have to unblock worker threads by poking them from ISRs. All of this is workable but I could definitely imagine it being more sane with first-class language support for the fundamental cooperative multitasking model that's at work.
Once in a while someone says that "yes, you can do logic programming in PHP or OOP in Haskell". Maybe, but syntax is more than a nudge, it's a finger on the scales.
I'd very much like a source for that, because I don't feel like it's very hard to understand.
> syntax like "async def" and "async with" makes you feel like you do, and stuff like FastAPI encourages you to type "async"
FastAPI's docs tell you to use "def" if you don't know what you're doing. [1] says "If you just don't know, use normal def."
It's also pretty simple, I think. If your function blocks, use "def", and let FastAPI run it in a thread pool to avoid blocking the event loop. If it doesn't block, use "async", and take advantage of concurrent execution without needing a new thread.
https://lucumr.pocoo.org/2016/10/30/i-dont-understand-asynci...
I haven't read the post in detail though, thanks for the link! Seems interesting.
Because that's the correct way. print("hello, world!") will block for a very short time, so it's not worth sending to a thread pool.
> In fact most, if not all, of the examples in the tutorial use "async def."
That's again because they show the correct way to use async.
Using async correctly, when it applies is preferable to a synchronous function. The docs encourage you to do it if you know how, and if your use case fits an async function.
> So I don't think it's surprising if developers new to FastAPI tend to use async by default.
It is to me, because the docs are clear:
- if you know your function can be made async properly, make it so,
- if your function blocks, use "def",
- if you don't know enough to choose, use "def".
People who complain that the event loop blocks too often simply haven't read the docs well enough.
There are trade offs for either but I think a lot of people come to python and don't understand event loops and how they work get bitten by it.
To this day I am not sure how anyone is getting better performance from using async over sync for something like Django. If you run 1000 sync gunicorn processes, I'm sure the OS scheduler is doing a better job in assigning them CPU time slices than what you could ever do with gevent.
There is a place for gevent/asyncio, for example if you want to fire off some http requests or DB calls in parallel but you'll have to do it manually no matter what gunicorn worker you run on.
I have used Gevent extensively in two places; long running eventsource/sse Django views (it’s just implicit due to the Gunicorn worker) and views with hundreds of db and http api calls that are made concurrently before returning a response. Both worked well from my experience, the latter needed coroutines to be spun up for each api/db call.
The problem with Gevent is no matter how clean the stack switching, the major liability is other code sharing that stack -- the complete universe of native Python extensions is huge, and some of it most certainly makes assumptions about the stack.
I've used Gevent/greenlet for years and never encountered issues, but it's important not to be under any illusions about the kind of hellish crashes that might be lurking just around the corner as more third party code gets added to a project.
Yep! And for those that didn't know, this strategy on HN and other paulg projects is actually what inspired Redis! See this antirez tweet https://twitter.com/antirez/status/1110468354542919681 and https://news.ycombinator.com/item?id=19498379
For the simple case, the throughput is not improved over sync workers. The scalability does improve. Basically, your code can magically cheaply wait for thousands of sockets to get ready, but it can't magically process more requests per second.
The sad thing is, this was all known because Twisted had been around for years before they made the decision to add async/await. Deferreds, waiting for deferreds, callbacks, deferred semaphores [1]. Not to mention the bifurcated state of libraries, a few supporting Twisted and most not supporting it. And then, when gevent/greenlet/gunicorn came around it was like a breath of fresh air. I remember the day I replaced all the Twisted stuff with a monkeypatch snippet + gevent with a good 20% of code reduction with all the boilerplate gone. And then, Python 3 rolled around and it all came back in the form of async/await.
By that point I moved on to a different ecosystem (BEAM VM Elixir/Erlang) so haven't followed much what was happening. But just point out, you're not the only one who noticed it. I remember commenting on it some years ago as well [2]
[1] https://twistedmatrix.com/documents/current/api/twisted.inte...
It was great at the time but ultimately tiresome. It introduced framework specific incantations all over the place and any framework using I/O had to be Twisted-aware in order to use it straight-forwardly. The exact same complaints I have today of asycnio.
I went to gevent 10 years ago and have never regretted it. If you have some reasonable knowledge of threading and synchronization then it is a really nice developer experience.
If you want to scale massively with async, do it Erlang way.
It works well, because in the end, it’s just passing a function that should be called when the process finished.
Guido listened attentively despite my rambling being essentially a poor rehashing of a proposal rejected 10 years prior! I was too starstruck to remember a concrete thing he replied with, but I remember it being kind and detailed. Was an incredibly formative experience for me about how those in a position of authority and privilege should behave.
I doubt I was "right" in any universal sense. Green threads make FFI difficult and costly, and cheap and easy FFI has always been one of Python's most compelling features.
I did switch to being a Go programmer as soon as I could find a job using it and haven't looked back. I don't think there's one right abstraction for concurrency, but I think I could live with Go's for the rest of my career and be happy.
All of this to say: +1 to gevent.
In general, I personally don't love the direction Python 3 has taken with asyncio, type hints, and more and more syntactic sugar. Meanwhile, no real improvements to packaging or to performance of pure Python code. But I'm just a bystander and Python is a big ship to steer, so I get it.
Languages like JavaScript that have been async from the start work well
This is one of several reasons I've enjoyed gradually switching over to Go from Python. Concurrency is more or less a pleasure in Go. It equals the great UX of Gevent (i.e. Goroutines) but baked into the language and multiplexed onto many real threads (versus only one thread in Gevent).
I also use Gunicorn in development too, it supports code reloading and also lets you enable Flask's debug middleware so you can use the interactive debugger. This is all configured in my Build a SAAS App with Flask course, a code example is here: https://github.com/nickjj/build-a-saas-app-with-flask.
A Django example is here: https://github.com/nickjj/docker-django-example
For anyone similarly disappointed here is an example: https://www.deviantart.com/kirawra/art/gUnicorn-Adopt-OPEN-8...
Go really simplified these decisions and basically all by using the std lib.
Erlang has been scaling massively in the telecommunications industry for a very long time now even if it's not as popular as other languages.
When selling Erlang, don't focus on "telecommunications industry" but sell based on Whatsapp: http://highscalability.com/blog/2022/1/3/designing-whatsapp....
More: https://www.erlang-solutions.com/blog/which-companies-are-us...
Makes me curious, is it that hard to clone its features into other languages? Is there something specific about Python that makes aspects of it possible there while not being possible in other languages? If so, is it a true case where dynamic languages win over static languages because certain things just can't be established satisfactorily in a statically typed language?
In hindsight I wonder if it’s just experience in the experience that has led to this differing experience or if the products are actually meaningfully different in their configuration and behavior that this is consistent for others.
Thanks for sharing, this made me feel better that I had so many issues with alternatives.
The main advantage in Python is that it's easy to learn. It just doesn't scale. Once things start taking off, there's always a migration to something more scalable like Go or Kotlin.
Gunicorn is probably the "best" tool for making a Python based website or API for Python users.
I've also found it's best to question anytime someone suggests a technology "doesn't scale."
Unless you have a very compelling reason to stick with Python (e.g., dependence on its excellent data science libraries), there are almost always better choices if your project has to coordinate many things happening at once.
I've seen folks run a $90,000 / month business through their SAAS app using gunicorn on a 2 CPU core box (4 workers).
Nothing wrong with being efficient!
Edit: somewhat typical issue with scaling python applications is that just throwing additional workers at the problem in the same instance of the wsgi container (including separate uwsgi processes) stops being effective fairly soon and makes the problem worse.
Otherwise the number of workers might limit the total throughput quite severely.
It even supports ASGI by specifying a particular Uvicorn worker class and thus can remain one go-to Python application server. This is also the reason I am using Gunicorn for the examples in my book Deployment from Scratch.
Do you need to put another server in front of it for https or has Gunicorn https handling built in?
mod_wsgi is a much less flexible setup when it comes to mixing and matching different python versions
So same number of moving parts either way.
Seems easier to just "apt install libapache2-mod-wsgi-py3" and put a "WSGIScriptAlias" directive into your apache config than to run another piece of software?
I wouldn’t leave Gunicorn “naked” for want of a better word, or use it for static files though.
> Although there are many HTTP proxies available, we strongly advise that you use Nginx. If you choose another proxy server you need to make sure that it buffers slow clients when you use default Gunicorn workers. Without this buffering Gunicorn will be easily susceptible to denial-of-service attacks.
But probably it's still a good idea to use them.
Previous: every project org-wide contained enough environment & setup information to get it running. Even if you didn't know anything else about it
Current: almost nothing
The initial "What magic incantations must I make?" step, before you're able to start really learning by experimenting, is invaluable.
https://github.com/multi-py/python-gunicorn
There's also a uvicorn one there for people who want the async features.
https://github.com/multi-py/python-uvicorn
And a container that combines both (although it's often better to just use uvicorn and add more containers to your API to scale up).
Why is that?
That's how I've heard others pronounce it too.
The disappointment....
I know people run monitoring scripts on their development systems that automatically restart Apache or whatever server they use every time they change a line of code. Still this is a cludge.
Development in PHP is:
- Change a line of code
- Alt+Tab to switch to Firefox
- F5 and see the result
So I still enjoy PHP development more when it comes to the web. Even though PHP's namespaces are crap compared to Python's module system.Change a line of code, refresh/rerun client and see the result.
- Change a line of code
- Alt+Tab to switch to Firefox
- No need to press F5 because site already updated.I have encountered very few development servers that lacked hot-reloading in decades. Many such stacks will even inject some JS or headers to automatically refresh the browser for you on detected changes.
I think OP may have an outdated or skewed experience of web-development stacks. Where PHP through apache+mod_php had "instant code refresh" back in early 2000s, and other stacks (Beans, CGI, etc) required a manual reboot everytime you'd want to check your changes in a browser. But this hasn't been true since Rails was released in 2004~5.
Ironically the most difficult one I recently encoutered was a PHP-FPM with APC, that required a "manually" restart of nginx and php-fpm in the right order after every code change. But this really is not suited as "development" server, as shown in this anecdotal issue.
Hey, now! Don't go lumping CGI in with those. :) CGI had the same "instant code refresh" as PHP since before PHP even existed. It even had a better "shared nothing" approach since you were guaranteed to start an entirely new process for every request.
The only really good hot reload system I've seen is Flutter's.
while python is gaining ground on such specialized areas (e.g things build on top of django) it is still very far behind in terms of adoption
I see this as a downside: hot-reloading is something that should be part of your developer toolbox, not a feature of the language runtime itself. Otherwise, deployments are harder because by the time you’ve completely uploaded all your updated PHP files, the server already started to execute the code in the ones uploaded so far. Also, I hope you have a caching mechanism otherwise this "read the code from disk every time you want to execute it" must be a performance nightmare.
Some of the JS kiddos from my previous org were pushing this crap on me from some random React module that I can't remember and it was terrible. We had to run a second server that proxied our personal dev webserver just so the thing could send window.reload()
Yes, there are some setups where the code reloads but the state stays there. I’ve worked like this with Clojure/ClojureScript in the past; for example if you’re working on a pop-in that appears after a click on a button, with window.reload()-based hot reloads you have to click on the button each time you reload the code, whereas with a state-preserving hot reload you don’t. It’s very useful on SPAs.
For web dev, there are countless "watchers" that auto restart on a file change. I pick any other language then PHP, on any day. Its a hot mess, and even now in PHP8.X multiple things are still broken (and will probably always be), its just a sad language to work with.