> (is decoding JSON really that intensive?)
in Python, everything is generally CPU intensive compared to what it would be in compiled languages, even though things like JSON decoding are usually happening in a C library, Python programs that do close to nothing still use way more CPU than you would if you were running in the JVM, or Go, C, whatever.
> Wouldn't this affect any language/framework that uses a cooperative concurrency model, including node.js and ASP.NET or even Python's async/await based frameworks? How is this problem specific to Python/Gunicorn/Gevent?
CPU bound-ness affects all of these platforms, yes. It affects Python and other intepreted languages the most however because these platforms get the most CPU-bound the most quickly. Also, applications that are written in scripting languages tend to have a lot of business logic going on in the first place; after all, if you just wanted to serve static pages you could use Apache with the Event NPM and if you wanted to proxy HTTP requests you'd use HAProxy; both event-based systems that are very much not CPU bound.
But yes, most importantly, Python's asyncio system is completely impacted by these same issues and I would have preferred she address that, as asyncio is part of the standard library now and is way more popular than gevent.
> What would be a better alternative? The author says something about using actual OS-level threads but I thought the whole point of green threads was that they are cheaper than thread switching?
I will grant she lost me a bit with the "use a real RPC system with <feature> <feature> <feature>" thing, and additionally the "load the application in the child process" thing is pretty typical, a worker process should obviously have either threads or greenthreads in use so that each process can handle multiple concurrent requests, but only as many as you'd want handled effectively by one core since the GIL is going to enforce that (another thing you wouldn't have to deal with in other languages such as the above mentioned compiled languages), but it's typical that child processes are going to have a mostly original copy of things.
But as far as the "context switching" thing, I've yet to see benchmarks that show the overhead of OS-level context switching actually being more of a performance burden than the less frequent, but more work intensive context switching that user-space schemes like asyncio have to use. If you are writing a logic-heavy, or even a logic-just-a-bit service that receives Python requests you will also have to worry about CPU-bound issues all the time. Using regular threads with processes, like what you get using something like mod_wsgi, will allow individual processes to attend to web requests more evenly. With mod_wsgi you can configure worker daemons that run multiple OS level threads and you can also have multiple daemon processes.
I'm not sure if the multi-process model used by mod_wsgi has solved the accept() problem, however in my experience the bigger problem is when a service configures itself to allow for 1000 greenlets within each process, while each process is realistically capable from a CPU perspective of handling maybe 5 or 10 concurrent requests, there's no mechanism that ensures that each process gets an even balance of requests. That is, you might have all your requests waiting in one process, because you told them it can process 1000 at a time, while other processes are idle.
TL;DR I'm in the "event based programming is extremely overrated in Python" camp.