We picked Python Asyncio based Tornado async for our services. As soon as we hit scale we were getting CPU bound.
Deep diving into profiling came up with JSON parsing being the culprit.
It's a painful problem to crack once you hit those limits. In some cases you could genuinely skip parsing the JSON (by returning a raw JSON containing string to client) such as when you are simply getting data from cache or db.
In other cases you simply can't skip it. For eg when you are interfacing with a 3rd party library that will only speak JSON. At that stage you are stuck.
You could try and use a wrapper around faster native JSON parser (say uJson) but it will be a trade-off between the parsing time and the time taken to copy the string to the FFI parser and copy back the results. And deal with all the complexity that that entails.
Or you could hand it off as a job to an async queue (this might be the canonical architectural approach to prevent blocking the event loop) but then you have just shifted the problem to a different place where you'll still need to throw more instances at the problem. And this adds extra latency.
I too was in the "don't optimize prematurely" camp but picking Python today for new services IMO would be taking that principle a bit too far.
Especially considering the ergonomics that modern languages like Golang or Rust offer.