I did enjoy the work overall, incl Graviton comparison. Likewise, OS X's painfully slow Docker impl has driven me to Windows w/ WSL2 -> Ubuntu + nvidia-docker, which has been night and day, so not as surprised that those numbers are weird.
In my example, I'm using an unmodified gunicorn runner to load the uvloop worker. So I'm still only using a single worker process. Once I start tweaking the `--workers` count, I get a much higher queries per second.
And you're correct - this is a narrow benchmark, not designed to test total TPS or saturate resources.
Interesting wrt workers -- how does TCP vs sockets start looking in a multiple worker + multicore scenario, esp. wrt peak QPS? That's more confusing for schedulers :) Also, FWIW, the GIL thing might even be true in the case of single core <> single vs multiple workers. Docker supports pinning (`cpuset`?), so should be pretty doable..
We actually have a fairly similar production setup, so up to GPUs confusing things, have been curious on how deep to look in upcoming scaling work.
As a fun extra wrinkle, we also used to have nginx as a poor man's api gateway: `request -> [ nginx_container -[tcp]-> app_container1 -[tcp]-> nginx_container -[tcp]-> app_container2 -[tcp]-> ... ]`, but had enough quality issues that we removed the internal nginx indirections.