We take tens of milliseconds to query Postgres and generate a rendered HTML page for our clients. Showing this to the vendor's devs we got a very surprised response.
Admittedly, we do not operate at their scale, but I am certain this $5 a month droplet will keep running this app for a long time yet even with many users :)
Edit: I did write an MVP in Python atop Sqlalchemy wrapping Postgres, but the performance was still not ideal when rendering hundreds or thousands of rows of data, and the primary developer was already using Dotnet Core.
A more even comparable rewrite would have been FastApi with an asynchronous library for Postgres (such as SQL Alchemy or TortoiseORM).
There are probably ways to achieve similar results with Django or Flask, but it’s pretty easy with FastApi.
They were returning a large number of rows from Postgres (which, if the DB is properly set up, should take at most tens of ms: of course, depending on the width of the rows too), and most (well, I know of none that don't) Python ORM libraries (SQLAlchemy included) have a huge "serialization" cost (turning raw data from Postgres into objects). I've done a benchmark once, and things like Django-ORM or SQLAlchemy were like 10-50x slower than fetching tuples with psycopg directly. SQLAlchemy-core was fastest when fetching tuples if you wanted to not do raw SQL (IIRC, a performance penalty of at most 100%, translated to a factor, up to 2x slower), but Django's fetch-me-tuples functionality was also a single digit multiple of psycopg.
So, the solution to that problem is to fetch tuples, and then pass them in for rendering the page.
Of course, this also points at the problem with all the ORM implementations in Python: they are being too "smart" and dynamic for their own good (if all are bad at it, it also means that Python is not doing something good either, so criticism is warranted).
After I had left that specific team, I came back and swapped out the json serializer with orjson. It was like 5 lines of code if I recall. The performance skyrocketed. The GUI was noticeably far more responsive in populating the various charts and plots. By "noticeable" I mean it was loading in less than a 1/3 the previous time. Definitely recommend it. It's written in Rust, and it inspired me to start learning the language.
You can definitely outrun I/O but you can never outrun the GIL.
Python does have a huge performance penalty for basic computation, which is why it has a bunch of C-based libraries that provide bulk-operations that avoid it. If properly used (you rarely need to roll your own with a number of compiled libraries present), Python itself is not a bottleneck. One can argue if that's still Python, but at the very least, it's idiomatic Python development: I hope you don't use Python to prove that pure dynamic languages can outperform compiled languages, but to develop and deliver applications faster using Python's expressiveness.
People bring up GIL as well: it will affect your application startup time and memory usage since you can trivially avoid it by running multiple Python processes. But performance of executing code itself will only be minimally affected if you switch to multi-processing (of course, if memory pressure is so high that all those Python libraries loaded multiple times in memory is affecting your app, that can hinder performance, but that's going to be pretty rare).