NumPy, SciPy, TensorFlow, PyTorch, JAX, Pandas, Pillow, lxml, cjson, PyCapnP, Tornado, fast-avro, etc. all get it right. They are wrappers around C (or in some cases: Fortran/assembly/CUDA) code, where the overflow of Python method dispatch is dwarfed by the hundreds of thousands of iterations of an inner loop that's in optimized, vectorized assembly. Django, Protobufs, and Avro get it wrong (often for portability or developer velocity sake), where they wrote the whole library in Python at the expense of performance.
I was briefly tempted to write an API-compatible reimplementation of Django with the core in C++ when I left Google, but by then Django (and server-side web programming) was already falling out of favor, and if you're just shipping JSON to a SPA you can use cjson with any number of fast wsgi or asgi gateways.