In the time Python can perform 1 FLOP, an A100 can chew through 9.75M FLOPS
twitter.com
twitter.com
Not sure what I'm supposed to take away from this tweet...
If anyone has a link to that please post.
[*] I like my joke, so I'll leave it be, but fwiw I missed an M ><
Python uses arbitrary precision integers, though, so it's doing more than just a single hardware add per Python "add." You can see the actual implementation here: https://github.com/python/cpython/blob/main/Objects/longobje...
You've got multiple for loops over every digit, array dereferencing, plus the actual adds. Thus, "overhead." Python takes care of a lot for you, making it possible to do whatever you want with thousand digit numbers and it'll just work, but is definitely not meant for doing lots of math with small numbers. Hence why something like PyTorch even exists.
I agree, the point isn't really about how Python is super slow compared to C++, or any language.
The point is about how easy it is for overhead to creep in when you're working with these unbelievably fast accelerators. Like, the point is that if you perform a single addition in Python, and then have your GPU crunching away at a matmul simultaneously, your GPU is going to get through 9.75 million FLOPS before Python finishes its single add.[1]
Even if you did it in C++, how many single-threaded adds could you do? Probably on the order of a couple billion (let's round (generously) up to 10 billion). In that case, in the time you finish one add, an A100 will finish 31 thousand FLOPS.
Of course, you can leverage multithreading/SIMD if it really is parallelizable. But say, if you're doing a ref-count bump, your A100 will finish 31 thousand FLOPS before you finish that ref-count bump.
I'm not meaning to badmouth python but I'm surprised it's still used as much as it is for numerical computing specifically because of these types of issues. It seems like it's only a matter of time before there's a switch to something else (Nim? Julia? Something not on most people's radar now?) or python gets a ground up overhaul.
There's a couple things that ameliorate that. For one, in many situations, it's possible to "trace" out the Python operations. The second is that, as in my blog post, you can "hide" the Python overhead by simply running asynchronously with the GPU.
But yeah, the Python overhead is annoying (see https://dev-discuss.pytorch.org/t/where-we-are-headed-and-wh...), just giving an explanation why it hasn't been a dealbreaker.