Go to try for self
Step 1 download 2.4GB of CUDA
As others have said, George Hotz is doing his best in reverse-engineering and skipping layers.
Instead people are trying to optimize install size of dependencies, which while maybe a fun hacking project...who really cares?
What I've seen is issues with the implementation of those libraries in a project.
I don't remember exactly, but I was playing with someone's wrapper for some kind of machine learning snake game and it was taking way longer than it should have on back of the napkin math.
The issue was using either a dict or a list in a hot loop and changing it to the other sped it up like 1000x.
So it's easy to think "yeah this library is optimized" but then you build something on top of it that is not obviously going to slow it down.
But, that's the Python tradeoff.
The programmer using the wrong data structure is not a problem with the language.
It's not like I had millions of items in that structure either, it was like 100. I think it contained the batch training data from each round. I tried to find the project but couldn't.
I was just shocked that there was such a huge difference between primitive data structures. In that situation, I wouldn't have guessed it would make a difference.
Go kind of cheats and has maps play double duty as sets.
Most of the time it doesn't matter because there's nothing hoy on the Python side, but if there is, then Python is going to be slowing your stuff down.