The article does skip the most important step for getting great inference speeds: Drop Python and move fully into C++.
The article does skip the most important step for getting great inference speeds: Drop Python and move fully into C++.
It's entirely valid to trade-off either a more straight-forward design or minimizing development time for performance and just throw hardware at the problem as needed.... companies do it all of the time.
Completely agree that almost none of the SoTA github repos are really ready for production and making this stuff work can be pretty hard.
Getting this done on C++ and moving up to the next level of performance is the focus of my next article :)
too bad such great ecosystems evolved around a language that can’t fully utilize the amazing hardware we have today.
Do you have any experience with that?
All the deep learning libraries are Python wrappers around C/C++ (which then call into CUDA). If you call the C++ layers directly, you have control over the memory operations applied to your data. The biggest wins come from reducing the number of copies, reducing the number of transfers between CPU and GPU memory, and speeding up operations by moving them from the CPU to the GPU (or vice versa).
This is basically what the article does, but if you want to squeeze out all the performance, the Python layer is still an abstraction that gets in the way of directly choosing what happens to the memory.
Do you have experience how single frame processing compares between Python and C++? I see that batched processing in Python gives me a huge speed boost which hints at inefficiencies at some point but I don't know if those are related to Python, Tensorflow or CUDA itself. (Or just bad resource management that requires re-initalization of some costly things in between evaluations.)
I am curious what the basis behind the idea that Python is the performance bottleneck for inference is.
The GIL and slowness of Python become a problem when processing multiple streams or doing further time consuming calculations in Python.
All because nobody has really provided off the shelf usable deployment libraries. That Bazel stuff if you want to use the C++ API? Big nope. Way too cumbersome. You're trying to move from Python to C++ and they want you to install ... Java? WTF?
Also, some of the best neural net research out there has you run "./run_inference.sh" or some other abomination of a Jupyter notebook instead of an installable, deployable library. To their credit, good neural net engineers aren't expected to be good software engineers, but I'm just pointing out that there's a big gap between good neural nets and deployable neural nets.
To me, seeing the GIL held for 40% of time and significant time spent waiting on GIL by other threads was a fairly strong indicator. Keen to hear your thoughts/experience on it.
I know a number of python frameworks (ie. detectron) that are fast.
I'd like to see the evidence that the performance bottleneck is python, esp. when asynchronous dispatch exists.