HNHacker News
TopNewBestAskShowJobs

Saurabh_29

37 karma · joined September 8, 2019

submissionscomments
Saurabh_29··on Fast approximate knn using random projections
Hi, C++ implementation of approximate nearest knn which doesn't suffer from problem of curse of dimentionality whilst giving theoretical bounds on approximation.
Saurabh_29··on Waveglow Inference in CUDA C++
+1, it is the biggest bottleneck at times
Saurabh_29··on Waveglow Inference in CUDA C++
Also, in some cases like small RNN/LSTMs, CPU's can be faster.
Saurabh_29··on Waveglow Inference in CUDA C++
I will agree with you. It is a combination of both. Plus I would like to add some extra points, pytorch/tensorflow essentially use the same CUDA/Cudnn libraries, the thing they are developed with the motive to catering wide corner cases, which tend to be robust but slower at times because of the memory access patterns/algorithm selection heuristics, some extra operations. Also, we can get rid of many memory read/write operations by clubbing kernels.
Saurabh_29··on Waveglow Inference in CUDA C++
The main bottleneck is the time spent in adding bias after Conv/dense. A well-optimized code can remove that bottleneck. I haven't done it in my code as it makes it less readable.
Saurabh_29··on Waveglow Inference in CUDA C++
The main bottleneck is the data transfer speed between the GPU and the SMs. Also, using tensor core doesn't necessarily apply using half-precision as now NVIDIA supports single-precision operation in Tensorcore too.