HIP – Convert CUDA to Portable C++
github.com
github.com
* I have 1024 threads per block
* I have 48KB of shared memory per block
* I have 32 threads per warp and need make sure that
global to local memory reads are coalesced.
* Think SIMD and avoid branching as much as possible
My kernels usually follow a specific pattern: (1) Read global memory into local memory: making sure that
if thread i reads memory[n], then thread i+1
reads memory[n+1].
(2) __syncthreads().
(3) Do computation in in most thread balanced way possible.
This very specific pattern doesn't really work elsewhere. In fact, optimizing in this fashion and then porting to C++ and elsewhere loses the specific optimization. Programming in more general way loses all the things that makes the program fast. Anyway, I definitely going to look this over more.If you simply do 64-thread per "warp" (AMD's groups are per-64) and 32KB LDS aka Shared memory per block, you would be able to write portable high-performance code between AMD GPUs and NVidia GPUs.
AMD seems like its a bit behind with regards to GPGPU adoption. But AMD's hardware seems to be a good bit cheaper. You can get HBM2 models at ~$2000 from AMD (Firepro WX9100, which is Vega architecture)
Although... as they say... hardware is cheap. I'm sure most datacenters will prefer the $8000 NVidia V100 instead, because there are more people using that hardware. In particular, its easier to get started with a V100 due to AWS and other cloud-compute offerings.
I considered using it for a project recently, but ultimately decided against it because I didn't need to be able to run on AMD systems.
AMD also supports OpenCL (which I prefer to both CUDA and HIP), but it's not connected to HIP.
That’s something that’s not only useful for pure graphics.
So it is (or was?) not a straight substitute for CUDA.
(Do NOT take my word for it, I have no idea about what I am talking about)
Google won out over Oracle in the end but it took a long time — from 2012 until 2016; four years of court cases — with some courts finding that API structures were copyrightable and others found that they were not, or that reimplementation was fair use. I guess we now have a precedent thanks to this but it could still be an issue? IANAL so I don’t know.
James Gosling interview, at 57:42
https://www.youtube.com/watch?v=ZYw3X4RZv6Y
"unwilling to help us pays the bills", so nice for the Do No Evil company.
Oracle does pay for their ANSI/ISO SQL certifications.
SQL is an international standard, that one needs to pay for, it isn't available for free.
Oracle already pays for SQL certifications, unlike Google does for Java.
It's a great start, but I'm sure a lot of cases are not yet handled (like asm instructions?)
and interesting
Tensorflow uses the cuDNN library, which is closed source. There is nothing for HIP to convert.
https://instinct.radeon.com/en/6-deep-learning-projects-amd-...
Edit: OK I was reading the docs, I think I got it: hipDNN is a wrapper that (once finished) will search and replace calls (from cuDNN to hipDNN), then hipDNN itself, in turn, will call MIOpen, not sure if that's right, I would appreciate if someone who knows more could confirm
If they'd chosen one approach 5 years ago and put decent resources behind it they might be competitive by now.