They should at least give rationale why GPU is not speeding up locality sensitive hash based algorithm. GPU's are used to compute fast hashes (they were used in Bitcoin mining once).
It's Intel sponsored research, but come on.
They should at least give rationale why GPU is not speeding up locality sensitive hash based algorithm. GPU's are used to compute fast hashes (they were used in Bitcoin mining once).
It's Intel sponsored research, but come on.
The bottleneck is the:
ptr = h(x); // trivial, ~0.5-2 cycles
bucket = *ptr; // --> ~200-300 cycles
edit: reading the paper, it's pointers to the data being stored, so I have to add the following as well: el_ptr = bucket[i]
el = *el_ptr
So that's two dependent random-access loads.... might have to...
- request to L1D cache, misses
- request to L2D cache, misses
- packet is dropped on the mesh network to access L3D, likely misses
- L3D requests load from memory from the memory controller, load is put in queue
- dram access latency ~100-150
- above chain in reverse
This is the best case scenario on miss, because there could be a DTLB miss on the address (which is why huge tables are crucial in the paper) or there could be dirty cache lines somewhere in other cores that trigger the coherency mechanism.
Obviously the memory throughput will not be as high as with matrix calculations, but the algorithm could still be optimized for the GPU. GPU's can do random access and have large and fast memories. What is the difference in memory lines? 64-bytes vs 128-bytes?
Also, SHA256(d?) employed by ethash is, actually, quite long - 80 cycles, at the very least (cycle per round). In mining you can interleave mining computation for one header with loading required by computation of mining of another header and, from what I know, this is what CUDA on GPU will do.
The sheer amount of compute power makes ethash mining faster on GPU.
On GPUs the atom size is 32/64 bytes. So GPUs are always better than or equal to CPUs when it comes to small reads/writes.
It's true that the compute power of ethash is not negligible, but to give you one more data point: on equihash there is even less compute spent on hashing, and GPUs still dominate CPUs
The thing is there's no benchmark for just neural-netting. Neural nets do well on a variety of image recognition tasks but the nets that do well as particular trained instances on particular data. And, moreover, it's not just doing well on a benchmark but generalizing well. Essentially, neural nets simply algorithms but paradigms (for good and ill).
Instead of the babble about CPU vs GPU, they should be talking the strengths and weaknesses of their approach to image recognition and related tasks. Their locality sensitive hash approach is certainly interesting and some parts of their approach might hypothetically be useful for dealing with the weaknesses of the neural net approach (fragility, black-box-unexplainability, etc).