Nvidia Unveils First Mobile Supercomputer for Embedded Systems
nvidianews.nvidia.com
nvidianews.nvidia.com
Having a board which works only with a certain vendor specific API is not really that useful unless I'd specifically want to develop a full fledged product using exactly that particular board.
My only gripe is that I decided to look at OpenCL and CUDA today and felt kind of.. underwhelmed. I have a lot of OpenGL experience and I guess it's not surprising that they are so similar to framebuffers and shaders. But I was really hoping for true general purpose computing, more like a hybrid between VHDL and say MATLAB. Ideally it would work like Go, where you would send a C-style function off to an execution unit or refer to other units by id, and use something like fork/join to tabulate the results, and maybe just give up on the notion of a global memory space. Instead it looked more like image processing kernels for doing convolutions and stuff, and I mean that's great, but didn't exactly wow me.
Maybe I'm missing something fundamental? Does anyone know of a site with more of a computer science approach, say like http://golang.org, instead of so much emphasis on graphics? Or possibly a compiler that would convert Go/MATLAB/Python to OpenCL/CUDA? Am I alone in feeling a little mystified here?
Oh well, that was somewhat predictable: parallela competed with nvidia in something that nvidia cared about and could execute.
NVIDIA is SIMD, a 32 thread warp doing the same instruction in lockstep (in 4 cycles usually on 8 SPs, the 192 SP chip can run 24 warps simultaneously, though specific numbers may change a bit from version to version) and heavily penalized for in-warp divergence.
Parallela - i.e. Adapteva (http://www.adapteva.com/epiphanyiv/) is 64 independent (execution-wise) RISC cores.
But I don't know how the GPU's memory caching infrastructure works. If the bandwidth is only for serial reads, that could be a problem.
The main thing I had in mind here was hashing for bloom filters. i.e. do the different algos in parallel on a GPU, then pass those values for the main lookup to be done by the CPU.
http://0b4af6cdc2f0c5998459-c0245c5c937c5dedcca3f1764ecc9b2f...
Am I wrong?