Hacking GCN via OpenGL
onedrive.live.com
onedrive.live.com
This makes me... uneasy.
Incredible work, though. I had no idea GPU shaders were run as Von Neumann programs. I always thought they were really tiny sets of math operations with minimal branching, because it had to scale to a bunch of little cores. But that's not true entirely, somehow!
program variables ↔ computer storage cells
control statements ↔ computer test-and-jump instructions
assignment statements ↔ fetching, storing instructions
expressions ↔ memory reference and arithmetic instructions.
Maps pretty much exactly to what you need.There's also the issue the gpus are not premptable which kind of makes preempitve multi-tasking hard
Intel's latest gpu architecture has an embedded OS running on the gpu for scheduling command batches, I'm not sure what AMD and Nvidia do.
I still wouldn't write a general purpose OS for it.
It will also be sold in stand alone chips soon
Same on AMD and NVidia, except it's been like this for the past 10-15 years (depending if you count at the bottom or at the top of the hardware release pipeline).
If you're curious, lookup the Intel Broadwell GPU specs, there's sections devoted to the various levels of preemption. If you're really curious look up the workarounds needed for the finest grained preemption (this would be preempting a single GPGPU draw call).
Then decide enabling fine grained preemption should probably wait for Skylake, unless you took too much Adderall and no challenge sounds impossible. Do I speak from personal experience? I plead the fifth.
I've no experience with how fine grained nvidia's preemption is.
The advantage is that most of the silicon can go towards the actual computation rather than stuff like branch-prediction and out-of-order execution. The disadvantage is that branching and looping is problematic: when only one item wants to go down the other branch of an if-else-statement, the GPU has to run through both branches for all items (and execution is masked off on a per-item basis).
This works extremely well for graphics and high-dimensional numerics workloads (linear algebra, finite elements). It doesn't work at all for, say, spell-checking.
Is it surprising that they are Turing complete or surprising that they are using a von Neumann architecture?
You seem to be referring to Turing completeness when talking about branches. von Neumann architecture means you can execute data as code, which seems to be more what the presentation is about (?)
Given that CUDA exists, I don't see why it is really surprising that you can do advanced things with OpenGL shaders, given that they are running on the same hardware. CUDA is definitely turing complete and I think OpenCL is the same.
At the end of the day, the optimization envelope is always shifting around, and industrial computing architectures chase that incrementally, so they'll always be looping around the wheel as today's "narrow fast path" gradually becomes tomorrow's "general purpose".
[0] http://cva.stanford.edu/classes/cs99s/papers/myer-sutherland...