GPU folks: we need to talk about control flow
medium.com
medium.com
There are some major issues at play:
1) Shipping an OpenGL-based product on multiple vendors' devices (especially on mobile) requires a huge amount of QA effort and per-device bug workarounds. Pretty much everyone makes their games on top of Unity/Unreal/etc because the engine vendors do most of the QA work.
2) Issue #1 is made worse by the fact that many consumer devices do not get GPU driver updates because the phone/tablet OEM end-of-lifes their products too soon and/or the mobile operators and other middlemen don't ship updates as they should. This means that if you ship a GPU-using product, you need to support devices with old, buggy drivers and this is expensive.
3) There's a huge untapped potential in GPUs, there are very few non-graphics apps taking advantage of the processing power but I can't see it changing for the better before GPUs become easier to target and verify correctness of operation.
It would be too easy to blame GPU vendors' software engineering practices, but I (as a GPU driver programmer) see that the bigger problem is issue #2 (not that the GPU drivers are faultless). Even if the drivers do get fixed and updated, getting the updates to the hands of the customers is still going to be an issue. The middlemen need to be cut out of the equation, we can't be dependent on the business requirements of OEMs and operators when shipping mission critical software infrastructure.
This is a bit of a chicken and egg problem, games are not important enough to OEMs and operators fixing their update delivery mechanisms, but no-one dares to use GPUs for anything more important before this issue gets sorted out.
Most of the vendors know they have serious issues here, but hide them. Literally. You never notice because they don't use open source compilers (even for things like CUDA, let alone shader compilers), so it's not obviously more than "just a bug" until it happens to you continuously. Instead, they pretty much never have to fix the bug until someone notices, and then they hack it some more and move on, instead of fixing underlying issues in their structurization/etc passes.
Most vendors i've talked to can't even tell me what control flow breaks their compiler (again, doesn't matter if we are talking shaders, cuda, you name it. it's all broke), they know plenty does, but are fairly ¯\_(ツ)_/¯ about doing more than working around whatever bug they get given.
Meanwhile, over in open source clang/llvm world, we can basically fuzz test CFG's, etc for CUDA.
The death of some of these compilers can't come fast enough.
Actually AMD recently completely open sourced their entire GPGPU compute stack. See https://github.com/RadeonOpenCompute
Compiler, assembler (LLVM based), linker, driver, etc. Everything is open to my knowledge.
and yes, AMD has done this because they have nothing to lose anymore :)
It is however my general impression that the graphics parts of GPU drivers are generally more buggy than the compute parts. Likely because the graphics stack is older and more complicated. Is this a correct assessment?
I'm not a graphics driver developer, but I know that game engines for modern games often are actually quite buggy and often the graphics driver stack tries to work around engine bugs. It's also well-known that for AAA titles the graphics driver internally even often replaces the game's shaders completely by hand-optimized ones made by the GPU vendor. I have read from an insider that in the last years NVIdia's graphics drivers often did not validate arguments (for a little bit more performance), while the AMD's did. This was as I heard the reason why one of the last Tomb Raider game had initially problems on AMD GPUs.
Also for OpenGL drivers you have to realize that in the past Intel drivers did not really support modern OpenGL versions. So at least for Indie developers the safest way was to target OpenGL 2.1 + extensions. Though OpenGL >= 3.1 should better be initialized using a Core context many developers still use(d) a Compatibility context - this adds a lot of legacy stuff that one has to stay compatible with and thus makes the whole OpenGL implementation of Compatibility contexts very error-prone (in the driver). Luckily we now have Vulkan (though writing your engine against Vulkan is at least at the beginning more complicated (time to first triangle) than targeting OpenGL or DirectX 11.x (not DirectX 12 - DX 12 is a lot more like Vulkan)).
In the closed-source and hardware worlds, it's a lot more common to implement a work-around wherever you are. (Alternatively, the entity paying money pressures the entity being paid to put in a workaround on their side.) Soon there are a bunch of workarounds in common use which in turn need to be worked-around. This is the world of commercial games and graphics drivers.
I used to work on a WiFi access point product, though I only touched the actual driver a few times (other team members focused there), and there was a lot of working around bugs, and working around work-arounds in other WiFi devices for other bugs in WiFi APs, in that world too. Some crazy stuff ...
The "experience" thing is optional afaik.
Also, as a consumer the "experience" seems to matter a lot more to me than ... whatever the other side of this argument is ... the hardship for programmers?
And since you're bringing up security: if there are security holes in GPU driver installations, I'd bet eliminating all the crapware would do a lot more for security than any number of updates could ever do.
Apple's OpenGL drivers (on desktop at least) are the worst in the industry, and that is not an exaggeration. I have found so many cases in which you can make shaders read random memory ([1] for example) that I'm no longer surprised.
[1]: https://github.com/servo/webrender/commit/96ac2f13b9b73dcfa0...
Getting rid of 3rd party installers and updaters should be a priority of Microsoft's anyway.
> if there are security holes in GPU driver installations, I'd bet eliminating all the crapware would do a lot more for security than any number of updates could ever do.
The "crapware" is just some unprivileged user space processes. The driver has kernel mode components, and security issues there could cause crashes, snooping memory of other processes or even kernel space remote code execution.
Not that I'm defending bundled (or optional) crapware, but I don't think that's a security problem.
A well written C code should run about 20-50 times slower than a well written CUDA code on the latest GPUs, and that's for a single CPU core. This was true with my hand-written implementations, as well as when using libraries (Numpy vs CuDNN).
Which is still pretty major, but nowhere near the wild claims that used to be common in academic papers. It's outdated, but [1] is still pretty good reading.
[1] http://sbel.wisc.edu/Courses/ME964/Literature/LeeDebunkGPU20...
Two conv layers (6 and 12 feature maps, 5x5 filters), one fully connected layer (120 neurons), activation=Tanh, pooltype: average (excluding padding), cost=Negative Log Likelihood.
Learning rate=0.20, minibatch size=100, dropout = 0.0, L2 lambda=0.0, momentum=0.0, initialization: normal
Training for 10 epochs:
CPU: 172 sec, GPU: 14 sec.
When decreasing batch size to 20 images, the numbers are: CPU: 318 sec, GPU: 38 sec.
So yeah, the CPU code your professors wrote was really crappy.
I'd be betting that someone with low level optimization expertise (comparable to what you appear to have done with the GPU) could get at least another 10x out of the CPU version. You are completely right that GPU's have the potential for great increases over the current normal, but there's also (typically) room for large improvement on modern CPU's as well.
Lets see it then. Most cpu programs have enormous amounts of fat left in them from cache misses and poor or no use of SIMD.
Just like them OpenGL is a text standard describing how a 3D API is supposed to behave.
It happens that what those papers state and what each team of developers at every card manufacturer understand is not always the same.
Then there is the set of card specific extensions, for every nice feature they want to sell on their graphics card, but yet to be adopted by OpenGL paper standard.
There are of course certification tests available, but they are costly and don't cover 100% of the API usage anyway.
So programming OpenGL, happens to be like trying to write web applications, while trying to make the code portable and bug free (with workarounds) across all graphic cards out there.
And it's much harder to test because you have to find the outdated graphic card, with the outdated driver, and have a machine where you can plug the offending card.
Ideally games would ship only shaders that are strictly valid GLSL but that's not quite the case.
If you're developing GL, do everyone a favor and start using the official GLSL validator in your build scripts.
https://github.com/mrdoob/three.js/pull/7556
Also some other code that seemed completely valid failed on Nexus 5 devices - so we had to revert it:
https://github.com/mrdoob/three.js/pull/9948
Basically mobile GPUs are a mixed bag of bugs all over the place. Three.JS is an attempt to find common path through those bugs in order to achieve reproducible results.
If you want to try your hand at this, there is still this outstanding issue that we haven't yet tracked down on the Nexus 5 devices:
I sincerely hope that you test our Linux drivers (I work for Intel on them). We'd be excited to learn about anything you find. Please feel free to file bugs here (https://bugs.freedesktop.org/enter_bug.cgi?product=Mesa&comp...) and ping us on #intel-gfx on Freenode.
I sincerely hope that they test out both Linux and Windows drivers. It is well-known that these drivers at least in the past were developed by two completely different teams, so the results might be quite different and thus interesting.
https://github.com/mrdoob/three.js/issues/10331
Linux Intel didn't reproduce, but Windows Intel did. Windows Intel acted like Adreno GPUs but Linux Intel behaved like NVIDIA GPUs. Fun times.
https://dolphin-emu.org/blog/2013/09/26/dolphin-emulator-and...
The vendors are Nvidia, AMD, Intel/Linux and Intel/Windows respectively. Yes, Intel maintains two completely separate OpenGL implementations.
It is pretty easy to capture the image from a canvas. The rendered garbage could reveal buffered adjacent/random memory (requires research). The memory can be extracted from the image with the right interpretation.
https://github.com/KhronosGroup/WebGL
There are already something like 2500 tests
AMD and nVidia seem to have this better worked out
AMD, for example: https://medium.com/@afd_icl/first-stop-amd-bluescreen-via-we...
They are doing the suppliers alphabetically, so they'll get to NVIDIA in good time :)
What mades you think that driver stability was a mobile problem? :D
Not an exclusive problem, but it seems to be worse on those platforms :)
As far as these go they aren't that bad, I've seen a highly regarded vendor's shader compiler die and bring down the whole Android stack which resulted in insta-reboot.
If it isn't in the rendering path for Android(HWUI) or Chrome then it's best to tread carefully on mobile GPUs.
This is true, but it eleminates at least one one the points where things can go wrong. Additionally this approach has the advantage that developers (with some practice) can read the SPIR-V "assembly" code to make sure it is correct. With existing solutions it was already hard to get and interprete the intermediate code to find out whether the problem is in the frontend or backend.
From the title, I thought this would be a post about warps/wavefront divergence :)
Bugs are annoying, but so are mismatched programmer expectations.
They are SIMD vector processors (Same Instructions Multiple Data).
CUDA and OpenCL make this a bit more explicit.
Wrong. Let's look at the instriction set of the AMD GCN3 ISA:
> http://gpuopen.com/compute-product/amd-gcn3-isa-architecture...
which links to
> http://32ipi028l5q82yhj72224m8j.wpengine.netdna-cdn.com/wp-c...
which even has a whole chapter about control flow: "Chapter 4 Program Control Flow".
I'd say it's more like "not quite". ~10 years ago, the statement was pretty much correct. GPUs were entirely SIMD, so they couldn't truly branch, but could fake it with predication. Longer branches would involve executing both branches on every thread and only committing the results of the active branch.
Modern GPUs can do much better, since individual warps/wavefronts can truly diverge. Within warps, however, it's still a bit of a mess.
If one looks at the capabilities of CPUs from 10 years ago, one can easily justify similar completely outdated statements for CPUs.
GPUs can implement control flow by calculating both results and then throwing away one of them based on a predicate. Also, GPUs typically do have a scalar unit which can perform branching but this is heavily discourage by CUDA and OpenCL due to the performance implications.