DirectX 11 vs. DirectX 12 oversimplified
littletinyfrogs.com
littletinyfrogs.com
"Creating dozens of light sources simultaneously on screen at once is basically not doable unless you have Mantle or DirectX 12. Guess how many light sources most engines support right now? 20? 10? Try 4. Four. Which is fine for a relatively static scene. "
For my Masters degree project at uni I had a demo written in OpenGL with over 500 dynamic lights, running at 60fps on a GTX580. Without Mantle, or DX12. How? Deffered rendering, that's how. You could probably add a couple thousand and it would be fine too.
"Every time I hear someone say “but X allows you to get close to the hardware” I want to shake them. None of this has to do with getting close to the hardware. It’s all about the cores"
Also not true. I work with console devkits every single day and the reason why we can squeeze so much performance out of relatively low-end hardware is that we get to make calls which you can't make on PC. A DirectX call to switch a texture takes a few thousand clock cycles. A low-level hardware call available on Playstation Platform will do the same texture switch in few dozen instruction calls. The numbers are against DirectX, and that's why Microsoft is slowly letting devs access the GPU on the Xbox One without the DirectX overhead.
DX12 does offer some potential CPU side performance benefits when it comes to updating large numbers of constants efficiently which may well help performance when dealing with lots of dynamic lights but it's not adding any new capabilities beyond what DX11 offers.
I don’t know what kind of 3D engine you wrote, but as a FPS gamer I have quite a bit of experience with 3D engines. From Doom 3 to Alan Wake, some of the worst performance hits occur in scenes with heavy use of dynamic lights. Did OpenGL 4/DX11 fix this?
Which of the modern APIs have you actually used? How do Mantle, DX12, Metal, and OpenGl Next compare?
EDIT: "as a non-3D graphics programmer, who's tried to build FPS levels for fun and familiar with all the modern engines". This article is for those of us with interest but not experts in the field, right?
There is a quite a good explanation on how deferred rendering works: http://gamedevelopment.tutsplus.com/articles/forward-renderi...
Someone figured out deferred rendering which was a new technique that allows lots of lights. Lots of games use it. I believe one of the first was Killzone
http://www.slideshare.net/guerrillagames/the-rendering-techn...
Here's a live demo of deferred rendering
http://threejs.org/examples/webgldeferred_pointlights.html
It's using only OpenGL 2.1 features (which is all that's needed to emulate OpenGL ES 2.0 which WebGL is based on). To do deferred rendering efficiently all you really need is support for multiple render targets.
It's like somebody claiming they can comment meaningfully on light bulb manufacturing standards because they've seen the lighting in a bunch of made-for-TV specials.
For stencil shaders threads might help, since the CPU has to calculate and upload them frequently (unless you're doing stencils in a geometry shader). Stencil shadows are pretty niche though, you only use them when you need pixel-perfect precision. Shadow maps are vastly more popular.
His entire point is that you CAN do this, but it all adds up to making it not look real.
He's setting up this notion that the problem with graphics is a lack of threading, which is ridiculous. Graphics code is incredibly parallel, and his assertion that graphics work spends most of it's time constrained by waiting for the CPU just does not add up to me. Except in poorly written systems, I just haven't seen this be the case often at all.
> For my Masters degree project at uni I had a demo written in OpenGL with over 500 dynamic lights, running at 60fps on a GTX580. Without Mantle, or DX12. How? Deffered rendering, that's how.
Indeed. For fixed function forward rendering you do have those limitations, though. However it's nowhere as low as 4: you can expect at least 8 and most often 16 lights. The catch is that DirectX 12 will do nothing for that limitation.
At the end of the day all that this comes down to is decreasing the cost of one thing:
Draw()
Which has immense amounts of CPU overhead due to abstractions. Anything else DX12 might do is really just a bonus.That's not how it works, the app developer has no control over individual GPU cores (even in DX12). At the API level you can say "draw this triangle" and the GPU itself splits the work across multiple GPU cores.
The whole post isn't just oversimplified, it's just wrong. Across the board wrong wrong wrong. The point of Mantle, of Metal, and of DX12 is to expose more of the low level guts. The key thing is that those low level guts aren't that low level. The threading improvements come because you can build the GPU objects on different threads, not because you can talk to a bunch of GPU cores from different threads.
The majority of CPU time these days in OpenGL/DirectX is in validating and building state objects. DX12 and others now lets you take lifecycle control of those objects. Re-use them across frames, build them on multiple threads, etc... Then talking to the GPU is a simple matter of handing over an already-validated, immutable object to the GPU. Which is fast. Very fast.
The whole digression on lighting is mostly just wrong too. Deferred renderers have been rendering with 100s of dynamic lights for years. DX12 may make it a bit more efficient to deal with the large amount of constant data that needs to be updated when dealing with 100s of dynamic lights but it isn't introudcing any fundamental changes to dynamic lighting.
http://www.anandtech.com/show/8526/nvidia-geforce-gtx-980-re...
The GPU is creating threads and tasks internally and it's not always easy to balance this workload so no parts of the GPU becomes saturated while following parts in the chip's pipeline are idly waiting for work.
The PowerVR chips we're working with have dozens and dozens of different profile metrics corresponding to the different areas of its pipeline, each one being a potential bottleneck.
You could do something as silly as render a ball with 12k vertices instead of 24 and expecting the vertex processing to be much slower, but after profiling you find out its the fragment part lagging way behind because the data sequencer is overloaded trying to generate fragment tasks. In both cases you're rendering about the same amount of pixels.
With unified shader architectures, its very frequent for vertex and fragment tasks from different draw calls to overlap simultaneously. We're even seeing tasks from different render targets overlapping! Such as fragment tasks from the shadow pass still running when the solid geometry pass is processing its vertices.
But that's not entirely correct either. Yes you can use it like that, but you can also use it as a single core with a vector length of 1536.
In the context of over-simplification these are better thought of as single-core processors. The reason being if you have method foo() that you need to run 10,0000 times, it doesn't matter if you use 1 thread, 2 threads, or 8 threads - the total time it will take to complete the work will be identical. This is very different from an 8-core CPU where using 8 threads will be 8x faster than using 1 thread (blah blah won't be perfectly linear etc, etc).
That was the point of the blog post. Maybe described wrong, but that seems to be the point. The possibility of uploading stuff to the GPU on multiple threads on both CPU side, and the ability of the GPU to store those uploads in parallel. Maybe even render shadow-maps in parallel?
I wonder is there already the OpenGL equivalent of this? One of the main hurdles of OGL was issuing all the calls from the main thread, if that is gone now, that would be awesome.
Rendering shadow-maps in parallel would be pointless. If you render 2 at the same time, then each map gets half the GPU so an individual render takes twice as long, resulting in the same total time as if you gave each render 100% of the GPU and rendered in sequence.
> I wonder is there already the OpenGL equivalent of this? One of the main hurdles of OGL was issuing all the calls from the main thread, if that is gone now, that would be awesome.
Yes, NV_command_list extension:
http://www.slideshare.net/tlorach/opengl-nvidia-commandlista...
Can someone please comment on that? Everything I know about GPUs suggests that even if it would be possible to run VMs on them, performance would be terrible. All those tiny GPU cores are not designed to do much branch prediction, all the fancy out-of-order execution, pre-fetching etc. and your average software will totally depend on that.
The sentence that you snipped seems to refer to this. "Right now, this isn’t doable because cloud services don’t even have video cards in them typically (I’m looking at you Azure. I can’t use you for offloading Metamaps!)"
Furthermore, this has nothing to do with rendering APIs; it's all in the driver and hypervisor.
"there’s nothing stopping a DirectX 12 enabled machine from fully running VMs on these video cards."
muizelaar's comment pointed out that the author probably wanted to say that the GPU can't be reasonably accessed _from_ VMs, yet. That makes much more sense.
GPUs are incredibly slow at doing the sort of work that a CPU does. If they were fast at doing general purpose work, then CPUs would be doing the same things.
GPUs are extremely fickle with memory access patterns, branches, etc... It takes a lot of care to get them to run fast on a given workload, and that workload better be identically parallel (as in, every "thread" takes the same branches, with memory access in a uniform offset+thread-id order, etc..).
EDIT: Rephrased it better.
With valve developing source2, having ported almost all source games, and other AAA titles like Borderlands with first class Linux support, openGL it's very alive in gaming.
In general, Direct3D is designed to virtualize 3D hardware interfaces. Direct3D frees the game programmer from accommodating the graphics hardware. OpenGL, on the other hand, is designed to be a 3D hardware-accelerated rendering system that may be emulated in software. These two APIs are fundamentally designed under two separate modes of thought.
[1] http://en.wikipedia.org/wiki/Comparison_of_OpenGL_and_Direct...
https://msdn.microsoft.com/en-us/library/windows/desktop/gg6...
With feature parity latest DirectX can be compared only to desktop OpenGL (although OpenGL ES 3.0 is up and coming)
Ogl is busy winning some ground but on the level the article is talking about... Slow gains.
But OpenGL is also targeting and benefiting from these advances. Also, since the article is mindboggingly dumb, watch this: http://gdcvault.com/play/1020791/ (It might be that the DX12 announcement the author read was the source of misinformation, but anyway, this submission should have been flagged hours ago.)
Except that this won't only depend on the programmers, but on the management too, and seeing what those guys do at EA, UbiSoft, Activison, etc, I am not sure it will be properly utilized for the first wave of Dx12 games.
Edit: 3D
It's pretty crazy how weird it is that there's no dead simple 3D rendering library that supports just bare polygons. When I look at an opengl tutorial I just flee when I see how much boiler plate code is necessary to make the mouse move a camera.
> 3D rendering doesn't really get any more simple than a blank OpenGL context window.
OpenGL is pretty low level. It's a bare standard bridge to make use of a GPU. What I mean is that there are either very high level engines, or a very low level graphics API, which requires a lot of work if you want to make anything decent with it. There are things like glfw glew, but nothing like a very simple to use 3D rendering library, that just does the bare minimum, like a camera, quaternions, some text rendering, inputs, offers some simpler access to what opengl has to offer, etc.
3D engines are always being made obsolete. Except irrlicht and ogre3d, which are still fat enough in my opinion, there are no general-purpose, light 3D renderer, that are not necessarily trying to do it all. This kind of engine might serve as a thin wrapper to opengl to just avoid the work with the high quantity of opengl calls.
The bgfx examples seem to contain a lot of boiler plate code... It seems to be an impressive engine, it's recent so everything is up to date with recent graphics API. Pretty cool.
I'm more about making something simple and quick without the bells and whistles.
https://developer.apple.com/library/mac/documentation/3DDraw...
Also, it's going to be exclusive to Windows 10. Not the end of the world since 7 and 8 owners get a free upgrade.
Take some time to watch these. I'm also not a hardware guy, but these problems are far-far from the hardware, and the solutions are very educational from a distributed systems standpoint.
http://gdcvault.com/play/1020791/
http://www.slideshare.net/tlorach/opengl-nvidia-commandlista...
oh wait it only runs on "windows"