Former Nvidia Dev's Thoughts on Vulkan/Mantle
gamedev.net
gamedev.net
So ... the subtext that a lot of people aren't calling out explicitly is that this round of new APIs has been done in cooperation with the big engines. The Mantle spec is effectively written by Johan Andersson at DICE, and the Khronos Vulkan spec basically pulls Aras P at Unity, Niklas S at Epic, and a couple guys at Valve into the fold.
This begs the question: what about DirectX 12 and Apple's Metal? If they didn't have similar engine developer involvement, that's clearly a point against DX12/Metal support in the long run.
Also, the end is worth quoting in its entirety to explain why the APIs are taking such a radical change from the traditional rendering model:
Personally, my take is that MS and ARB always had the wrong idea. Their idea was to produce a nice, pretty looking front end and deal with all the awful stuff quietly in the background. Yeah it's easy to code against, but it was always a bitch and a half to debug or tune. Nobody ever took that side of the equation into account. What has finally been made clear is that it's okay to have difficult to code APIs, if the end result just works. And that's been my experience so far in retooling: it's a pain in the ass, requires widespread revisions to engine code, forces you to revisit a lot of assumptions, and generally requires a lot of infrastructure before anything works. But once it's up and running, there's no surprises. It works smoothly, you're always on the fast path, anything that IS slow is in your OWN code which can be analyzed by common tools. It's worth it.
I wonder if/when web developers will reach a similar point of software layer implosion. "Easy to code against, but a bitch and a half to debug or tune" is an adequate description of pretty much everything in web development today, yet everyone keeps trying to patch it over with even more "easy" layers on top. (This applies to both server and client side technologies IMO.)
The thing is while gaming is seen as a "performance critical" application, where it is natural to want the maximum performance to be squeezed out of your hardware, Web development is just a behemoth of inefficiency, and we think that that is the norm, an inevitable drawback of developing for the web, an inescapable trait of web applications. Web apps are dozens or even hundreds of times slower than native applications, and we think nothing of it, because it seems to be an inherent characteristic of the web about which we can't do anything.
Your comment made me reconsider that. Maybe it doesn't have to be so. What if there was a Vulkan for the web, instead of the layers upon layers of inefficiency which are all the rage these days? What if I could run a medium complexity web page or app without it bogging my computer several times harder than a comparable native app?
Edit: Seriously, how feasible is that? Creating a GUI framework on top of WebGL and creating an abstraction layer (which the original article advises against) would be a huge performance boost comparing to current webpages, wouldn't it?
[0] http://zephyrosanemos.com/windstorm/current/live-demo.html
Edit: To be clear, I don't think writing a web rendering engine meant to do the same think as the browser on top of WebGL will be faster in any respect, as really the HTML and CSS are just inputs fed into super optimized low level rendering engines. I don't see how JS could compete in that domain.
No file systems, no network protocols, no processes; just a very thin layer that lets userspace implement those if/how it wants to.
HTML/CSS was designed for requirements that don't apply to many modern websites. I would not be surprised if it were possible to implement a rendering engine that performed better for relevant use cases, e.g. perhaps assuming a grid as a first-class part of the layout.
We've started creeping that way with a lot of the newer HTML5 stuff, but we have a long way to go.
It's in very early stages. Feedback is very welcome:
If you do see any performance benefits it likely will be because you support less features than what a modern day browser supports with HTML/CSS. But in my opinion you could probably achieve faster speeds with Canvas in a similar manner.
It depends on what you are doing. Rendering thousands of sprites with independent rotation and scale is easy. Rendering text is a lot of work, but once you do it you can do all sorts of things with the text. Rendering simple things like bezier curves is very difficult. There are some things in the 2D world that HTML/CSS/GDI/GDI+ are very bad at, compared to OpenGL, and vice versa.
All web content on OS X and iOS is rendered using OpenGL. That's not to say it was easy to make efficient, but it definitely outperforms the default CoreGraphics implementation on most content.
You have to be ambitious, right? :)
Though as an aside, I think voice control is very underutilized in modern GUI development...
I just wonder how well text performance will be. Would be awesome to get near native speed with editors like Brackets and Atom.
Let's imagine a world where a browser is just a secure place to execute code. You get vulkan, io, networking, input, camera access in a cross platform way.
So, how do search engines find any info? How do you share links to content? How do you support braille readers? Browser extensions?
I'd argue those things are pretty essential to the web as we know it (well maybe not the extensions) and I might also argue that some of the inefficiencies of the web environment are a direct result of making the decisions that the most important part of the web is a common semi scannable and linkable format, HTML.
javascript:document.querySelector('meta%5Bname=viewport%5D').setAttribute('content','width=device-width,initial-scale=1.0,maximum-scale=10.0,user-scalable=1');Desktop integrated search.
> How do you share links to content?
Extensions, contracts and intents.
> How do you support braille readers?
OS Accessibility APIs
> Browser extensions?
Plugins
Indies are going to increasingly look at the web and other interfaces because of the easy layers.
(And professional studios don't have to license the ancillary tools and services and source code access of the Pro editions in order to ship a game.)
It will make it to indie engines that aren't companies funded as well eventually.
This actually gives a ton of power to current engine developers in the market so getting them involved was a big win.
I am a huge fan of OpenGL ES / WebGL (also based on OES) and was happy that we finally had mobile win that battle but with new graphics rendering layers like Metal and optimization for mobile more needs to be done. The driver mess is a big problem and limits innovation as well.
Now that the feature set of graphics hardware is very stable and uniform across different vendors' hardware, the APIs are needed to solve an entirely different problem from what they were invented to do.
There's several parts to this. DirectX 12 is going to run on Windows 10, and its adoption by developers is probably going to be entirely driven by how much of their target audience can be persuaded to upgrade from Windows 7. If Windows ships with better DirectX 12 drivers than Vulkan drivers (and I would bet that it will), then you'll probably see more AAA games written to DirectX12. Metal, likewise, is the preferred graphics API for reaching iPhone users, which are plentiful and (compared to Android users) lucrative.
I've been a full-time GPU engineer at a couple of major companies for about 6 years now. My colleagues at my current company have all been GPU engineers (and video, prior to that) in the 20--30 year range.
While a lot of the technical details are correct, there are a lot of red-flags in this guy's essay. For instance:
> Although AMD and NV have the resources to do it, the smaller IHVs (Intel, PowerVR, Qualcomm, etc)
Intel owns, what? 60? 70 percent of the desktop market? More? PVR, Qualcomm, etc. own 90% of the mobile market? Intel has an enormous number of GPU engineers---far more than AMD, and easily comparable to (or more than) NV.
Now, to OP's statement about DX12 & Metal: obviously they were developed cojointly with AAA game developers. The major title developers are in constant contact with every major IHV and OS vendor (less so for open-source) at all stages with respect to driver- and API- development.
Furthermore, it's not like the major architects and engineers of these driver teams are total tools; these guys live and breathe GPU architecture (and uarch), and are intimately familiar with what would make a good driver.
The impetus for these low level APIs is differentiation and performance. When those couldn't be had from GPU HW, the next place to look is "one up the stack": driver & API. It surprised no one in the industry that Mantle/Vulkan/Metal/DX12 all came out at about the same time: we've all been pushing for this for years.
for me it was 3 years ago in web. And 6 years ago in Enterprise java.
It's hard to imagine someone working with Enterprise Java and thinking "This isn't helping. This isn't making things more reliable".
I still love JS, and coding for the web. But I don't use frameworks now.
At one point you want full control of the HW, much like you did with game consoles..
On the other you want security: This model must work in a sandboxed (os, process, vm, threads, sharing etc.) environment, along with security checks (oldest one that I remember was making sure vertex index buffers given to the driver/api must not reference invalid memory, something you would make sure is not the case for a console game through tests, but something that the driver/os/etc. must enforce and stop in a non-console game world - PC/OSX/Linux/etc.)
From little I've read on this API, it seems like security is in the hands of the developer, and there doesn't seem to be much OS protection, so most likely I'm missing something... but whatever protection is to be added, definitely would've not been needed in the console world.
Just a rant, I'm not a graphics programmer so it's easy to rant on topics you just scratched the surface...
----
(Not sure why I can't reply to jeremiep below), but thanks for the insight. I was only familiar with one I posted above (and that was back in 1999, back then if my memory serves me well, drawing primitives on Windows NT was slower than 95, because NT had to check all index buffers whether they were not referencing out-of-bounds, while nothing like this was on 98).
Invalid data would simply cause a GPU task to fail while the other tasks happily continue to be executed. Since they are pure and don't interact with one another there is no need for process isolation or virtualization.
Basically, its easy to sandbox a GPU when the only data it contains are values (no pointers) and pure functions (no shared memory). Even with the simplified model the driver still everything it needs to enforce security.
1. https://software.intel.com/en-us/blogs/2013/07/18/order-inde...
2. http://beta.ivc.no/wiki/index.php/Xbox_360_King_Kong_Shader_...
This is just hn adding a delay until the reply link appears related to how deeply nested the comment is. The deeper the longer the delay. It's a simple but effective way to prevent flame wars and the likes.
Well, perhaps a better way of putting it is that "easy" APIs try to model a domain. Sometimes the API designer just nails it, and folks forget eventually take it for granted. The problem is now "solved". For example, the Unix userspace filesystem API (open, close, read, write, etc.) is pretty damn solid for what it is. I admit that I took it for granted until I worked at a place where someone who'd never seen a filesystem API (!!) was tasked with creating one from the ground up. You can guess how that went.
The downside case is where the API's model either never cut it in the first place, or as @wtallis illuminates above with OpenGL, where the model just ages poorly[1].
The Unix filesystem api isn't complete, if it were ioctl wouldn't exist. Assumptions all around, and ioctl is the way you adjust those assumptions.
Much like the simultaneous ubiquity of C/C++ and Python, in their respective spaces.
I don't think anything will stop someone from being able to run vulkan directly on to of libdrm like you can right now with GL. At least for open source drivers like Intel's.
Will doing things the "good but hard" way force me to write tons of boilerplate? Because as a (currently iOS) programmer, as soon as I start having to write generic, structural code, I find it difficult to sustain my interest in the project. I want to go from idea to something that works as quickly as humanly possible; having to remake the universe from scratch just because somebody somewhere needed the performance is not what I got into the game for.
(I'm currently working with OpenGL for the first time, and, ugh, all the state management stuff and shader boilerplate is driving me crazy — especially after years of the relative ease of dealing with UIKit! Am I to understand that Vulkan will only make things worse in this regard? I guess it's good for engine devs, but I would not want to code like this in my day-to-day.)
Whether or not the Extensible Web Manifesto is successful or not remains to be seen.
However! Having written OpenGL 3 code before, the idea that Vulkan is going to be a lot more complex, frankly scares the hell out of me. I'm by no means someone that programs massive 3D engines for major companies. I just wanted to be able to take a bitmap, optionally apply a user-defined shader to the image, stretch it to the size of the screen, and display it. (and even without the shaders, even in 2015, filling a 1600p monitor with software scaling is a very painful operation. Filling a 4K monitor in software is likely not even possible at 60fps, even without any game logic added in.)
That took me several weeks to develop just the core OpenGL code. Another few days for each platform interface (WGL/Windows, CGL/OSX, GLX/Xorg). Another few days for each video card's odd quirks. In total, my simple task ended up taking me 44KB of code to write. You may think that's nothing, but tiny code is kind of my forte. My ZIP decompressor is 8KB, PNG decompressor is another 8KB (shares the inflate algorithm), and my HTTP/1.1 web server+client+proxy with a bunch of added features (run as service, pass messages from command-line via shared memory, APIs to manipulate requests, etc) is 24KB of code.
Now you may say, "use a library!", but there really isn't a library that tries to do just 2D with some filtering+scaling. SDL (1.2 at least) just covers the GL context setup and window creation: you issue your own GL commands to it. And anything more powerful ends up being entire 3D engines like Unity that are like using a jack hammer to nail in drywall.
And, uh ... that's kind of the point of what I'm doing. I'm someone trying to make said library. But I don't think I'll be able to handle the complexity of all these new APIs. And I'm also not a big player, so few people will use my library anyway.
So the point of this wall of text ... I really, really hope they'll consider the use case of people who just want to do simple 2D operations and have something official like Vulkan2D that we can build off of.
Also, I haven't seen Vulkan yet, but I really hope the Vsync situation is better than OpenGL's "set an attribute, call an extension function, and cross your fingers that it works." It would be really nice to be able to poll the current rendering status of the video card, and drive all the fun new adaptive sync displays, in a portable manner.
Actually it's par for the course, if you care about backwards compatibility. Raymond Chen has many, many stories about the absolutely heroic lengths Microsoft has gone to ensure that popular, yet broken, programs continue to work smoothly between Windows upgrades [1][2][3][4][5].
On Depending on Undocumented Behavior:
[1] http://blogs.msdn.com/b/oldnewthing/archive/2003/12/23/45481...
[2] http://blogs.msdn.com/b/oldnewthing/archive/2003/10/15/55296...
Why Not Just Block Programs that rely on Undocumented Behavior?
[3] http://blogs.msdn.com/b/oldnewthing/archive/2003/12/24/45779...
Who cares about backwards compatibility? (A lot of people):
[4] http://blogs.msdn.com/b/oldnewthing/archive/2006/11/06/99999...
Hardware breaks between upgrades, too:
[5] http://blogs.msdn.com/b/oldnewthing/archive/2003/08/28/54719...
Joel Spolsky's (somewhat dated) "How Microsoft Lost the API War" gives a good overview of the above concerns:
Worse, they rarely check the limits, only the basics. They also they didn't check anything related to multiple contexts. The point being the drivers are/were full of bugs on the edge cases.
Testing works. If there were tests that were relatively comprehensive and that rejected drivers that failed the edge cases they'd gone a long way to mitigating these issues because the dev's apps wouldn't have worked.
There's also just poor api design. Maybe poor is a strong word. Example, uniform locations are ints so some apps assume they'll be assigned in order then fail when they get to a machine where they're not. Another example, you're allowed to make up resource ids. OpenGL does not require you to call `glCreateXXX` just call `glBindXXX` with any id you please. But of course if you do that maybe some id you use is already being used for something else. So id=1 works on some driver but not some other.
I'm excited about Vulkan but I'm a little worried it's actually going to make the driver bug issues worse. If using it in a spec compliant way is even harder than OpenGL and your app just happens to work on certain drivers other driver vendors will again be forced to implement workarounds.
https://gitorious.org/bsnes/bsnes/source/1a7bc6bb8767d6464e3...
The platform abstraction layers are wgl.cpp, cgl.cpp and glx.cpp
Go inside the opengl/ subfolder to find all the platform-agnostic OpenGL code. opengl/surface.hpp gets particularly fun using matrix multiplication to compute model view / projection / texture coordinates by hand (which you need to do for the new GL3 / no-fixed-function-pipeline stuff.) Also has lots of required internal allocations.
In my own case, I allow user-defined additional shader passes, and that adds to the code a bit. But note that another thing about GL3 is that you do actually have to create and execute at least one vertex + one fragment shader.
GL2's FFP, while still difficult, was a whole lot easier. The driver did a lot of the work you have to do manually if you want to follow GL3.2-core/GL ES/etc.
While not a 2D library, https://github.com/bkaradzic/bgfx is a very sweet abstraction layer over platform graphics APIs
These are the vanishingly few people who have actually
seen the source to a game, the driver it's running on,
and the Windows kernel it's running on, and the full
specs for the hardware. Nobody else has that kind of
access or engineering ability.
One option is to release the source of the drivers, making it possible for motivated engine developers to do this without explicit access to AMD developers.If they did this on Linux as well, then developers would have access to the full stack, in order to be able to learn about how the lower levels work and more easily track the problems down without simply a whole lot of trial and error on a black box.
I guess with Vulcan/Mantle, the drivers are just trivial hardware abstraction layers, so the GPU manufacturers no longer have to make up for the cost of developing complex drivers by doing part of the work on every AAA title.
We might see Nvidia and AMD releasing open source Vulcan drivers, and if not, it would at least be easier for reverse-engineering projects like Nouveau to produce competitive drivers. This at least means you don't have to have non-free code running in ring-0 to use a discrete graphics card on Linux.
I've never actually dived into the source code for the open source video drivers but I'm now curious how much time the devs of the open-source drivers have to spend on anticipating and routing around the brain damage of the programs calling them. Do they similarly try to find a way to correctly do the right thing despite the app crashing if it finds out it didn't get the wrong thing as asked for? Or is the attitude more akin to 'keep your brain damage out of our drivers and go fix your own damned bugs'?
Open source driver developers don't have the time or resources to pull off a stunt like that!
I think that in some cases the developers (at least Intel) have reached out to game developers first, to get them to fix their own bugs/out of spec behavior.
While at GDC I saw a DirectX 12 API demo (DX12 is more/less equivalent to Vulkan from an end-goal perspective). On a single GTX 980:
DX11 was doing ~1.5 million drawcalls per second. DX12 was doing ~15 million drawcalls per second.
This API demo will ship to customers, so I am pretty sure we can easily verify if these are bunk figures. But a potential 10x speedup, even if under ideal conditions, is notable.
I'm wondering though, if the demo also uses the CPU for other things - physics, audio, collision, path-finding or some other form of ai, state machines, game script, game logic. My point is that 10x might be possible (on a 10 core cpu) if the cpu's are only used for graphics, but there are other things that come into play... But even then, even if only half the cpu's are used for graphics, it's still better.
The bigger question to me, is how would game developers on the PC market (OSX/Linux included) would scale their games? You would need different assets (level of detail? mip-mapped texture levels? meshes?) - but tuning this to work flawlessly on many different configurations is hard...
Especially if there are applications still running behind your back.
E.g. - you've allocated all cpu's for your job, all to be taken by some background application, often a browser, chat client, your bit-coin miner or who knows what else.
The bigger question to me, is how would game developers
on the PC market (OSX/Linux included) would scale their
games? You would need different assets (level of detail?
mip-mapped texture levels? meshes?) - but tuning this to
work flawlessly on many different configurations is
hard...
This isn't really any different from what it has been until now. All AAA games have different levels of detail for meshes/textures/post-processing etc. Even when not exposed to the user as options in a menu these different levels of detail exist to speed up rendering of for example distant objects or shadows where less detail is needed. DX12/Vulkan is not going to change anything in that regard.Doing a good PC port is not as easy as it may seem at first glance. Different hardware setups and little control over the system cause lots of different concerns that simply don't exist on consoles, which means nobody bothered taking that into account when the game was originally built. These new APIs will help though; the slow draw calls on PC are a pain compared to lightning fast APIs on consoles (even Xbo360/PS3!).
Its possible that these perf gains are actually working in a single thread, and the gains are from eliminating the default driver overhead that would be in these drawcalls. To some extent this could be mitigated by batching but it is still ideal to have the option to do far more drawcalls in a frame.
Valve was showing DOTA 2 running on Vulkan with Source 2 at greate framerates too (with many peons on screen) running on... Intel HD graphics.
So it does seem that we will get a large performance boost with Vulkan / DirectX 12.
So true and yet, we have all those crazy JavaScript frameworks trying to abstract everything away from developers. There's a lesson in there.
From the javascript frameworks - I use only backbone and jquery - first as a simple router, second for easy dom traversal.
This is so true. You can, however, successfully abstract away a bunch of boilerplate for common actions, and enforcing patterns. The key to a framework (and skill of the developer), is knowing when and how to avoid the abstractions.
Believing you're somehow avoiding abstraction because you only use a couple of additional libraries on top of JS and the DOM is like insisting on a 99th storey apartment instead of an 100th storey one, because you prefer being close to the ground.
You can by all means argue that all the big JS frameworks are poor abstractions. But a sweeping statement like "you cannot abstract away complexity" completely ignores the fact that web development as a field is only possible because of the successful abstraction of huge quantities of complexity.
Similarly, since with web browsers we get what we get, we might as well consider that our reasonable base for web development, it's not like anyone is going to do that part differently any time soon.
The browser severely limits what you're able to do with the network connection, but at the end of the day dealing with the network is a complex beast.
If you doubt, look no further than the HTML5 caching API's. The tooling around them is absolutely terrible, but even if it weren't there's an inherent complexity with that sort of thing.
Often it's also a question of how things are composed, not just how much they're abstracted away. If you're use case dictates that you need to separately control something that's been composed into one thing then you're usually out of luck
No, you have a selection bias, you don't notice the abstractions that have worked.
Oh sure you can. It just costs something. The most common tradeoff is that you abstract away complexity in favor of performance.
that makes no sense.
Perhaps what you meant is that there is a limit to how much a problem can be simplified.
More companies are starting to support OpenGL, but I'm just curious as to why uptake is so slow. It seems like poor API design may be a part of it. I'd like to see more games written with OpenGL support, and I think it's happening slowly but surely. We even see weird hacks add Linux support at this point... Valve opensourced a D3D -> OGL translation layer[0], though hasn't supported it since dumping it from their source.
To be fair, driver vendors have gotten a lot better. And the spec itself has matured tremendously, with some great features that don't even have direct equivalents in Direct3D.
Sadly as a whole the API still sucks, especially if you care about supporting the vast majority of the audience who want to play video games. Issues I've hit in particular:
Every vendor's shader compiler is broken in different ways. You HAVE to test all your shaders on each vendor, if not on each OS/vendor combination or even each OS/vendor/common driver combination. Many people are running outdated drivers or have some sort of crazy GPU-switching solution.
The API is still full of performance landmines, with a half-dozen ways to accomplish various goals and no consistent fast-path across all the vendors.
At least on Windows, OpenGL debugging is a horror show, especially if you compare it with the robust DirectX debugging tools (PIX, etc) or the debugging tools on consoles.
Threading is basically a non-starter. You can get it to work on certain configurations but doing threading with OpenGL across a wide variety of machines is REALLY HARD, to the point that sometimes driver vendors will tell you themselves that you shouldn't bother. In most cases the extent of threading I see in shipped games is using a thread to load textures in the background (this is relatively well-supported, though I've still seen it cause issues on user machines.)
OpenGL's documentation is still spotty in places and some of the default behaviors are bizarre. The most egregious example I can think of is texture completeness; texture completeness is a baffling design decision in the spec that results in your textures mysteriously sampling as pure black. Texture completeness is not something you will find mentioned or described anywhere in documentation; the only way to find out about it is to read over the entire OpenGL spec, because they shoved it in an area you wouldn't be looking to debug texturing/shading issues. I personally tend to lose a day to this every project or two, and I know other experienced developers who still get caught by it.
I should follow with the caveat that Valve claims GL is faster than Direct3D, and I don't doubt that for their use cases it is. In practice I've never had my OpenGL backend outperform my Direct3D backend on user machines, in part because I can exploit threads on D3D and I can't on GL.
So, essentially, no validation, you have to manage your own buffers (with some help in DX12 I think), you can shoot yourself in the foot all day long. But if you manage to avoid that, you are able to reduce overhead and use multithreading.
I always was interested in game programming, but never was able to really get interested enough in graphics programming, I guess having a messy API is not an excuse, but you can really sense that CPUs and GPUs really have different compatibility stories, and that's maybe why it's not attracting enough programmers.
I still hope that one day there might some unified compute architecture and CPUs will get obsolete. Maybe a system can be made usable while running on many smaller cores ? Computers are being used almost exclusively for graphic application nowadays, I wonder if having fast single core with fat caches really matters anymore.
Enough for what? It's not as if the world is short of games, or game engines.
Game companies have, in general, never had difficulty attracting programmers, especially young programmers. That's why they have such a reputation for low pay and crappy work conditions: because they can.
I haven't heard of DX12 getting overhauled for efficient multi-threading or great multi-GPU support. DX12 probably brings many of the same improvements Mantle brought, but Vulkan seems to go quite a bit beyond that. Also, I assume DX12 will be stuck with some less than pleasant DX10/DX11 legacy code as well.
It's possible Khronos will leapfrog DX12 with Vulkan in regards to threading but I find it highly unlikely. We'll know when one of the two actually publishes documentation (likely not soon, based on how long it took for Mantle to become available to the public)
For example, in OpenGL, you can upload a texture with glTexImage2D(), then draw with glDrawElements(), then delete it glDeleteTextures(). The draw command won't be complete yet, but the driver will free the memory once it's no longer being used.
It sounds like with Vulkan, you'll need to allocate GPU memory for your texture, load the memory and convert your data into the right format, submit your draw commands, and then you'll need to WAIT until the draw commands complete before you can deallocate the texture. At every step you're doing the things that used to be automatic. So it's harder to use, but you're dealing with more of the real complexity from the nature of programming a GPU and less artificial complexity created by the API.
In general most of the additional steps are things people already do. e.g. you shouldn't be deleting textures that are in use anyway because not all drivers have always handled that well, etc, managing the equivalent of a command buffer is common, even when the actual command buffer isn't something you have access to...
There is no waiting to free resources however, unless either the CPU or GPU is starving for work. We have a triple-buffering setup and on consoles you also get to create your own front/back/middle surfaces as well as implement your own buffer swap routine. This provides a sync point where we can mark the resources as safe to release.
It's definitely more complexity on the engine part, but as mentioned in the forum post it makes everything much, much easier when you get to debug and tune things. Also having to implement (or maintain) all of that engine infrastructure gives us a better perspective into how the hardware works and how to optimize for it.
However, even with Vulkan or DirectX12 I doubt NVidia or AMD will expose their hardware internals publicly which is critical in optimizing shader code. On consoles we get profilers able to show metrics from all of the GPU's internal pipeline. It makes it easy to spot why your shader is running slow without having to send your source code to the driver vendor.
I haven't had to profile an AAA title for the desktop so far and therefore don't know much about the state of tools there. However, I heard only good things about Intel's GPA.
I'm sure there's a counterargument that it raises the bar for indie game devs, but when's the last time an indie game directly programmed against DirectX or OpenGL (barring webGL)? This should let engine developers better use their development time.
I don't work in indie, but my understanding is that this is still pretty common, especially for games with any budget at all.
Most indie games can happily use Unity or whatever, but there are cases where they are not suitable.
This is really sad. Imagine if someone were pushing a new file API with the justification that storage-device drivers were full of bugs.
I get a similar feeling whenever I read one of those "CSS trick that works on all browsers!" articles. Yes, it's nice that you can build a thing of beauty on top of broken abstractions, but ....
The open source drivers have frequently been written by volunteers, or professionals who do it as only a small part of their job, have had to spend more time and effort reverse-engineering rather than having documentation and the ability to talk to hardware engineers, and are perpetually at least a few months to a few years behind the proprietary developers who had access to all of this before the hardware even came out.
It's impressive what the open source drivers have managed to accomplish, and they are improving, but they are severely handicapped compared to the proprietary drivers.
This excludes, of course, Intel, but Intel focuses on lower end integrated graphics rather than high-end discrete graphics, so it's not quite comparable.