Why would I blame NVIDIA? If it wasn't for them, we'd still only have needlessly cumbersome APIs and ecosystems. They did what Khronos always failed to do: They created something that is both easy, powerful and fast. Khronos always heavily neglects the easy part.
Not happening. WGSL wants to support the lowest common denominator, so it'll always mainly be a 5-year old mobile-phone API. Also if you want to beat CUDA, you'll need some functionality that's completely missing in compute shaders, especially WGSL. Like pointers and pointer casting (and that glsl buffer reference extension is the worst emulation of that feature I've every seen).
What's wrong with CUDA? I avoided it for years because it's proprietory but about one year ago I started using it because all the alternatives (OpenGL/Vulkan compute, OpenCL, WebGPU, ...) couldn't quite do what I wanted, and it turned out to be a game changer. Nothing comes close to it. Now I'm hooked because there simply isn't an alternative that's as easy to use, yet powerfull and fast.
I wish there was an open alternative, but NVIDIA did several things right that others, especially Khronos, do not: The UX is top-notch. It makes the common cases easy yet still fast, and from there you can optimize to your hearts content. Khronos, however, usually completely over-engineers things and makes the common case hard and cumbersome with massive entry barriers.
It's used in one of the fastest sorting approaches - counting sort / binning - to compute the location of where to store the sorted/binned items. First you count the number of items per bin, then you use prefix-sums to compute the memory location of each bin, then you insert the items into the respective bins. Some radix-sort implementations also utilize counting sort under the hood, and therefore prefix-sums. (Not sure if all radix-sort implementations need it)
It's not just micro-triangles, it works for triangles that span multiple pixels. And that's not a special case, that's the standard nowadays, except for games targeting very low-end devices.
The dedicated hardware is nice for general-purpose support for arbitrary triangle-soups. But if you structure triangles a certain way and have a certain amount of density (which you want for modern games), you can specifically optimize for that and beat the general-purpose hardware rasterizer.
Huh? It had other options longer than it had async/await. Async/await is a fairly recent addition. E.g. before fetch with async/await, there was XmlHTTPRequest with callbacks. It also had Web Workers as a means for parallel&concurrent processing for way longer than it had async/await.
The standard rendering pipeline is still fairly fixed and mandates shaders for vertices and fragments. With software rasterization, there are no more vertex or fragment shader, only compute. And you don't use the hardware rasterization units - you rasterize the triangles yourself and write the results to the framebuffers with atomic-min/max operations. You decide for yourself how and when you compute the shading in your compute shader. This can be multiple times faster for small triangles, and 10-100 times faster for points. And once you do things that way, there isn't much point for graphics APIs anymore - everything is just buffers and functions that process them.
Not really trivial if you have to support any input set of triangles and don't know much about them. Software rasterizion exploits things like localized chunks of triangles, but the hardware rasterizer does not know about that in advance. Also, these kinds of software rasterization algorithms add limitations, which aren't much of an issue with your own rendering pipeline that specifically knows how to deal with these limitations and works with them.
It's the exact reason why I've avoided CUDA for years, but I hit a dead end with OpenGL and Vulkan, and CUDA happened to be a fantastic, easy and fast solution. Of course I don't want graphics programming to be NVIDIA-only, but I want it to be like CUDA, just for all platforms.
Exactly. Things like that already work with a workaround: You can use Cuda-OpenGL interop to expose an OpenGL framebuffer in CUDA, then you can simply write into that framebuffer from your CUDA kernel, and afterwards you get back to OpenGL to display it on screen. Just directly integrate that functionality in CUDA by providing a CUDA native framebuffer and a present(buffer) or buffer swap functionality.
Yeah, but OpenGL doesn't get updates anymore. My timeline goes like: I needed pointers(and pointer casting) for my compute shaders so I checked the corresponding GLSL extension, which was only available in Vulkan so I tried switching from OpenGL to Vulkan. After a week I gave up - the pointer/Buffer reference extension did not look promising anyway - and I tried out CUDA instead. That's when I found out that CUDA is the greatest shit ever. That's what I want graphics programming to be like. Since then I just render all the triangles and lines in CUDA, because it easily handles hundreds of thousands of them in real-time with a naive, unoptimized software-rasterizer, and that's all I need. In addition to the billion points you can also render in real-time in CUDA with atomics.
I meant the development/learning overhead. With Vulkan you can do incredible low-level optimizations to squeeze every last bit of performance out of your 3D application, but because you are basically mandated to do it that way, you have a very harsh learning curve and need lots of code for the simplest tasks. I'd rather prefer approaches that make the common things that everyone wants to do easy (draw your first simple scenes), and then optionally gives you all the features to sqeeze out performance where you really need to. Because at least for me, I don't work with massive scenes with millions of instances of thousands of different objects. I do real-time graphics research, mostly with compute, and I'd just like to present the triangles I've created or the framebuffers I created via compute (like shadertoy).
I'll readily admit, I'm neither smart nor patient enough for Vulkan so I quickly gave up and learned CUDA instead, because it was way easier to write a naive software-rasterizer for triangles in CUDA, than it was to combine a compute shader and a vertex+fragment shader in Vulkan. I'm just rendering a single buffer with ~100k compute-generated triangles, and learning Vulkan for that just wasn't worth it.
Yes, sorry for the confusion. I'd just rather have an API where the common things are easy, and the super powerful low-level optimizations are optional.
> Vulkan and OpenGL both already support mesh shaders, which is a compute-oriented alternative to the traditional rasterization pipeline.
Mesh shaders are a step in the right direction, but they are still embedded in all that unnecessary Vulkan fluff. I would want these things in CUDA because it does the opposite approach of Vulkan - it makes the common things easy, and the hard/powerful things optional. Just let me draw things directly in CUDA, and maybe give access to the hardware rasterizer via a CUDA call.
I know it won't happen overnight, but since it's already possible to do software rasterization for small triangles faster than hardware, having a graphics API framework starts losing its purpose. After all, we want to have the detail of small triangles anyway. Just let us draw to the screen in CUDA without the need for OpenGL/Vulkan interop, and I believe we'll soon see a shift to serious compute-based real-time rendering.
Basically, instead of graphics being a framework, I want graphics to be a straightforward library you include and use in your CUDA/HIP/SYCL/OpenCL code.
Part of Nanite is software rasterization by rendering triangles with 64 bit atomics. You can simply draw the closest fragment of a triangle to screen via atomicMin(framebuffer[pixelID], (depth << 32) | triangleData).
Glad to see more support for OpenGL, but I really hope we'll soon move to a compute-only way of handling graphics. The overhead of vulkan is absolutely insane (and not warranted, in my opinion), and OpenGL is on its last legs.
Things like Nanite spark a little hope, since they've shown that software-rasterization via compute can be faster than the standard graphics pipeline with the hardware rasterizer for small, dense triangles. Seems like a matter of time until everything goes compute, even large triangles. Maybe the recent addition of work-graphs in DirectX is one step in that direction?
I strongly disagree. Async/Await is one of the nicest, cleanest ways to deal with asynchronous tasks, and asynchronous tasks are everywhere. Loading data from disk without blocking and doing something once this is done -> async. Memcpying data from CPU to GPU without blocking -> async. Memcpying data from GPU back to CPU without blocking -> async.
Sure, there are other ways to handle these things like polling state, but async makes it trivial and readable.
But this mostly applies to JS which has a kickass implementation of async/await. Whatever C++ tried to do, it's an awful mess so there I still use threads with busy loops, polling, etc., whatever makes sense for the task at hand.
I think that's the problem. Khronos isn't known for good UX, and being from Khronos is exactly the reason why I'm not even bothering to check it out. I want an alternative to CUDA, but I also want it to be as easy to use as CUDA.
Many of the best papers appear an Arxiv first. In some fields, it is customary to put your preprint on Arxiv before/during the submission to the peer reviewed venue.
Arxiv is vital for quickly developing research fields.
The issue with Vulkan is that it's so cumbersome, I just don't want to use it at all. Tried to switch, but went back to OpenGL. There isn't anything in Vulkan that warrants that insane amount of overhead.
WebGPU is based on similar modern paradigms but it's an actually usable, sane, not overly verbose API. It's a pitty that WebGPU is deliberately gutted to make it a 5 year old smartphone graphics API because I'd prefer it over Vulkan any day.
I'm glad that Vulkan is taking some steps back on render passes, pipelines, etc. If they keep removing unnecessary entry barriers, I might eventually try the switch again in a few years.
Because browsers are by far the easiest, safest, and fastest way to distribute applications. Operating Systems still don't have any sort of meaningful sandboxing so downloading and executing binaries from any source is out of question. With web applications, you can do that. Instantly and uncomplicated. This is not going to change anymore, it's just way too useful.
The problem is that GDPR doesn't go far enough. It should completely prohibit third party tracking without the possibility to agree to it. Third party cookies/tracking simply should not exist.
One reason is that WebGL isn't a 2010s technology, but more of a 2005 technology (it doesn't even have compute shaders). WebGPU will finally bring the Web to the state of 2010.