Capturing the WebGPU Ecosystem
developer.chrome.com
developer.chrome.com
https://machengine.org/pkg/mach-gpu-dawn/
It's also possible to use wgpu-native in C/C++ projects as prebuilt library, see:
[1] https://forum.babylonjs.com/t/why-havent-3d-web-games-reache...
The only thing good about it, is moving the Web 3D beyond GL ES 3.0 subset.
Don't expect nanite or Fortnite on the browser anytime soon.
WebGPU still has a nice balance between convenience and a somewhat modern feature set, but going down to the native platform 3D APIs (Metal, D3D12 and Vulkan) will always offer a much bigger feature set and less CPU overhead (mainly because WebGPU also needs to support mobile GPUs, and some of those are indeed 10 years behind the curve).
[1] https://registry.khronos.org/vulkan/specs/1.3-extensions/man...
[2] https://registry.khronos.org/vulkan/specs/1.3-extensions/man...
[3] https://registry.khronos.org/vulkan/specs/1.3-extensions/man...
[4] https://registry.khronos.org/vulkan/specs/1.3-extensions/man...
WebGPU outside of the browser doesn't really matter, middleware engines already cover that, and with much better tooling.
Besides WebGPU outside of the browser cannot deviate from the WebGPU on the browser, otherwise it is exactly the same pain as porting between OpenGL ES, OpenGL and WebGL, API that are only compatible in name.
In my opinion, the main strength of WebGPU is it removes the cost of re-implementing something when targeting a different platform. If that was present in 2015, maybe we'd have more games, more interesting applications, and less pressure on people who feel locked into one vendor's ecosystem.
I have a PR that does a lot of the same things (meshlets, visbuffer, material depth, two pass occlusion culling) open for Bevy https://github.com/bevyengine/bevy/pull/10164 that I've been working on, which uses WebGPU.
WebGPU is actually a pretty good API imo. It's missing some advanced features like raytracing, mesh shaders, and subgroup operations (coming soon!), but it can still do a lot.
The much bigger missing feature is "bindless" support (non-uniform arrays of bound resources). BindGroup overhead (and ergonomics) is a significant downside.
If you're ever looking to join a team to advance Unreal Engine 5 and graphics in the browser, please reach out. We're tackling Nanite and Lumen in the browser, and contributing to the WebGPU spec in collaboration with Google.
You can reach me on Discord - my username is astlouis44
I'm also on LinkedIn: https://www.linkedin.com/in/alex-st-louis-3986a5102/
Best of luck though, and always glad to see more people work on the WebGPU spec! Maybe you can help push for raytracing and bindless extensions :)
As a non game developer, though, I am wondering if migrating existing renderers would allow me to create UIs, or parts of them, using graphics that is not just html+CSS without the massive overhead that we currently have. I mean, I can draw the frames just fine but these extra cycles suck power and produce heat
https://felt.com/blog/svg-to-canvas-part-2-building-interact... and https://github.com/servo/pathfinder/ might be of interest.
Obfuscated assembly is often changed in such a way that there are no corresponding language constructs to "reverse" back to, if it originally came from a language like C++. Sure you can see the actual assembly code, but you can't recreate the higher level C++ constructs in a sane way, because the obfuscation technique scrambles things too much.
minified JS can be turned into reasonable JS, yes, but you're probably not going to get TypeScript code back, so the same sort of challenge exists there.
Assembly -> high-level language is harder, but there are absolutely binary -> C decompilers that are very popular/used in the RE community to make changes to existing programs.
But that doesn't even matter, WASM is much higher level than assembly, it's a stack machine, there is no arbitrary control flow / labels / `goto`, there are pre-defined data types, etc. all of this means it's easier to convert WASM -> high-level language than it is with a generic x86/arm binary.
There are WASM decompilers[0][1] which can convert WASM binaries into C code and back.
In both cases (minified JS and WASM), you're not going to get out exactly what you put in, but WASM doesn't really change the situation very much given the widespread adoption of 'compile to JS' languages like TypeScript these days.
[0] https://chromium.googlesource.com/external/github.com/WebAss...
You might not get the original C++ code back, but it is relatively easy to get an alternative version in C code back, which is already good enough as starting point.
And currently some high-frequency WebGPU calls have an even higher CPU overhead than similar WebGL2 calls (I wrote about that here a bit: https://floooh.github.io/2023/10/16/sokol-webgpu.html).
WebGPU allows more flexibility of course, compute shaders and storage buffers are more flexible and cleaner than trying to shoehorn the same stuff into vertex/fragment shaders and using textures as "poor man's storage buffer", but that doesn't mean that one approach is slower or faster than the other.
Most people are better served by using GL 4.6, DX 11 and such, or a middleware engine.
Even console vendors have multiple APIs because of that, not everyone needs all little details of the GPU.
Unity is supposed to have good DirectX 12 coming up, by the way.
"Achieving Real Time Ray Tracing on Xbox with Unity and DirectX12"
I've heard this repeated a lot in internet forums, but my experience working through hundreds of pages of Frank Luna's 800+ page DX12 book before concluding it pointless for 2D was that DX12 is actually fundamentally very similar to DX11 with most of the API focused on graphics rather than general compute. Compute shaders are just one chapter (13) of Luna's book, roughly 40 of the 800+ pages. I did some of the LunarG Vulkan tutorial and browsed the Khronos ref pages and reached a similar conclusion for Vulkan. I played with CUDA a bit and that's what real GPU programming looks like, almost no mention of graphics for much of the documentation. The "hello world" program isn't drawing a triangle, it's adding two arrays. Whereas a great deal of DX12, Vulkan, etc. is all about pixel formats, pixel shading, swap chains, geometry and tessellation, blending, depth and stencil, mipmaps and cube maps, clip coords, triangle winding and culling, the perspective Z divide, viewports, indexed and instanced draw calls, ... you know, graphics. But in chapter 9 of the DX12 book, end of 9.4, Luna writes, "Texture atlases can improve performance because it can lead to drawing more geometry with one draw call." So the conclusion I reached is that DX12 doesn't offer some fancy GPU compute way of writing a GPU program that can use 1000 distinct textures to draw 1000 distinct quads using only one CPU function call to launch this GPU program.
Now I've been doing more research and there is some sort of new feature called bindless textures not covered in Luna's DX12 book that might accomplish what I want (I'm not sure), but it seems to be Win11 only, WDDM 3.0 only, shader model 6.6 only, very new cards only. With this feature I might be able to set up 1000 distinct integer ids for my 1000 distinct textures, and then, with one single CPU draw call, have those 1000 textures applied to the correct 1000 quads, with no need to pack those 1000 textures into an atlas. Doing more web searching just now, this possibly can also be done in OpenGL on cards that support NV_gpu_shader5, but only semi-recent nVidia cards might support this. (I'm finding it difficult to get quick, quality answers to these sorts of questions using either web searches or LLMs.) Anyway, a gamedev forum or DX-focused reddit might be a better place for me to ask these sort of technical questions.
You will need DirectX 12 Ultimate or Vulkan for them.
https://developer.nvidia.com/blog/introduction-turing-mesh-s...
https://microsoft.github.io/DirectX-Specs/d3d/MeshShader.htm...
https://www.khronos.org/blog/mesh-shading-for-vulkan
https://devblogs.microsoft.com/directx/d3d12-work-graphs-pre...
https://gpuopen.com/learn/gpu-work-graphs/gpu-work-graphs-in...
So for the sake of future web searchers who discover this thread: there are only two proven ways to efficiently draw thousands of unique textures of different sizes with a single draw call that are actually used by experienced graphics programmers in production code as of 2023.
Proven method #1: Pack these thousands of textures into a texture atlas.
Proven method #2: Use bindless resources, which is still fairly bleeding edge, and will require fallback to atlases if targeting the PC instead of only high end console (Xbox Series S|X...).
Mesh shaders by themselves won't work: These have similar texture access limitations to the old geometry/tessellation stage they improve upon. A limited, fixed number of textures still must be bound before each draw call (say, 16 or 32 textures, not 1000s), unless bindless resources are used. So mesh shaders must be used with an atlas or with bindless resources.
Work graphs by themselves won't work: This feature is bleeding edge shader model 6.8 whereas bindless resources are SM 6.6. (Xbox Series X|S might top out at SM 6.7, I can't find an authoritative answer.) It looks like work graphs might only work well on nVidia GPUs and won't work well on Intel GPUs anytime soon (but, again, I'm not knowledgeable enough to say this authoritatively). Furthermore, this feature may have a hard dependency on using bindless to begin with. That is, I can't tell if one is allowed to execute a work graph that binds and unbinds individual texture resources. And if one could do such a thing, it would certainly be slower than using bindless. The cost of bindless is paid "up front" when the textures are uploaded.
Some programmers use Texture2DArray/GL_TEXTURE_2D_ARRAY as an alternative to atlases but two limitations are (1) the max array length (e.g. GL_MAX_ARRAY_TEXTURE_LAYERS) might only be 256 (e.g. for OpenGL 3.0), (2) all textures must be the same size.
Finally, for the sake of any web searcher who lands on this thread in the years to come, to pack an atlas well a good packing algorithm is needed. It's harder to pack triangles than rectangles but triangles use atlas memory more efficiently and a good triangle packing will outperform the fancy new bindless rendering. Some open source starting points for packing:
The "max number of sampled texture per shader stage" is a runtime device limit, and the minimal value for that seems to be 16. So texture atlasses are still a thing in WebGPU.
WebGPU has render bundles, which allow to pre-record command sequences, but even with that you don't want to change resource bindings thousands of times per frame (or even hundreds of times).
It might make sense though to build texture atlases dynamically (basically use one very big texture as "tile cache") and update that via writeTexture() calls (just don't rebuild the entire atlas each frame).
Not excited, and very few of those will materialise. Next gen games require next gen graphics. You can check for yourself how much those graphics weigh.
Wouldn't that make Chrome unusable if you don't have a recent GPU?
https://github.com/gpuweb/gpuweb/issues/4266
Also I guess Chrome could just fall back to a "legacy rendering backend" if WebGPU isn't supported.
IMO that's only really necessary if you are targeting hardware >5 years old, or embedded hardware, though.
I feel like this would only really get used for simple mobile games and not for actual titles with gameplay... even if the API is there, most people don't have powerful enough devices to drive modern console or PC quality games. And if they did, they'd be looking on Steam or an app store, not the Web.
What's the use case?
It also gets more interesting when you consider that the world is increasingly moving to ARM. Take Qualcomm's new ARM chips for example; they're competitive with Apple Silicon already. WASM/WebGPU represents a straightforward way to convert x86 games to run on these new chips.
Below is a blog post about our platform-as-a-service and tools, and our current development focus to support Unreal Engine 5 in WebGPU/WASM:
I enjoyed your Western Town demo. I always wished they would make virtual tours, Assassin's Creed-style, of old ghost towns like https://www.parks.ca.gov/?page_id=509
I wish we could agree on something slightly more ergonomic than a C FFI. The largest common denominator has to be slightly more advanced than C (if we decide to not accommodate C).
If you want to use WebGPU from as many languages as possible, a C API is pretty much the only option.
There's also a C++ API, but just as with the JS API, this is limited to a single language.
For usage from Rust you can also use gfx-rs/wgpu, which has a Rust-idiomatic API (e.g. see: https://github.com/gfx-rs/wgpu/blob/trunk/examples/hello-tri... - which to be honest doesn't look more "ergonomic" than the C API to me).
Also, it's not like the C API is "unergonomic", it works with C99 designated initialization which makes it quite nice to use. The only downside is COM-style manual lifetime management via ref/release calls, but if you are used to Direct3D that's a no-brainer (it works exactly the same), and the C++ API has RAII wrappers for this stuff.
Thanks.
Was the title a freudian slip or something?
As much as we all distrust Google, this kind of tangential comment is pretty tiresome when it has no relevance to what's being posted.
I’m surprised there isn’t much support for even a bytecode representation…