HNHacker News
TopNewBestAskShowJobs

jms55

547 karma · joined August 16, 2020

submissionscomments
jms55··on Programming SDF Animations of Rick and Morty
Ehh from the user's perspective sure, but under the hood it's different.

Vulkan you're doing GLSL -> SPIR-V

WebGPU you're doing GLSL -> WGSL -> (HLSL->DXIL) / (MSL->IR) / SPIR-V / GLSL (for the compact backend)

Then the driver takes GLSL/DXIL/Metal IR/SPIR-V/etc and produces it's own bytecode. Different copies of LLVM are involved a few different times in different places. It's a complex and frankly fairly crappy pipeline.

jms55··on Programming SDF Animations of Rick and Morty
Pretty much the same. Both Vulkan and WebGL can use GLSL directly (well, GLSL -> SPIR-V for Vulkan). WebGPU technically can't if you run it in a browser, but native WebGPU implementations can take GLSL, you can transpile, and finally you could just write WGSL as it's basically the same as GLSL, just with more Rust-inspired syntax rather than C-inspired.
jms55··on Machine Learning in Production (CMU Course)
Reservoir sampling as in the stuff that's used in ReSTIR for graphics?

It's funny to me where statistics ends up sometimes.

jms55··on The impact of competition and DeepSeek on Nvidia
Great article, thanks for writing it! Really great summary of the current state of the AI industry for someone like me who's outside of it (but tangential, given that I work with GPUs for graphics).

The one thing from the article that sticks out to me is that the author/people are assuming that deepseek needing 1/45th the amount of hardware means that the other 44/45ths large tech companies have invested were wasteful.

Does software not scale to meet hardware? I don't see this as 44/45ths wasted hardware, but as a free increase in the amount of hardware people have. Software needing less hardware means you can run even _more_ software without spending more money, not that you need less hardware, right? (for the top-end, non-embedded use cases).

---

As an aside, the state of the "AI" industry really freaks me out sometimes. Ignoring any sort of short or long term effects on society, jobs, people, etc, just the sheer amount of money and time invested into this one thing is, insane?

Tons of custom processing chips, interconnects, compilers, algorithms, _press releases!_, etc all for one specific field. It's like someone taking the last decade of advances in computers, software, etc, and shoving it in the space of a year. For comparison, Rust 1.0 is 10 years old - I vividly remember the release. And even then it took years to propagate out as a "thing" that people were interested in and invested significant time into. Meanwhile deepseek releases a new model (complete with a customer-facing product name and chat interface, instead of something boring and technical), and in 5 days it's being replicated (to at least some degree) and copied by competitors. Google, Apple, Microsoft, etc are all making custom chips and investing insane amounts of money into different compilers, programming languages, hardware, and research.

It's just, kind of disquieting? Like everyone involved in AI lives in another world operating at breakneck speed, with billions of dollars involved, and the rest of us are just watching from the sidelines. Most of it (LLMs specifically) is no longer exciting to me. It's like, what's the point of spending time on a non-AI related project? We can spend some time writing a nice API and working on a cool feature or making a UI prettier and that's great, and maybe with a good amount of contributors and solid, sustained effort, we can make a cool project that's useful and people enjoy, and earns money to support people if it's commercial. But then for AI, github repos with shiny well-written readmes pop up overnight, tons of text is being written, thought, effort, and billions of dollars get burned or speculated on in an instant on new things, as soon as the next marketing release is posted.

How can the next advancement in graphics, databases, cryptography, etc compete with the sheer amount of societal attention AI receives?

Where does that leave writing software for the rest of us?

jms55··on ROCm Device Support Wishlist
PyTorch and Jax, good to know.

Why do they have ROCm/CUDA backends in the first place though? Why not just Vulkan?

jms55··on ROCm Device Support Wishlist
As someone from the rendering side of GPU stuff, what exactly is the point of ROCm/CUDA? We already have Vulkan and SPIR-V with vendor extensions as a mostly-portable GPU API, what do these APIs do differently?

Furthermore, don't people use PyTorch (and other libraries? I'm not really clear on what ML tooling is like, it feels like there's hundreds of frameworks and I haven't seen any simplified list explaining the differences. I would love a TLDR for this) and not ROCm/CUDA directly anyways? So the main draw can't be ergonomics, at least.

jms55··on Physically Based Rendering: From Theory to Implementation
There is no standard.

Filament is extremely well documented: https://google.github.io/filament/Filament.html

glTF's PBR stuff is also very well documented and aimed at realtime usage: https://www.khronos.org/gltf/pbr/

OpenPBR is a newer, much more expensive BRDF with a reference implementation written in MaterialX that iirc can compile to glsl https://academysoftwarefoundation.github.io/OpenPBR

Pick one of the BRDFs, add IBL https://bruop.github.io/ibl, and you'll get decent results for visualization applications.

jms55··on The Missing Nvidia GPU Glossary
I'm not too informed on the details, but iirc drivers _do_ try and optimize shaders in the background, and then when ready swaps in a better version. But I doubt it does stuff like change threadgroup size, the programmer might assume a certain size and their shader would be broken if changed. Also drivers doing background work means unpredictable performance and stuttering, which developers really don't like.

Someone correct me if I'm wrong, maybe drivers don't do this anymore.

jms55··on The Missing Nvidia GPU Glossary
The weird part of the programming model is that threadblocks don't map 1:1 to warps or SMs. A single threadblock executes on a single SM, but each SM has multiple warps, and the threadblock could be the size of a single warp, or larger than the combined thread count of all warps in the SM.

So, how large do you make your threadblocks to get optimal SM/warp scheduling? Well it "depends" based on resource usage, divergence, etc. Basically run it, profile, switch the threadblock size, profile again, etc. Repeat on every GPU/platform (if you're programming for multiple GPU platforms and not just CUDA, like games do). It's a huge pain, and very sensitive to code changes.

People new to GPU programming ask me "how big do I make the threadblock size?" and I tell them go with 64 or 128 to start, and then profile and adjust as needed.

Two articles on the AMD side of things:

https://gpuopen.com/learn/occupancy-explained

https://gpuopen.com/learn/optimizing-gpu-occupancy-resource-...

jms55··on Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
Raster is believe it or not, not quite the bottleneck. Raster speed definitely _matters_, but it's pretty fast even in software, and the bigger bottleneck is just overall complexity. Nanite is a big pipeline with a lot of different passes, which means lots of dispatches and memory accesses. Same with material shading/resolve after the visbuffer is rendered.

EDIT: The _other_ huge issue with Nanite is overdraw with thin/aggregate geo that 2pass occlusion culling fails to handle well. That's why trees and such perform poorly in Nanite (compared to how good Nanite is for solid opaque geo). There's exciting recent research in this area though! https://mangosister.github.io/scene_agn_site.

jms55··on Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
* MegaGeometry (APIs to allow Nanite-like systems for raytracing) - super awesome, I'm super super excited to add this to my existing Nanite-like system, finally allows RT lighting with high density geometry

* Neural texture stuff - also super exciting, big advancement in rendering, I see this being used a lot (and helps to make up for the meh vram blackwell has)

* Neural material stuff - might be neat, Unreal strata materials will like this, but going to be a while until it gets a good amount of adoption

* Neural shader stuff in general - who knows, we'll see how it pans out

* DLSS upscaling/denoising improvements (all GPUs) - Great! More stable upscaling and denoising is very much welcome

* DLSS framegen and reflex improvements - bleh, ok I guess, reflex especially is going to be very niche

* Hardware itself - lower end a lot cheaper than I expected! Memory bandwidth and VRAM is meh, but the perf itself seems good, newer cores, better SER, good stuff for the most part!

Note that the material/texture/BVH/denoising stuff is all research papers nvidia and others have put out over the last few years, just finally getting production-ized. Neural textures and nanite-like RT is stuff I've been hyped for the past ~2 years.

I'm very tempted to upgrade my 3080 (that I bought used for $600 ~2 years ago) to a 5070 ti.

jms55··on Show HN: I've made a Monte-Carlo raytracer for glTF scenes in WebGPU
This was a recent presentation from SIGGRAPH 2024 that covered using neural nets to store baked (not dynamic!) lighting https://advances.realtimerendering.com/s2024/#neural_light_g....

Even with the fact that it's static lighting, you can already see a ton of the challenges that they faced. In the end they did get a fairly usable solution that improved on their existing baking tools, but it took what seems like months of experimenting without clear linear progress. They could have just as easily stalled out and been stuck with models that didn't work.

And that was just for static lighting, not every realtime dynamic lighting. ML is going to need a lot of advancements before it can predict lighting whole-sale, faster and easier than tracing rays.

On the other hand ML is really really good at replacing all the mediocre handwritten heuristics 3d rendering has. For lighting, denoising low-signal (0.5-1 rays per pixel) lighting is a big area of research[0] since handwritten heuristics tend to struggle with such little amount of data available, along with lighting caches[1] which have to adapt to a wide variety of situations that again make handwritten heuristics struggle.

[0]: https://gpuopen.com/learn/neural_supersampling_and_denoising..., and the references it lists

[1]:https://research.nvidia.com/publication/2021-06_real-time-ne...

jms55··on Show HN: I've made a Monte-Carlo raytracer for glTF scenes in WebGPU
To add to this, DLSS 2 functions exactly the same as a non-ML temporal upscaler does: it blends pixels from the previous frame with pixels from the current frame.

The ML part of DLSS is that the blend weights are determined by a neural net, rather than handwritten heuristics.

DLSS 1 _did_ try and and use neural networks to predict the new (upscaled) pixels outright, which went really poorly for a variety of reasons I don't feel like getting into, hence why they abandoned that approach.

jms55··on Spherical Harmonics
Obligatory useful SH paper for 3d rendering: http://www.ppsloan.org/publications/StupidSH36.pdf

Also lots of other cool research around SH in rendering, e.g. the recent ZH3 paper.

jms55··on Spherical Harmonics
Very very high level explanation that I'm trying to paraphrase from memory, so there's a good chance parts of it are wrong or use the wrong terminology:

A polynomial is a function like `f(x) = Ax^3 + Bx^2 + Cx^1 + Dx^0`

You can approximate most(all?) continuous functions using polynomials. The more "parts" (e.g. Bx^2 is one part) of the polynomial you have, the better you can represent a given function.

The previous example was a 1d function, but polynomials can be over any number of dimensions. E.g. a 3d polynomial `f(x, y, z)`.

Spherical harmonics are just a form of 3d polynomials, but with some special "parts" (called a basis function), with the amount of parts you have called a "band".

As for what they're good for, they're a fairly compact way of representing and filtering 3d signals.

In 3d rendering, they're really good at storing light hitting a point from different directions. You have an incoming ray of light on a unit sphere/hemisphere with origin x, y, z going towards the given point. You can then take your list of light rays with various (x,y,z) origins and (r,g,b) intensities, and then form a spherical harmonics approximation over it (basically a fitted 3d polynomial), which can be stored as just a few coefficients (A, B, C, D, etc...) using 1 or 2 bands, and cheaply computed (querying the light value r,g,b for a given ray) by plugging in the ray's x, y, z into the formula.

Besides being cheap to store and query, because you're only using 1-2 bands, you only capture the "low frequency" of the lighting signal, e.g. large changes in the light value get dropped, since it's just an approximation of the original signal. While this is normally bad, for 3d rendering, it's free denoising, giving you a smoother output image!

jms55··on Intel announces Arc B-series "Battlemage" discrete graphics with Linux support
> Isn't it insane to think that rendering triangles for the visuals in games has gotten so demanding that we need an artificially intelligent system embedded in our graphics cards to paint pixels that look like high definition geometry?

That's not _quite_ how temporal upscaling work in practice. It's more of a blend between existing pixels, not generating entire pixels from scratch.

The technique has existed since before ML upscalers became common. It's just turned out that ML is really good at determining how much to blend by each frame, compared to hand written and tweaked per-game heuristics.

---

For some history, DLSS 1 _did_ try and generate pixels entirely from scratch each frame. Needless to say, the quality was crap, and that was after a very expensive and time consuming process to train the model for each individual game (and forget about using it as you develop the game; imagine having to retrain the AI model as you implement the graphics).

DLSS 2 moved to having the model predict blend weights fed into an existing TAAU pipeline, which is much more generalizable and has way better quality.

jms55··on Bevy 0.15 released
Ah yeah pipeline compilation is a very non-trivial problem[1]. Unreal is also struggling with this a lot. I hear you that it's an issue.

The Bevy issue you want to follow is https://github.com/bevyengine/bevy/issues/10871.

[1]: https://therealmjp.github.io/posts/shader-permutations-part1 + https://therealmjp.github.io/posts/shader-permutations-part2

jms55··on Bevy 0.15 released
You're not at all wrong, but like you said it's the (sometimes unfortunate) reality of open source. We have a lot of rendering contributors, but less so for every other area except probably the core ECS.

Here's to hoping we get more boring contributions in the future!

What is pipeline events though? I haven't heard of that, and google isn't bringing anything up.

P.S. I'm actually the author of the virtual geometry feature, so that was a particularly well chosen example :)

jms55··on Bevy 0.15 released
I mean, kind of. It's a bit of a vague comment. Dynamic dispatch is not bad by itself, and as not-an-ECS-developer, I couldn't actually answer how much it is or is not used in the ECS internals. Probably more than I expect, but not enough to be an issue. Performance is plenty fast, we profile that pretty often.

As for things that can't be caught by the type system, it's not like we can enforce that the logic in system_b that expects a given entity to be alive won't accidentally break if you introduce some system_a later on that despawns the entity at some point. We can't compile-time validate that kind of safety for game logic - it's just not possible.

On the other hand there are things we can do that C++ based engines can't thanks to Rust. Compile-time variable mutability means we can know exactly what set of systems modify which pieces of data, which means we can automatically schedule different systems in parallel without any safety issues, or detect what when and who touched a piece of data last.

In terms of features, Bevy is still early in development. Feature parity with existing, 10-20 year engines with many paid developers is going to take a while. But I think we have some pretty compelling features on our own even today, foremost our focus on ECS and modularity.

In terms of what Rust brings, it has the performance of C++, but with better memory safety (mostly in terms of engine internals - not like user code touches lifetimes that often), thread safety, a way better and standard build system, some modern features like pattern matching and enums, etc. The usual ways in which Rust is better than C++.

The answer isn't so much "why would I want to use an engine in Rust" (although I personally love Rust and would absolutely choose it over a C++) but more "we want to make a new engine". You want C++ levels of performance, but without having to use C++, so Rust is an obviously compelling choice for a game engine.

jms55··on Bevy 0.15 released
> My impression of Bevy hasn’t always been the greatest for a number of reasons but I could possibly be convinced to see the light on this one!

As one of the Bevy contributors, I'd love to hear what you didn't like in the past, and what you liked from this release. User feedback is super useful - I don't see the same set of pain points most users see since I work on the engine internals.

In general building a game engine is an _extremely_ large task - we're under no illusions that we're going to get it perfect the first time, or even the first few rewrites. But I'm fairly confident in the long term direction of the project, and the community we built and amount of developers contributing means I'm confident that we'll get there eventually :)

jms55··on What's Next for WebGPU
I don't disagree, but good luck getting vendors to standardize on anything...
jms55··on What's Next for WebGPU
> But WGPU has other contributors with other priorities. For example, WGPU just merged some additions to its nascent ray tracing support. That's not a Mozilla priority, but WGPU took the PR. Similarly for some recent extensions to 64-bit atomics (which I think is used by Bevy for Nanite-like techniques?), and other areas.

Yep! The 64-bit atomic stuff let me implement software rasterization for our Nanite-like renderer - it was a huge win. Same for raytracing, I'm using it to develop a RT DI/GI solution for Bevy. Both were really exciting additions.

The question of how performant and featureful wgpu is is mostly just a matter of resources in my view. Like with Bevy, it's up to contributors. The unfortunate reality is that if I'm busy working on Bevy, I don't have any time for wgpu. So I'm thankful for the people who _do_ put in time to wgpu, so that I can continue to improve Bevy.

jms55··on What's Next for WebGPU
I don't necessarily disagree. But I don't agree either. WebGPU has given us as many positives as it has negatives. A lot of our user base is not on modern hardware, as much as other users are.

Part of the challenge of making a general purpose engine is that we can't make choices that specialize to a use case like that. We need to support all the backends, all the rendering features, all the tradeoffs, so that our users don't have to. It's a hard challenge.

jms55··on What's Next for WebGPU
We do when there available, but I think the way browsers implement limit bucketing (to combat fingerprinting) means that some users ran into the limit.

I never personally ran into the issue, but I know it's a problem our users have had.

jms55··on What's Next for WebGPU
It's partly because WebGPU has very conservative default texture limits so that they can support old mobile devices, and partly it's a problem for engines that may have a bunch of different bindings and have increasingly hacky workarounds to compile different variants with only the enabled features so that you don't blow past texture limits.

For an idea of bevy's default view and PBR material bindings, see:

* https://github.com/bevyengine/bevy/blob/main/crates/bevy_pbr...

* https://github.com/bevyengine/bevy/blob/main/crates/bevy_pbr...

jms55··on What's Next for WebGPU
Timestamp queries will give you essentially time spans you can use for profiling, but anything more than that and you really want to use a dedicated tool from your vendor like NSight, RGP, IGA, XCode, or PIX.

> Right now I feel like the only way to write efficient WebGPU code is to deeply understand specific GPU architectures. I hope some day there's a dev tools tab that shows me I'm spending too much time sampling a texture or there's a lot of contention on my atomic add.

It's kind of the nature of the beast. Something that's cheap on one GPU might be more expensive on another, or might be fine because you can hide the latency even if it's slow, or the CPU overhead negates any GPU wins, etc. The APIs that give you the data for what you're asking are also vendor-specific.

jms55··on What's Next for WebGPU
> This is one reason we don't see high-performance games written in Rust.

Rendering is _hard_, and Rust is an uncommon toolchain in the gamedev industry. I don't think wgpu has much to do with it. Vulkan via ash and DirectX12 via windows-rs are both great options in Rust.

> After four years of development, WGPU performance has gone down, not up. When it dropped 21% recently and I pointed that out, some people were very annoyed.[1]

Performance isn't most of the wgpu maintainer's (who are paid by Mozilla) priority at the moment. Fixing bugs and implementing missing features so that they can ship WebGPU support in Firefox is more important. The other maintainers are volunteers with no obligation besides finding it enjoyable to work on. Performance can always be improved later, but getting working WebGPU support to users so that websites can start targeting it is crucial. The annoyance is that you were rude about it.

> Google pushing bindless forward might help get this unstuck. Although notice that the target date on their whiteboard is December 2026.

The bindless stuff is basically "developers requested it a ton when we asked for feedback on features they wanted (I was one of those people who gave them feedback), and we had some draft proposals from (iirc) 1-2 different people". It's wanted, but there are still major questions to answer. It's not like this is a set thing they've been developing and are preparing to release. All the features listed are just feedback from users and discussion that took place at the WebGPU face to face recently.

jms55··on What's Next for WebGPU
> The issue is, it's likely that a company with $2 BILLION spent on product development and a very deep relationship with Apple, like Unity, will have success using WebGPU the way it is intended, and nobody else will.

Not really. Bevy https://bevyengine.org uses WebGPU exclusively, and we have unfortunately little funding - definitely not $2 billion. A lot of the stuff proposed in the article (bindless, 64-bit atomics, etc) is stuff we (and others) proposed :)

If anything, WebGPU the spec could really use _more_ funding and developer time from experienced graphics developers.

jms55··on What's Next for WebGPU
Bindless is pretty much _the_ most important feature we need in WebGPU. Other stuff can be worked around to varying degrees of success, but lack of bindless makes our state changes extremely frequent, which heavily kills performance with how expensive WebGPU makes changing state. The default texture limits without bindless are also way too small for serious applications - just implementing the glTF PBR spec + extensions will blow past them.

I'm really looking forward to getting bindless later down the road, although I expect it to take quite a while.

By the same token, I'm quite surprised that effort is being put into a compatibility mode, when WebGPU is already too old and limiting for a lot of people, and when WebGL(2) is going to have to be maintained by browsers anyways.

jms55··on Unreal 5.5 is a big deal [video]
Just realized my comment's formatting got completely messed up. I can't seem to fix it, but hope it's readable enough.

> With megalights, could you for each pixel, pick multiple random lights to increase accuracy?

You could, and megalights probably does. That's what I meant by 0.5-2 rays. Each ray you pick a different light source.

The problem is that you usually needs thousands of rays for good results in a single frame, and even a single ray more is very expensive.

Hence the reason people go as low as 0.5 per pixel (i.e. rendering at half resolution and then upscaling), and rely so much on a spatiotemporal denoiser, among other techniques.

← PreviousPage 2 of 6Next →