11,952 karma · joined October 4, 2008
In the GeForce3 era, the registers got complicated enough they resembled tiny "pixel shaders", but under the hood it was still a small struct held in registers. Vertex shaders were 1 to 128 asm instructions executed strictly linearly.
In the G80 era we got "general purpose shaders." But, they still depend heavily on the fix function pipeline for their dispatch/scheduling and I/O.
These days, everything is basically a dressed-up compute shader. The shared-memory SRAM is front-and-center in your attention. Dispatch and scheduling are manual and complicated. GPUs are transitioning into tensor evaluators.
So, maybe today we can start considering talking about planning committee meetings about stability. But, what I've observed is that this has been a request for a few decades now. And, in hindsight it would not have worked out in the past. Moving forward, maybe it would work out OK today for a while. But, I don't see the rate of change in GPUs slowing down any time soon. Wouldn't be surprised if we're racing towards some Cerebras + Tensor Cores + FPGA near future.
https://news.ycombinator.com/item?id=49736660
https://www.reddit.com/r/LocalLLaMA/comments/1wjieap/made_th...
Papers: https://arxiv.org/abs/2503.23303 https://arxiv.org/abs/2510.01237
Model: https://huggingface.co/DeepMostInnovations/sales-conversion-...
Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sal...
I can't compare it to changes I've seen in CPU architecture since then. Maybe like: Compare the NES with its 6502 and per-cartridge mappers vs. a IBM 386 PC. Now repeat that shift 2 or 3 more times.
A flood of VR mods for popular games have been coming out lately.
That can't recognize what you say. But, it can ID where you are in a specific show out of zillions of hours of shows.
So, draw a curve from the ground to the sky. Draw some explosion shape. Change pen color to the sky color. Draw the curve and the explosion again to erase them. Change pen color and start the next firework.
Me. I get why people like small phones. But, I've have the iPhone Jumbo Huge edition for a long time now. I use my phone as a pocket tablet 10X more than as a phone.
Looking at the Duo specs https://www.apple.com/iphone-duo/specs/ vs my current 12 Max Pro specs https://support.apple.com/en-us/111874 it looks like the differences I care about are
1) 1878 pixels over 4.64" vs. 1284 over 3.07" 2) Default camera mode is 48 MP vs 12. Other modes are surprisingly similar. 3) A much faster CPU
I'm sure there are 1,000 more small differences. But, my 12 Jumbo Huge from late 2020 still runs great. If I were to upgrade, I'd get the Duo. We'll see how much longer I keep waiting, I guess.
I don't know when that ability faded. I think it was early because didn't really notice until people started talking about it recently. I can paint and draw. I can think about how things look. But, I don't see them. If I'm on the edge between awake and asleep I can force visualization as a lightweight lucid dreaming.
When I code, I think about data structures and algorithms using a mental proprioception. Putting things in space, moving them around. But, I have to remember where they are. Like playing the shell game with your eyes closed.
Best I can do is link you to https://x.com/RadianceFields and https://radiancefields.com/ They have all the news about GS every day.
On the GPU however, the hyperthreads are just a round-robin execution queue to take advantage of instruction pipelining. The register bank of a single GPU core is huge and can be flexibly divided across a variable number of thread contexts when a kernel is launched. Many thread contexts can be held in registers simultaneously in a single GPU core. That makes stalling on memory latency much less of a problem. The hardware can focus on delivering raw bandwidth with high latency and get great overall performance. This throughput-instead-of-latency trade-off extends to many other aspects of GPU design.
The promotional material likes to label the individual lanes as “cores” because it sounds more impressive. And, it’s not entirely incorrect.
Even the dev docs use the marketing terminology. The description I gave above needs a bit of piecing together.
Instead, the "modern" APIs allow you to leverage your pre-existing knowledge of "allocate arrays of structs and start indexing them." That's not trivial. But, it is already familiar. And, it ends up a lot better in the end compared to "Invoke whole lot of functions to manipulate a hidden state machine."
You can find a little info on Modern OpenGL at https://github.com/fendevel/Guide-to-Modern-OpenGL-Functions, https://juandiegomontoya.github.io/modern_opengl.html, https://ktstephano.github.io/, https://patrick-is.cool/posts/2025/on-vaos/
If you are going to use Vulkan, check out https://howtovulkan.com/ Vulkan started with a lot of compromises to make mobile hardware happy at the expense of making everything overly complicated on desktop. Over the past decade, desktop devs have managed to get a lot of features added to make Vulkan on desktop more sane. How To Vulkan covers that newer approach. This recent video "It's Not About the API" https://www.youtube.com/watch?v=7bSzp-QildA shows how simple it can be if you let it.
And, if you are on a Mac, folks who use Metal like it a lot. Don't worry about lock-in. Once you learn the basics, knowledge is easily transferable to DX12 and Vulkan. You should plan to write 3 or 4 renderers to throw away anyway :P
As @petecordell put it: "Telling a programmer there's already a library to do X is like telling a songwriter there's already a song about love."
Sometimes I'm not interested in the inner workings of a technology, I just want a one-off tool. Ex: No interest in web tech. But, vibe-coded a tool to scrape and cross-reference a couple of online reports. Now I have the info I need for other projects.
Sometimes I know exactly what I want and it's quicker to hand-hold the AI through making it for me. Ex: Put together yet another SIMD math lib recently. SIMD intrisics are an obnoxious API. But, a non-SIMD implementation is easy to verify and many SIMD implementations are easy to validate vs. the non-SIMD reference. It's just a huge amount of obtuse code.
Sometimes I don't know exactly what I want and it's great to have a always-online partner to bounce questions off of, do rapid-prototyping for me, do code reviews for me. AIs aren't perfect. But, being instant, patient and often right makes them a huge improvement over online forums.
Lots of people start out with a non-programmable config format. But, as their situation becomes more complicated, they end up shoehorning in programmable-ish features until they realize they are running straight into https://en.wikipedia.org/wiki/Greenspun%27s_tenth_rule and decide to do it properly.
https://www.uni-konstanz.de/mmsp/pubsys/publishedFiles/LuCoD...