Neural processing is ubiquitous in phones (all the flagships do neural enhancement on the image, and iirc apple does gesture recognition neurally), and all Ice Lake and newer
laptop chips (plus lol Rocket Lake), as well as all Zen4 chips have neural instructions too. NVIDIA's had it for 2 gens now. It's pretty well across the stack at this point. And it's getting used for different things in different places. Things like face recognition or gesture recognition are natural fits in a low-power environment, reduces power consumption of those features hugely. PCs and servers can do content generation or some other larger "inference" tasks.
Neural game upscaling is a huge win too. The neural-TAAU upscalers (DLSS 2.x and XeSS) perform a lot better than FSR still, even compared to FSR 2.0/2.1. And there will be other things they can figure out how to ML-accelerate as adoption continues, I'm sure. It enables some solutions to hard problems that don't have good deterministic algorithms, and you can run a lot bigger models than you can without acceleration.
also fabrice bellard (of course) wrote a neural compressor... https://bellard.org/nncp/
it's also likely going to be useful in level design and asset creation as well, although of course as we're seeing now with Stable Diffusion there's some interesting legal questions with content creation.
optical flow engine (not neural) is another big win for computer vision stuff, too, offloads a few tough tasks to hardware, like object tracking and motion estimation. hooking into that with zoneminder or something would be awesome.
I'm also stoked for shader execution reordering too. This is similar to what Intel calls "ray binning", it basically is a best-effort "re-alignment" of the thread (state,etc) to the most similar execution group for memory read alignment/coalescing. I think one or another raytracing implementation (Intel or NVIDIA) they were coalescing based on material.
Intel talks about theirs here: https://www.youtube.com/watch?v=SA1yvWs3lHU
Ada whitepaper: https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvid...
I am interested to hear what AMD is doing with RDNA3, supposedly there is a new RDNA3 ISA instruction for matrix acceleration but rumors are it's not a full high-performance matrix unit like CDNA and >=Turing? I don't know why you'd add an instruction without some hardware acceleration though. So maybe less than the other implementations... which is a little disappointing.