In what way has Nvidia “forgotten” gamers with the rise of their datacenter business?
It used to be that $200 would get you a low-range GPU, $400 a midrange, and $600 a pretty darned good one. The 1070 launched at $379 (or $480 in 2023 dollars), the 2070 at $499, the 3070 at $499, then all of a sudden the 4070 now is $599. That's a big difference, even accounting for inflation.
The 1080 Ti was $699, an already-unfathomable price back then that cost more than a whole console setup. Today? The 4080 is $1200. That's often more than the rest of the system put together.
On the other hand now that it’s stopped, so has raw performance scaling. There was a few generations of cleanup, but you can’t squeeze blood from a stone forever. Maxwell arguably cut too far and pascal had to start adding functionality back. Gpus are like e-cores, large powerful cores mean you get fewer of them and that often works out to a lower PPA. There aren’t many opportunities for cool tricks and the model doesn’t favor using lots of area. The coding is written for extreme parallelism already, so, that’s not a problem, it’s inherent to the platform.
Transistor per $ growth hasn’t completely stopped but it’s certainly nowhere near what it was 10 years ago, even if you factor in things like packaging/stacking the total wafer area still runs up a big bill that offset most of the density gains. And that’s what we’ve been seeing over the last 10 years in gpus too. 4070 is about the same area and cutdown as GTX 1070 but wafers cost like 8x as much.
What you have to do in this operating regime is find ways to get more performance per transistor - and that’s exactly what Jensen made a big bet on 5 years ago with dlss. 7% more transistors that with DLSS 2.5 give 30% speedup at native quality and 50-70% speedup at iso-quality with FSR2 Quality mode.
That’s what the future looks like - rewriting your applications to take advantage of new accelerators that provide large speedups. And the rewriting is very minimal for any application that uses TAAU already. It sucks but if cost/transistor is not going to come down you have to get more out of the transistors you have. Work smarter not harder.
Yeah the upscaling is nice, but I dunno about $1200 nice...
It was cool when they instead focused on things like creating mobile smaller versions of the cards with laptop level power draws, or better cooling systems that weren't as noisy etc. While staying within reason price points. Just the mundane stuff, not necessarily real time raytracing (gimmicky and not very noticeable IMO, even as someone who uses 4080s on Geforce Now). Something like the Steam Deck package with its integrated console-like APU is much more consumer friendly, but also much much weaker than a 4090 I guess. Different priorities that Nvidia might've reconsidered if not for the AI supersampling stuff also bleeding into crypto and ML, justifying their bet.
Edit: Yeah their AI profits are skyrocketing, while gaming is a has-been. Good bet for the company, bad omens for gamers :/
https://www.pcgamer.com/nvidias-record-breaking-profits-are-...
Because moore's law wasn't just about transistor count but about the economic impact of exponential growth in transistors-per-$. In a world without moore's law, using more transistors will result in a higher-cost product. If you want to hold product cost fixed, or even contain the cost spiral, you need to do more with the same amount of transistors - performance-per-transistor is the metric that matters now.
AMD and NVIDIA have already stripped down their pure-raster implementation as far as they can go, with RDNA1 and Maxwell respectively. Maxwell actually cut too far (software scheduling, "minimal" DX12 support, etc) honestly. So where do you keep making perf/tr gains after that?
The gaming world has already pretty well settled on TAA (although some people will never accept it) and upscaling is already common in the console world. So, do TAA upscaling better such that you get the performance gains but not the reduction in visual quality that usually comes with it.
Tensor makes up a relatively small amount of die area (5.9% of total Turing die area, based on comparisons between Turing Major/RTX and Turing Minor/GTX SM engine die shots). And that gets you to about 30% faster than FSR2 for a given level of visual output quality. So the perf-per-transistor metric increases. Also, unlike a fixed-function accelerator, it can be used for all kinds of other stuff too. It's basically a whole programmable sub-processor, an accelerator for your accelerator.
https://www.reddit.com/r/hardware/comments/baajes/rtx_adds_1...
Now, why ML as opposed to just running it on shaders? Same logic as adding an AVX unit, math density is a lot higher and it can do a lot of work for applications that are specifically tailored to it. DLSS2 uses a relatively standard TAAU (similar to FSR2) but determines the weighting of the samples using a neural net. This produces a lot higher quality than a procedural algorithm currently can - especially under "bad conditions" like higher degrees of upscaling, low framerate/limited sample count, or temporally unstable/high-temporal-frequency areas of the image.
http://behindthepixels.io/assets/files/DLSS2.0.pdf
https://raw.githubusercontent.com/NVIDIA/DLSS/main/doc/DLSS_...
FSR2 does ok at 4K quality mode, but at 1440p and (especially) 1080p output resolutions and in performance modes it does much worse. FSR2 quality 1080p is more like DLSS2 performance mode or maybe balanced mode, so NVIDIA gets more speedup at a given level of visual quality. And DLAA can produce a better-than-native image when running with a native input quality.
The neural weighting just is a lot more efficient at using its samples, it understands what is going on in the scene (moving edges/occlusion etc) and can extract a higher signal-to-noise ratio from the samples and the textures. It's like an op-amp, the ratio of input:output pixels is the "gain factor", and FSR2 and other traditional TAAU algorithms are simply noisier at any given level of gain, whether that's unity or extreme gain, and have other edge-cases like turn-on threshold (bad performance with low samples). ML is the "schottky diode" of graphics amplification (dangerously mixed metaphor, lol), it's simply a lot more agile at shaping the signal than what came before.
(and while on paper plenty of people have argued that procedural programs should be able to do anything ML can, it's not like AMD and others haven't tried to improve TAAU with FSR2, and many others before them. DLSS2 is better, just like LLMs and Stable Diffusion are a lot better than procedural algorithms in their own niches.)
--
All of this exists completely orthogonally to actual wafer costs or packaging or other things. Packaging may boost that transistors/$ metric a little bit but that just gives you a little more to play with. It does allow you to make chips with twice the transistors at twice the cost and have them yield at high rates, but fundamentally 2x400mm2 is still 800mm2 of silicon even if you yield at 100% - you're using more wafer, which drives up costs. Wafer costs have been increasing nearly as fast as density (and predicted to match/pass at 3nm) but there has been a small gain in tr/$, certainly nowhere near the rate of moore's law days. But if wafers cost 8x what they did for 28nm, and are continuing to increase at ~50% per generation, and you keep using more wafer area to compensate for slowing shrinks, then costs will go up (even more than they have).
There is no direct link between "running deep learning" and costs going up. That is just happening independently and affects AMD too, even when they didn't go in on deep learning. TSMC prices keep going up (as do their margins, even now) and even if they went to zero, the design+validation costs are still going up too. A lot of these costs are driven by hard physics problems and not just TSMC profit margin (although it doesn't help).
--
> It was cool when they instead focused on things like creating mobile smaller versions of the cards with laptop level power draws, or better cooling systems that weren't as noisy etc.
I think a 30% performance boost at native visual quality, without power increase is pretty cool. Don't laptops benefit from having 30% higher perf/w just from turning on a setting? And Ada itself is a ~60% perf/w increase over previous generations too.
Like Ada is one of the most efficiency-focused generations ever, much moreso than Ampere or Turing with their trailing nodes. DLSS just stacks on top of this - and unlike FSR2, NVIDIA doesn't fall apart at 1080p resolutions that laptops tend to be using.
Cost is higher than people want, but on the other hand (a) that's going to be the reality unless there is a breakthrough in transistors-per-$, you can't make a fixed number of transistors infinitely fast, there is some asymptotic limit. And (b) people are cherrypicking favored examples or comparing against trailing-node products that had larger dies on slower, less energy efficient nodes to keep costs down.
GTX 970 at $329 was an outlier and the lowest x70 product of all time, on a trailing node (28nm again after 20nm fell through). GTX 670 launched at $399 for a similarly sized die over 10 years ago. GTX 1070 launched at $449 7 years ago with another similarly-sized (~300mm2) die. Turing, Ampere, and Maxwell were all abnormally cheap due and large due to the trailing-node but you paid for this with worse efficiency. There has never been a x70 product launched at $299 and you are welcome to check this!
First x70 product: https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
And yes I think the consensus is that nodes like 8nm probably are "good enough" especially if they are much cheaper, especially at the low end where PHY size is becoming a problem. PHYs don't shrink, so you can't scale a product arbitrarily small - the logic may shrink by 70% but those PHYs are just as big as ever. So there is a de-facto "minimum die size" that is ever worth producing, because there is a fixed PHY area that you simply cannot eliminate. And in a world where wafer costs are going up, that area costs more and more every generation.
A 3060 Ti 16GB wouldn't even need clamshell (it has 8 PHYs, 8x2GB per module=16GB) and could probably have hit $299 or $329 launch cost, if NVIDIA had gone down that road. And it's Good Enough for 1080p, and avoids some weird compromises that shake out of the need to trim PHY area.
AMD already did exactly this with the 7600 - which is held back on 6nm (N7 family) rather than using N5P (N5 family) like the rest of the RDNA2 lineup. Why? Cost.
GPUs will be bought like TVs. It took 4060, about 6 years to double the Passmark score over 1060.
BUT that comes at a price - Nvidia consumer chips are also notoriously expensive, but if you want best-of-breed for gaming, it does come at a price.
I am hoping that AMD and Intel will be able to compete with Nvidia someday but I'm not holding my breath.
The hoarding isn't going to stop unless there are either efficient alternatives that are competitive on performance and price.
Perhaps that is why I keep seeing gamers crying over GPU prices and unable to find cheap Nvidia cards due to the AI bros hoarding them for their 'deep learning' pet projects.
So they settle with AMD instead.
Supply issues should be gone too; I got a 4070 Ti shortly after launch, no problem.
Thus only gamers that care about RT and only in the games that make a good use of it (virtually none) have any serious benefit.
So raw performance has stagnated and gains have been concentrated in things like DLSS 2.5 that let you get native quality at a 30% speedup or better than native DLAA at 0% speedup, or FSR2 Quality level quality at 50-70% speedup. Cause that gets you more performance out of a small increase of transistors/cost.
My main gripe is that at 4k resolution, top of the line GPUs shouldn't be using AI frame scaling to get decent fps unless you are taking the raytracing penalty for funsies.
Seems like you’re really dismissing the massive speed ups these past few years. Agreed that ray tracing in games is only at the beginning. A lot of that is gated by the consoles/AMD but that’s generally how it goes. Would love to see Nvidia in one of the powerful consoles to accelerate adoption of these technologies.
My fear is that Nvidia, seeing record-high profits for their datacenter cards, will gradually phase out budget/midrange gaming cards in favor of the higher end stuff. Maybe AMD and Intel will step up their game to tackle that segment, who knows. Or maybe Apple Silicon gaming will gradually take off. But if neither of those happen, there might just be a huge void left in the "affordable PC gaming" segment (which used to be most of it, not so long ago).
On the other hand... I do have to give them credit and say that GeForce Now is AMAZING -- limited library aside. If they can keep growing that segment, hell, maybe the future of PC gaming is just in the cloud, like everything else, and nobody would want to buy expensive GPUs for home use that just become obsolete in a couple years anyway.
But aren't most hardware producers these days fab-limited, with Apple, Samsung, Nvidia, etc. all competing for the same few foundries, especially at the smaller processes?
If they only have X amount of production available a year, it seems like they'd want to focus those on the super-high-margin high-end datacenter cards rather than the leftover low-end stuff. Or maybe the high/low stuff are manufactured on different pipelines altogether, that don't suffer from the same production limits? I'm not sure how that works...
I no longer buy nvidia hardware but I do enjoy stock price getting higher. I just wish I had the sense to buy more, a lot more, stock when it was much cheaper. How does a chicken shit like me make big money :(
That’s a product designed and manufactured by AMD, eight?
The gamers paved the way!