New GeForce RTX 4080 Super, RTX 4070 Ti Super, RTX 4070 Super
nvidia.com
nvidia.com
It's weird that they're only comparing the new cards to the RTX 30's and 20's, and not the "v1" 40's. I assume the 4080 SUPER is faster than the 4080 (based on name?) but it seems cheaper and there's absolutely no comparison data
Please send donations to the Aboveground Miners fund in your choice of shitcoin.
What, precisely, is the point of making "actually you mean crypto, not bitcoin" posts, other than demonstrating that you are, indeed, "very smart"? Like, this person doesn't even exist, it's just a "heh aren't those no-coiners dumb" strawperson that you imagine to be some big dummy.
This article https://beebom.com/nvidia-rtx-4080-4070-ti-super-gpu-specs-r... says "the RTX 4070 Ti Super is up to 10% faster than the non-Super 4070 Ti on average"
So, on average, up to 10% faster, yeah seems pretty incremental.
While the TFLOPS of the Super variant does only see a ~10% increase as you note, memory bandwidth jumps by 42% and the memory capacity jumps by 33%, while the launch price is the same in my currency.
It basically bridges half the distance between a 4070 Ti (non-Super) and a 4080 (non-Super) for the same launch price as a 4070 Ti (non-Super).
Great card for memory intensive workloads like LLM inference with big context windows, IMO.
EDIT1: 4070 Ti Super TDP is 320W (same as 4080), higher than the anticipated 285W
EDIT2: launch price confirmed to be same as the 4070 Ti (non-Super), lower than anticipated!
Ergo, there's a decent chance it won't sell for MSRP.
I got a 4090 a few months ago before the prices increased, and I'm beyond stoked with the performance for (typically triple qhd simulation) gaming. It's just a beast.
I have a 2nd PC I'd like to upgrade too though, and the 4070 TI looks like it would be fantastic in this.
But yes, very weird that none of the comparions actually show them!.
Not everyone is buying the best and latest. Plenty of people wait for previous gens to drop in prices or enter the secondary market.
You're not getting what I'm saying - people stay in the same tier when they upgrade. They might do every generation, every other generation, skip every two generations, but the point is that people who have xx70 (or xx80) will buy xx70 (or xx80) from a newer generation.
10xx to 20xx upgrade made little sense to most gamers because RTX was a thing you turn on, look at pretty reflections and turn off to regain the performance. 10xx generation was a weird generation for NVIDIA and doubt they would make such a consumer friendly generation ever again.
Ive a lot of AMd Nvidia machines - two high-end gaming machines.. the naming conventions of Nvidia cards are just odd to me and I can tell what anything actually means..
Most people upgrade from a 10, 20 or less so 30 series card.
They are selling upgrades to older generations.
At least 4070 Ti and 4080 have become completely obsolete when their Super variants are much better, and in the case of 4080 Super, even cheaper too.
I suppose that they have stopped producing the non-Super variants, as nobody would want those where the Super cards are available.
I am still using a 2060 Super from 2019, and the same has happened in that year when the RTX 2000 Super series has replaced the previous RTX 2000 series.
Card A is old, and card B is new.
The availability of original and super cards that compete in the same price point will be very short, and most people won’t really have that choice.
EDIT: For training
AMD and Intel GPUs do not have the software ecosystem for AI workloads that Nvidia does, though AMD is rapidly improving. Nvidia has had an effective monopoly on the AI hardware space for the last year or so, and continues to have an effective near-monopoly, but that won't last forever as AMD and Intel catch up.
The VRAM is one of the largest differentiators of their cards. Sufficient VRAM allows you to run huge LLMs like 65B in-memory, which is orders of magnitudes faster than system RAM + CPU. Smaller amounts of VRAM require swapping between VRAM and system RAM and incur a major performance penalty.
Businesses are fighting to fork over $50k+/card for 40/80GB cards with the same processor as the 24GB consumer cards - it doesn't make economic sense for Nvidia to offer more on the consumer cards, lest they start cannibalizing demand for the enterprise cards.
ADA 6000 RTX (48GB) — an enterprise (workstation) cars — is about $10K sticker price. The differentiator between those and data center cards in the 40GB-80GB range is, obviously, not just VRAM.
The RTX ADA 6000 actually slightly performs worse than a consumer-grade 4090 (~10% less performant), even though it retails for 4-5x more ($8000 for ADA 6000 compared to $1500-2000 for a 4090).
Also, as samspenc points out, that RTX 6000 is using the same AD102 chip found in a consumer RTX 4090, just with marginally more CUDA cores, TMUs, ROPs, etc.
The most substantial difference between the two is that the $10k card has twice the VRAM of the $2k card.
I know it sounds outrageously oversimplified, but Nvidia can indeed more or less print money by attaching a few extra memory chips to what is otherwise a flagship consumer-oriented graphics processor.
The profit margins on the H100, for reference, are estimated to be around one thousand percent (1000%), i.e., they sell for ~10x as much as it costs Nvidia to make them, and demand for them still grossly exceeds the total supply. See: https://the-decoder.com/nvidias-h100-gpu-sells-like-hot-cake...
The xx90 cards (3090, 4090, etc) have always been aimed at the gamer with a lot of cash. But games aren't really designed with that hardware in mind, so you won't see games that will take advantage of that much VRAM, so NVIDIA isn't inclined to increase the memory on them.
I have a 3080 right now, so haven't really been able to play around with training a model, but I'd like to see the 5090 have 32 GB, but I'm having my doubts.
The real reason is that it doesn't deteriorate with regards to the input length in case of text, or far neighbourhood in case of vision. It's just a universal, new, building block that allows for shallower neural networks to perform more like their bigger versions
I suppose I could be getting a biased impression though, as of course many more people are in a position to recommend the more accessible models.
What sort of things are you running that take full advantage of that 24GB?
> If you’re interested in ML training
Training - at least the one I tried - requires to be run in fp16 mode. So a 7b net needs 14 GB for the model weights alone, plus some extra for the context and the stuff I don't really understand (some gradient values, oh that makes sense now that I've written it)
https://nonint.com/2022/05/30/my-deep-learning-rig/
and land a job at OpenAI as a side effect. Of all places, I wasn't expecting such pushback here
NVidia no longer supports this for the 40-series, I think this is because they want anyone interested in using their GPUs for LLMs to buy the pricier models with more VRAM.
(Which is of course how CUDA built its success more generally, vs the "you have to buy the $5k workstation card to get started" strategy from ROCm.)
More generally you'd call this optimization and targeting the hardware that's available. No sense releasing crysis when everyone is running a commodore 64, after all.
Having said that, I've trained/finetuned image models just fine on an RTX 2070 Super with 8 GB of VRAM. This was back when doing so was more fruitful than simply training a more robust model in the first place. Given that is the current status quo - I'm curious what sort of training you're doing whole-network that actually produces results that are noticeably better than doing something few-shot during inference or doing LoRA finetuning? The latter brings you back into the realm of tuning on low-VRAM configs.
In general, a single GPU's memory constraints are one of many when training a model _from scratch_. In that case, you're bottlenecked by data and data parallelism. You don't need one or a few GPU's, you need more than would fit in a consumer setup in the first place.
Give the consumers / gamers a consumer-priced GPU with a max of 16-24 GB VRAM for the high-end models. By consumer-priced, I mean $500-2000.
And make anyone interested in AI / ML / LLM / 3D / creatives pay $3000-10000 for GPUs that are similar in performance but have much higher VRAM.
Then top it out with six-figure (or higher) priced GPUs for the FAANG companies which can afford them for their data centers and currently contribute the most revenue (and profit) to NVidia.
It costs $2,000 and might get some people someplace interesting.
[1]: https://developer.nvidia.com/embedded/learn/getting-started-...
Here's a direct amazon link: https://www.amazon.com/dp/B0BYGB3WV4
And a running demo: https://forums.developer.nvidia.com/t/llama-2-llms-w-nvidia-...
For running sparsified/quantized llama2 it might be good, not sure about for fine tuning. I didn't see any FP16 numbers.
https://ir.amd.com/news-events/press-releases/detail/1176/am...
IMO this pretty well displaces the 4060 8GB and 16GB - it's cheaper (than even the already below-MSRP street prices) on 4060 Ti 8GB, it's way cheaper than the 16GB model. $50 over 7600 MSRP for twice the VRAM is a very fair deal, and street prices will probably float just as much as 7600 street prices have.
Clearance-priced 6700XT is a great deal but 7600/7600XT is ultimately a 6600/6600XT replacement and it's not a knock on the 7600 that it doesn't have the wider memory bus/etc - it is a lower-tier product that is only in the same price tier due to clearance markdowns.
I maintain that people are just mad about the whole last 5 years (since RTX launched) at this point and pretty much just give automatic thumbs-down to anything that isn't an absurdly out-of-band good product. The pandemic shortages and mining boom have embittered a fair number of people to the point I don't think they're coming back to hobby, and instead they sit on social media and complain.
But future games are likely to run better in >8GB simply because the PS5 and XBox Series X have more than 8.
I think 8GB is going to continue to be a long-lived target especially at 1080p resolutions (with whatever gains can be squeezed from upscaling etc too - although generally DLSS needs to inference against the full-quality textures etc). Series S has 8GB (of fast-partition ram, the rest is GTX 970 style slow-partition) and even Series X only has 10GB.
People also aren't giving enough credit to mesh shaders etc, the GTX 1650 is actually still in the game with 4GB in Alan Wake 2[0], it does make a difference. The "but a 2060 super isn't relevant anymore!" argument relies on the assumption that you're deciding not to turn on upscaling etc. 1650 can run AW2 on lowest-settings 1080p with 4gb with FSR2 and it looks fine, and it'd be even better with RTX/DLSS. 2060 with DLSS can do a console-like experience on AW2 zero problem.
[0] https://www.youtube.com/watch?v=vFf8NsOi-HU&t=356s
Consoles have always been a mixed bag. Yeah, they get a lot of specific tweaking and they also have special hardware which helps somewhat. But overall you're working from a (later in the gen) fairly low baseline. 6700 non-XT performance is ok but nothing stellar, and optimization doesn't save that. But honestly what they are good at is removing "paralysis of choice", having too many choices really hurts people and the emotional feeling of having to turn the setting onto lowest hurts people, even if that's what the console does itself! It's at least pre-tweaked lowest settings etc (although often worse tech etc - FSR3 is blown away by DLSS 3.5 image quality let alone 4.0 and future iterations which aren't far away). You don't have to think about it, you just say "framerate or quality" and you probably know which you want.
That is the problem that people will struggle with. 8GB will still work. It just also will be a Series S level experience, modulo things like mesh shading that occasionally differentiate the consoles (PS5 lacks it iirc, as well as DP4a). And that can still often look fine. Will you get more from spending more? Yes. But it also doesn't take that much - series X is the 6700, series S is like APU territory. A 3080 blows away the series X, etc. But you will have to stomach through moving that slider from "native" to "performance" and the texture quality from "ultra" to "medium". Etc. People have lost touch of the world of yesterday when "can you run crysis" was an actual question and not a meme, slamming every setting to ultra is not a given when you buy an entry-level card, and people also can't handle the fact that $200-300 is now entry-level. Midrange is $500-700, high-end is $800 to "how much have you got".
And that's not NVIDIA, that's really just wafer costs. If you want to compare die sizes and MSRPs against 10+ years ago (look up GTX 670/GK104 lol), you have to bear in mind that a given die size might cost 5x what it did back then. And it increases ~30% every node-family since 28nm, more or less. It's gonna go up over time, if you aren't moving up in price you're moving down in product-design-bracket and are going to have to deal with more design compromises to hit those lower price-points in the face of rising costs. It sucks, but nobody has any better ideas - to paraphrase what someone once told me, "the industrial and creative poles of several societies and continents are laser-focused on pushing this backwards, and yet the problems only become more difficult after each success". There is no easy answer, lots of smart people are working at this.
EDIT: even people on the 4-series are experiencing it https://www.nvidia.com/en-us/geforce/forums/game-ready-drive...
Sucks for me, but overall I'm glad that Nvidia is getting prices under control to some extent.
to be blunt, that's because you bought into ayymd propaganda. it is so endemic that people don't even see it for what it is anymore, people are constantly bombarded with absurdly pro-AMD and absurdly anti-NVIDIA takes, it's just the sea in which we swim on social media.
you should take it as a learning experience and not constantly buy into the ayymd bandwagon of the week next time. because there will absolutely be a next time - probably people will move onto the next insane thing within a few weeks here.
Last year it was that the 4090 was going to be >900W... people talked themselves into thinking that a two-node shrink was going to result in zero efficiency gain. This ada gen is a dud, just wait for AMD, the 7900XTX is gonna blow the doors off!
https://www.techpowerup.com/294261/nvidia-allegedly-testing-...
And it's happened to RDNA3, Vega, Fury X, Zen2, etc, and against every single technology deployed via RTX or DLSS. The flip on framegen the day AMD released FSR3 was amazing, and instantly all the complaints about latency etc vanished within a single day, despite being significantly worse latency because of forced vsync/incompatibility with VRR, let alone the reflex-only baseline. "Possibly the best part of FSR now" etc.
it's like the runup to the iraq war or something, there were counter-voices, but why would you want to listen to them when everybody knows the truth already? Going against the grain constantly is tiresome and frames you as an iconoclast, and even if you're right people still think you're a troublemaker for having contradicted them earlier. The people who blocked you are not gonna unblock you just because you were right. It's like trying to be the voice of reason in a failing project, even if you save the project you're still a troublemaker. So eventually the dialogue just fades into an echo chamber. It is a fast road to what was eulogized as "epistemic closure" - aka "we bandwagoned too hard and drowned out all the opposing voices, and it turns out they were correct".
https://en.wikipedia.org/wiki/Epistemic_closure#Epistemic_cl...
So here we are: green man bad, everyone knows it, and this exception really only proves the rule. Now if you'll excuse me I've got some very important posts to make about how you'll never be able to buy one for MSRP anyway, like it's still 2020 or something.
In fairness you are not alone, fun reading from only a few days ago etc: https://news.ycombinator.com/item?id=38804502
> The GeForce RTX 4080 SUPER arrives January 31st, starting at $999
> The GeForce RTX 4070 Ti SUPER launches January 24th, starting at $799
> The GeForce RTX 4070 SUPER launches January 17th, starting at $599
But it's really more a reflection on how shitty the 4070 was than anything else tbh
SUPER series has been a response to rival products offering better raw performance/price released afterwards.
Power consumption is a separate issue which may or may not be a concern depending on where you live.
Most of the rest of the 40xx stack was either unexpectedly slow or unexpectedly expensive (or both), such that performance per dollar stayed flat or regressed
here's a review from a traditionally pro intel and nvidia outlet: https://www.techpowerup.com/review/nvidia-geforce-rtx-4080-f...
Or one of us sees what they want to see.
| Game | 3090 FPS | 4080 FPS | Delta | % Change |
| Cyberpunk 2077 | 83.3 | 114.6 | 31.3 | +38% |
| Doom Eternal | 261.2 | 378.4 | 117.2 | +45% |
| Forza 5 | 108.2 | 147.3 | 39.1 | +36% |
| Halo Infinite | 95.7 | 107.0 | 11.3 | +12% |
Maybe all those years where Intel was stuck on 14nm made me forget how big leaps could be generation to generation, but to me those jumps of more than 30% are huge especially considering that while the 4080 is a gen ahead of the 3090 it's also a tier down.Also if you look at 4K performance, the % gaps for all these games are even larger (and I'm not looking at 4K now because it's better for my point. My next monitor will be one of the 4K QD-OLEDs that were announced at CES, so those charts are now more relevant to me than the 1440p ones)
Very much a "I'll believe it when I see it" scenario.
I’m wanting to know how this would compare to a 4090…as I’m thinking of upgrading my 3080FE.
The improvements are very small in this space right now, for the most part.
2060Super GPU is starting to feel dated though!
That being said, these look competent enough, just stingy with VRAM still making them less desirable for longer use (4+ years) in either playing games or training models.
I feel the exact opposite. There are too many! There are now nine models of the 4000-series, and that's not including laptop models or the cancelled 4080 12 GB.
IMO, there should be 5 models at the maximum. I shouldn't have to sit here and do a bunch of research to find out if the 4070 SUPER is faster than a 4070 Ti, and whether I should go for a 4070 SUPER Ti.
4060, 4070, 4080, 4090. That's all they ever needed. Budget, mid-grade, enthusiast, top-of-the-line. That's all that's needed.
Though given that NVidia seems to release its consumer GPUs once every 2 years (20 series in late 2018, 30-series late 2020, 40 series in late 2022), I wonder if this is more of a marketing ploy - release the main series once every 2 years, but bring our "super" refreshers in the middle of the cycle to make sure you're still in the news, and get some segment of consumers / gamers to upgrade to those.
That said, I'm curious to know how they stack up against a 4090, and if these new cards can melt wires and burn down a house as easily.
3060 - no
3060 Ti - yes, it still a great card that I use today
However, if you can't find a 3060 Ti at a lower price point than the 4060 Ti... then I reluctantly have to say you are probably going to be better served with the 4060 Ti.