3090 is priced at $1500 for its 24GB RAM which enables ML & Rendering use cases (Nvidia is segmenting the market by RAM capacity).
AMD's 6900XT has the same 16GB of RAM as 6800XT, with less support on the ML side. Their target market is gamers enthusiasts who wanted the absolute fastest GPU.
RTX Titan for example was not as good in FP64 calculation compared to RTX Quadro. While first Titan was good at both. Nvidia kept changing what Titan means in order to extract most amount of money without hurting Quadro sales.
Even at its peak certain things were software limited for Titan, but not for Quadro. Nvidia has tons of artificial limitations for each of their graphics card once you go beyond gaming.
I also expect been a kind of struggle between their consumer side who need some other reason to sell their top end GPUs beyond just the whales when game developers aren't really that interested in pushing out features that will require it given the niche ownership, while the datacenter people want to protect their Quadro margins. This to me seems their most likely reason for the on-again/off-again relationship with pushing their high end cards for compute usage vs gaming usage.
I'm saying the only reason Nvidia didn't call it a Titan was because they knew it wasn't going to have an unassailable advantage over AMD.
A Nvidia card is Titan if it has the name and the drivers, and not if it doesn't.
AdoredTV made an entire video[0] to demonstrate this is absolutely baseless conjecture that NVIDIA Marketing tricked people into inferring.
It absolutely destroys the Titan idea, leaving no room for doubt.
curious what is that?
An official firmware release could change this but they're likely saving it for the Titan cards down the road.
The Reddit AMA implies the Ampere architecture supports this configuration but is software limited under the FP32 section: https://www.reddit.com/r/nvidia/comments/ilhao8/nvidia_rtx_3...
To epistemologically correct, I don't believe one will ever see a statement from NVIDIA confirming that this is a software limitation. I believe it's just inferred.
edit: This is close to confirmation: https://imgur.com/a/RH8vyz9 -though I may suspect there are non reversible hardware fuses at play too.
I mean, I really hope that it's just in the driver so that an enterprising reverse engineer can hack the driver and re-enable full FP16/FP32 accumulate :)
TITAN X got access to pro drivers shortly after AMD announced the Vega Frontier Edition.
So maybe not 100% designed for ML, but better than anything else out there unless you want to sell a kidney.
Here’s a dirty secret: there is a ton of ML prototyping done on GeForce level cards, and not just by enthusiasts or at scrappy startups. You’ll find GeForce level cards used for ML development in workstations at Fortune 50 companies. NVIDIA would love everyone to be using A100s to do their ML work (and V100s before that,) but the market isn’t in sync with that wish. The 2080TI remains an incredibly popular card for ML even with only 11GB. Upping to 24GB, even with the artificial performance limitations for certain use-cases, enables new development opportunities and use-cases to explore.
When it comes to product stratification, the hard rule according to driver EULAs is that GeForce cards can’t be used in data centers. For serious ML development at scale, NVIDIA has their DGX lineup. In the middle are the Quadro cards, but they tend to be a poor value for ML. The cost differential with Quadro is largely due to optimizations and driver certification for use with tools like Catia or Creo (CAD/CAM use,) which don’t intersect with ML.
The Titan RTX may not have the gimped drivers, but the 3090 beats the Titan in many benchmarks nonetheless. Is the 3090 the best NVIDIA PCIe form factor card for ML? No. The A100 is still king of the crop and is the only Ampere card with HBM memory, and even the A6000 will outperform for many use-cases with 48GB of RAM. Still, the 3090 will be the optimal card for many.
I’m one of the lucky few to have a 3090 in my rig. I lead of team of volunteers doing critical AI prototyping and POC work in an industry give-back initiative, and price was not a leading factor in my decision to procure a 3090 over a Quadro. I chose the 3090 principally because I didn’t want a loud blower card in my computer (and I don’t need 48GB.) If someone donated an A100 to our efforts, I’d gladly take it, but it wouldn’t replace the 3090. It’s not a graphics card and it won’t play games, which indeed is an important value-added benefit of the 3090 :)
Aka, it's a more expensive 2080 Ti, not a cheaper RTX Titan.
AMD performance is on par or slightly higher. Power efficiency is much higher.
And costs $500 less.
A slower card that uses way more power and costs $500 more is really hard to sell, even with NVIDIA marketing team being as strong as it is. At those prices, few people are going to automatically buy the product without exploring their options.
Are there benchmarks showing this?
And they've obviously used the same preset for both their cards and the competitors'.
There's some more details in the press kit, but I do agree with the principle that decisions should be withheld until NDAs expire and third party benchmarks are available.
What's clear is that, with the information in hand, buying NVIDIA Ampere cards is simply not sensible. Waiting for RDNA2 reviews is.
Still it's more of a Zen 2 moment than a Zen 3 moment. From the lack of comments on RT performance compared to the competition (just that it was added to the hardware) it seems extremely unlikely the RT performance is at the same level. The cards also lack the dedicated inference hardware for features like DLSS or voice/video filtering. And the card still has less VRAM than the 3090. These are all minor but if you put them together it seems really unlikely we'll call the 6900 XT the absolute best performing GPU of the generationjJust like Zen 2 didn't topple Intel's claim of "best gaming CPU". We'll have to see 3rd party reviews and benchmarks to find out for sure though. What it does represent though is a huge upset in the 3070/3080 area where most cards are sold and a hint that there may be an Zen 3 moment coming for GPUs in the next generation where AMD really drives top tier performance to a new level after long stagnation instead of "just" coming close to taking the crown dead even.
Personally (and this part isn't going to be reflective of the average person) I was going to be a 3090 whale and I probably would still be if it weren't for Nvidia's shit stance on open drivers in Linux (one of my biggest gripes with my 2080 Ti). However with AMD being so close this round and me not having liked DLSS or RT on the 2080 Ti I'm willing to trade off for the 6900 XT. The $500 is a nice bonus but not really what's coming into play, like I said if I were trying to get perf/dollar the 6800 XT makes WAY more sense. This is similar to what happened with Zen 2, I was planning on getting the better Intel CPU for the couple extra FPS but I was fed up with meltdown type issues and Zen 2 was really damn close. Now I'm really excited for Zen 3 though :).
Which it doesn't. Thanks to AMD having some new, large "Infinity Cache" feature which they adapted from their Zen CPU architecture.
At the end, performance is what matters.
Edit: should clarify that I’d really love to get a quadro or one of their data center cards, which aren’t gimped in certain non-gaming workloads... but I’m not made of money :)
Would you still do it without hesitation if the money was coming of your own pocket?
Now on the other side of the world, where a 3090 is several times your rent, you really need to think thrice about buying one.
AMD GPUs have near zero support in major ML frameworks. There are some things coming out with Rocm and other niche things, but most people in ML already have enough work dealing with model and framework problems that using experimental AMD support is probably a no go.
Hell, if AMD had a card with 8GB ram more than nvidia, and for 500$ cheaper, I would still go with nvidia. Everyone wish AMD would step their game up w.r.t ML workloads but it's just not happening (yet), Nvidia has a complete monopoly there.
* It seems there was not a single commit in the past ~6 months which by itself is already a deal breaker.
* Documentation is lackluster
* You need to use Keras, I use PyTorch. This is not a deal breaker, but a significant annoyance.
* Major features are still lacking. E.g. No support for quantization (afaik), which for me is fundamental.
* Most importantly there seem to be no major community around it.
It feels a bit bad to say that, because clearly a lot of work went into this project, and some people have to start adopting it to drive the momentum, but from an egoistical point of view, I just don't have the courage to deal with all the mess that comes with introducing an experimental layer in my workflow. Especially in ML where stuff can still appear to "work" (as in, not crashing) despite major bugs in the underlying code, leading to days or week of lost work before realizing where the issue is.
Check out keras-helper.. I made it to switch between various backend implementations which are non NVidia specific.
Pytorch may eventually need porting, but for now I don't need it. I've been trying out Coriander and DeepCL now but I decided to stick to Keras, which seems to be a decent compromise. Not using 2.4.x though, do not need it.
OpenCL based backends are cutting it for me, running production workloads without needing to install CUDA/ROCm is the best way to go.
Spending a week/year working around ROCm would already cost you 5k$ plus the opportunity cost. For a whole team that’s a money sink.
The catch is that ML software stacks have had hundreds if not thousands of man-years of effort put into things like cuDNN, CUDA operator implementations, and Nvidia-specific system code (eg. for distributed training). Many formidable competitors like Google TPU have emerged, but Nvidia is currently holding onto its leadership position for now because the wide support and polish is just not there for any of the competitors yet.
The analogy is that everyone should get NVIDIA Ampere units (non consumer) units worth $30k because it's fast and you'd rather be spending less time in a lab with millions of dollars in funding. insert don't be poor T Shirt reference
PlaidML is not ROCm. Nobody needs ROCm, what people need is just linear algebra well implemented with OpenCL primitives. That's what PlaidML is. And it works quite well, even on those integrated Intel GPUs on most laptops.
Have you also looked at DirectML and WSL2? They seem to be running tensorflow quite well too. Those things may be the key to bringing these in adoption outside the well paid class of data scientists you came up with.
_If_ AMD has made a sufficiently powerful GPU, that will add a lot of incentive to ML frameworks to support it. But it's going to have to be a big difference, I imagine.
Given how active AMD is in open source work, I'm a little surprised they haven't been throwing developers at ML frameworks.
So AMD got a lot of heat for not supporting Navi (RDNA) with ROCm, but it seems that they are weeding out the things keeping people from running it ( https://github.com/ROCmSoftwarePlatform/pytorch/issues/718 and the links in that look like gfx10 is almost there for rocBlas and MIOpen). We'll see what ROCm 3.9 will bring and what the state of big navi is.
Why would I get a Radeon VII when used nvidia cards for machine learning are extremely cheap, and then I don't have to worry about experimental stuff breaking one day before deadline lol
That is incorrect - I have been running Tensorflow on RX5*0 cards for close to 2 years now. I even transitioned to TF2 with no problem. Granted, I have to be extra careful about kernel versions, and upgrading kernels is a delicate dance involving AMD driver modules and ROCm & rocm-tensorflow. My setup is certainly finicky, but to say AMD GPUs have near zero support is false.
gimme 48 GB gimme 128
edit: Also CUDA is just too important to switch to AMD.
This would make sense if these were Titan cards with ML drivers. Instead, they are not, only the regular drivers are available, and FP32/64 performance is artificially capped.
These aren't cards for ML.
The ML driver and FP32/64 performance capping aren't really issues since in reality we rarely hit those limits.
Are individuals buying graphics cards for ML? I would think that it makes more sense to provision scalable compute on a cloud-plaform on an on-demand basis than buy a graphics card with ML capabilities?
So yeah, the huge gap ML capabilities between AMD and Nvidia are a selling point, but probably for a small enough group that it doesn't make a difference.
Of course this calculation depends on how bursty your compute requirements are and how much you pay for electricity (datacenter cards are more power efficient)
I guess if you're rolling your own compute clusters, you've probably rolled your own storage solution too?
Using cloud services for work would mean using the exorbitantly expensive services from Azure, to say nothing of the painful and annoying experience that using Azure is at the best of times.
Instead, I can spend a moderate amount, build a machine that is more than adequate for work requirements will last ages, write it off on tax and still end up spending way less than I would have renting a cloud machine for a couple of months.
[1] https://www.nvidia.com/content/dam/en-zz/Solutions/geforce/a... (Appendix A)
BTW - Whether it "could" be faster is indeed relevant because some of us are holding out for a Titan GPU next year with this unlocked. If you have unlimited budget or are under time constraints then by all means get the 3090, it is a beast. But if one has a 2080 TI then it's an important consideration.
If you can wait until next year you should always wait until next year, because there will (almost) always be something better than what is currently out. That's unrelated to whether or not the 3090 is good for doing ML research; it objectively is.
The 3090 will probably stay at a premium due to a few factors - I don't think ML performance plays into this at all though:
DLSS - There's a reason AMD cites "raster performance" since with DLSS enabled the 3090 has a major advantage.
AMD-specific optimizations - near the end of the presentation AMD disclosed that with all the optimizations on (including their proprietary cpu->gpu optimizations, only available on the latest gen cpus)- they could pass 3090's raster performance in some games.
I think for these two reasons, and the fact nvidia can't seem to get cards to vendors (and customers), there won't be a price drop on this SKU. They may however release a watered down version as a 3080ti and compete there.
That's also the target market for the RTX 3090, all the Nvidia marketing material describes it as a gaming card, it's Geforce branded and Ampere based Quadro's will be a thing.
There was/is also this whole thing: https://www.reddit.com/r/MachineLearning/comments/iz7lu2/d_r...
Same with CUDA vs ROCm.
And getting ROCm set up is still a buggy experience with tons of fiddling in the deep inner workings of Linux so it is nearly impossible for the average machine learning engineer to use.
It is compelling for certain bespoke projects like Europe's shiny new supercomputer, but for the vast majority of machine learning, it is totally unusable. By now in ML world the word "gpu" is synonymous with "nvidia".
We actually had more issues with nvidia drivers messing up newcomers' machines during updates than with setting up AMD GPUs, but then again n is small (and AMD GPUs were for playing around rather than real work).
Still, a Titan Xp has CUDA support and plenty of memory, but it's better, IME, to upgrade to a model with less memory but higher cuda compute and access to tensor cores.
For the amount invested in the hardware development, the amount AMD have been investing in the software side has been shocking.
Also, would the TI versions actually help? Even a 100$ markup on a 3080TI would bring it close to the 6900 XT pricing. And the 3080TI cannot come close the 3090/6900 XT performance for that markup or otherwise it would risk cannibalizing Nvidia's own products.
Nvidia's only hope at this point is that either AMD fudged the benchmarks by a large margin or AMD gets hit with the same inventory issues.
(In the way they use their devs to help AAA games to fix some issues under some circumstances, there had been cases where optimizations speed up Nvidea but hindered AMD due to architectural differences but, surely it was all accidentally).
https://techreport.com/news/14707/ubisoft-comments-on-assass...
"Radeon HD .. gains of up to 20%.... Currently, only Radeon HD 3000-series GPUs are DX10.1-capable, and given AMD’s struggles of late, the positive news about DX10.1
Ubisoft’s announcement about a forthcoming patch for the game. The announcement included a rather cryptic explanation of why the DX10.1 code improved performance, but strangely, it also said Ubisoft would be stripping out DX10.1 in the upcoming patch
Ubisoft decided to nix DX10.1 support in response to pressure from Nvidia after the GPU maker sponsored Assassin’s Creed via its The Way It’s Meant To Be Played program."
https://techreport.com/review/21404/crysis-2-tessellation-to...
"Unnecessary geometric detail slows down all GPUs, of course, but it just so happens to have a much larger effect on DX11-capable AMD Radeons than it does on DX11-capable Nvidia GeForces. The Fermi architecture underlying all DX11-class GeForce GPUs dedicates more attention (and transistors) to achieving high geometry processing throughput than the competing Radeon GPU architectures."
GameWorks slowing down ATI/AMD users by up to 50% https://arstechnica.com/gaming/2015/05/amd-says-nvidias-game... https://blogs.nvidia.com/blog/2015/03/10/the-witcher-3/
Edit: Here’s the video: https://youtube.com/watch?v=TY4s35uULg4
Something is very definitely wrong.
It is not new and untested, it's been in use since 2018.
But on a serious note - why are they doing so? Surely, they have good predictions of supply and demand - why are EBay resellers reaping the profits from the gap between the two? Why aren't these cards retailing for $1000, with price reductions happening as demand drops below supply?
And I do not feel sorry for them.
Personally, I think its less due to these and more due to people using bots to buy up stock for scalping. The same stocking issues have happened for the new Xbox and Playstation, and even for non-tech hardware, like the newest Warhammer 40k box set. The bot scalping issue is just becoming more pervasive.
Back then these highend-GPUs used to be prestige projects that mostly existed for the marketing of "We have the fastest GPU!", the real money on consumer markets was made with the bulk of the volume in the mid-range.
How much is this one really about TSMC 7nm vs Samsung 8nm?
Then again, I'm rooting for AMD all the way until they become the new evil.
Isn’t A14 a “5nm” process chip? Why would it be compared to intel and their 14nm++++?
I'd say the crowning achievement for this architecture is the "Infinity Cache": https://twitter.com/Underfox3/status/1313206699445059584
"This dynamic scheme boosts performance by 22% (up to 52%) and energy efficiency by 49% for the applications that exhibithigh data replication and cache sensitivity without degrading the performance of the other applications. This is achieved at a area overhead of 0.09mm²/core."
See also this presentation: https://www.youtube.com/watch?v=CGIhOnt7F6s
TSMC has to spit out Zen3, RX6000, plus the custom versions for XBX and PS5, all around the same time...