The new GPU world order is beginning to take shape
theregister.com
theregister.com
Regardless of what AMD does (although I hope the Reg is right on their pricing thoughts), the market is highly disrupted on the low end due to Intel. If you run DX12 games (and as time goes by, more games will be), then Intel will be a nice price point.
Assuming that Battlemage is something akin to a 3080, but at the price of Arc, then they could cause real problems due to the price anchoring effect. [0]
[0] - https://research.stlouisfed.org/publications/page1-econ/2021...
If it's the game, all bets are off. No truly popular game is abandoned though. There's even a FreeCiv. If can be done in the drivers, expect nearly every even remotely popular report coming in to eventually be fixed. This hardware is going into Intel's CPUs sooner than later. Support won't be lacking. The support for this card may dwarf AMD and Nvidia's own support.
Intel is using D3D9On12 and perhaps D3D11On12. I would expect any extremely popular titles like CSGo (DX9C) and LoL (DX11 and offers a legacy DX9 mode) to get hand-tuned by Intel.
For most recent games and going forward, most games target engines, not APIs. Between that trend and DX12, I expect Arc to be just fine. I can find you driver bugs with both Nvidia and AMD today. I've experienced visual bugs with both. I'm using an older branch of Nvidia's drivers right now due to stuttering that occurs on newer branches. There's no outrage because this happens all the time with GPU vendors. Intel puts out a very good dGPU and people are hyper-critical.
Well, for every FreeCiv and Devilution there are 10K other titles that no one is reimplementing.
Nvidia started development on this card before the crypto crash. This card made sense for a market in which users expected to be able to spend money to make money, it does not make sense for a market in which
The Steam hardware survey says that only 0.52% of users have an RTX 3090[1]. Pasting the data [2] into a spreadsheet, some quick selection groups say that more than half of all Steam users have hardware older or less powerful than a 1080. They've all been priced out; why would you build a game that most users can't play?
[1] https://store.steampowered.com/hwsurvey/videocard/?sort=name [2] https://pastebin.com/w7rneTFc
You sure about that? It still feels like we're being bent over and Nvidia is just laughing.
And yet I still can't buy a two year old 3080 for its launch price MSRP.
As comes up in these threads over and over again, Nvidia does set a price floor their partners aren't allowed to sell below.
Intel Arc cards are dead in the water because they have extremely unusably unstable drivers. To the point of the word "broken" being applicable. This hasn't gotten better with the A750 (in fact, it's gotten worse it seems?).
The 4090 is extremely expensive, but it is ludicrously stupidly comically powerful. It's 10x-20x more powerful than a 2080, nearly 100% more powerful than a previously more expensive MSRP 3090ti and this is BEFORE DLSS3, which is not a gimmick. Check out the GN review of the 4090, the metrics are ridiculous. F1 at 4k gets 230 FPS vs 135 for a 3090ti and 41 for the 2080. Total War at 1440p gets 235fps vs 135 for a 3090 ti and 47fps for a 2080.
Yes, its extremely expensive, but it's Nvidia's HERO card. You're going to see downward movement with the 4080 models (one basically being a 4070 in disguise) next month.
These cards are too powerful for anything out there today (they do cyberpunk at 4k at 80fps without DLSS and 135 fps with it) but this is a card you buy if you're looking at 3-5 year upgrade cycles.
Comparing this card to Intel's Arc is crazy. It's like comparing a Ferrari to a Civic. They both serve a purpose, but their markets are completely different.
That being said, I bet you the 4090 will outsell the intel arc 750 by a stupidly large margin. It has to suck if you spent $2000 MSRP on a 3090ti in March though.
I think AMD is going to somewhat come to the rescue because chiplets are going to get way more out of each wafer compared to the absolute massive dies that Nvidia fabs but we're not returning to $800 cards in the penultimate slot like before. A lot of people are probably about to get disappointed.
The timing would have been perfect one year ago, as was originally scheduled. You couldn't actually buy a graphics card unless you paid at least double. That would have been the perfect point in time to enter the market, EVER. Instead they are now too late with a card that was designed to compete against Ampere when Ada is coming out. They missed the perfect opportunity.
I still hope that Ada's price will come down once Nvidia has cleared its considerable inventory of unsold Ampere SKUs. This will not happen until next year, unless AMD is pricing RDNA3 aggressively, which they would be stupid to do.
Intel has always had good drivers for Linux (compared to Nvidia) which are scheduled to land in the upcoming 6.x kernels, I'm wondering if they're planning to attack the CUDA AI/ML moat that Nvidia has built up. AMD hasn't done much to shake CUDA with ROCm but I'm sure Pat Gelsinger has plans to capture all that margin that Nvidia is reaping currently in the datacenter space, those Nvidia DGX A100's are so insanely pricey!
It's really a price/performance king for compute applications and has my attention. Gaming/drivers is a concern as it is a 1st generation product though.
Is there a compatibility layer for CUDA, or is it OpenCL only (or is there something else)?
It apparently works well enough that Blender's cycles works on SYCL.
In my understanding GPUs do mostly parallel worloads: SIMD with limited code size, operating on the same batch loaded arrays, but also with massive local memory.
And CPUs do mostly serial workloads: code size can be gigantic, lots of indirection, can pipeline a handful of operations (computations but several inflight small data loads), complicated cache herarchy; with massive silicon areas dedicated to shortening this serial critical path (branch prediction, instruction reordering, speculative execution)
Is there a new paradigm that we could see emerge and go mainstream in a few years? We do away with branch prediction, instruction reordering, speculative execution, and instead we can have massive amount of small dumb independent cores. These could do their share of small independent loads and would have small amounts of local working memory, could share some data with nearby cores. It would be common for these to halt for a long time waiting for a load from RAM, but that'd be ok.
Could we see that appear? Would it get the performance increases GPUs get? What language and coding style would be best suited for that silicon? What would be the market drivers (certainly not matrix multiplication/AI, nor single-threaded performance which are GPU and CPU's unfair advantage respectively)?
You're gonna have to explain why a GPU is illsuited for this task. GPUs even have __shared__ memory space that passes data between local threads at extreme speeds (comparable to GPU L1 cache).
> but also with massive local memory.
On the contrary. CPU has more "local" memory if we're talking about L1, L2, and L3 caches.
GPUs have massive register space, but very little high-speed memory (aka: cache). CPUs probably win if the data fits inside of say, 10MBs with hot-data within 512kB.
GPUs win if you can get the entire computation inside of register space. But if you have lots of lookups, GPUs have very small caches, and the latency to GPU cache is an order of magnitude slower than CPU caches.
GPU vRAM is just RAM. GPUs have much faster vRAM. GDDR6x has well over 500GBps bandwidth, while DDR4 and DDR5 will be in the 50GBps to 100GBps on typical machines, maybe 200GBps to 400GBps on servers.
The per-core (or SM / WGP) L1 or L0 cache of GPUs is quite anemic. The GPU designers clearly "intend" for the programmer to hold as much state in registers as possible, rather than in cache, for their computations. You really want to keep the state at 1024 bytes or less in practice for high speed GPU computations.
CPUs with 2MB L2 cache per core (alder lake from Intel), or 1MB cache per core (Apple M2 and/or AMD Zen4), are simply designed to handle bigger "states per thread". Using the full 1MB to 2MBs for this L2 cache is reasonable.
... And if you can successfully "hide the latency", then you kind of don't care about the latency...
I'd estimate CPU RAM latency to be 50ns to 100ns and GPU vRAM latency to be 100ns to 500ns, depending on architecture. (There's many more GPU architectures out there, and they all have very different latency characteristics)
But again, GPUs "don't care" about the latency to some extent due to their massive number of concurrent wavefronts/blocks. While CPUs also "don't care" because of L1/L2/L3 cache and out-of-order execution.
Its probably more important to understand the different strategies of latency-hiding between CPU vs GPU, rather than knowing the actual latency figure.
I suppose one possible gain could come from using integrated GPUs that can access CPU cache directly. Reducing the data upload to the GPU would really open up the more workloads to GPU type work.
What developers want are SIMD co-processing cards with good software support that lets you push a framebuffer over a wire occasionally. Nothing more. That's going to commodify the graphics card market--a place where NVIDIA really does not want to be.
If you look at modern game engines, they are taking more and more of the responsibility for rendering the graphics and basically want the GPU drivers out of the way.
Intel could slaughter NVIDIA right now. It won't because you have to throw an enormous amount of resource at the software, and that's just not in Intel's DNA.
From the physical appearance of the Intel Arc, it looks to be a much more power efficient card. I hope they keep evolving it and maybe their next, or next-next, generation will be in my shopping cart.
In my experience, all modern cards idle at very low power usage (less than 10W), and scales up: using more and more electricity the more difficult the video games you throw at them.
One of the better settings to come by (especially for laptops) is FPS limiters.
We have to hope AMD has the price, and just as importantly volume, to make the GPU market competitive.
That might be ok. The integrated graphics on Intel chips are getting better - probably 100% thanks to the discrete GPU effort.
Apple has shown with the M1/M2 that integrated graphics can be really quite good even without high-end Nvidia performance. If Intel matches that, they could own the low-to-mid range just by selling CPUs and leave Nvidia in a pickle with no profitable market for their binned chips.
Don't think about the present, think about what the landscape could look like in 3-5 years ;)
Don't get me wrong, it's stunning performance, and having it on the same package as the CPU and GPU has other benefits.
The GN review really shows the harsh reality of Intel ARC
https://www.youtube.com/watch?v=nEvdrbxTtVo
Even Linux support is limited to Linux 6.0, it might help that the DirectX to Vulkan stack in Linux is Better, but it still going to be a worse buy compared to AMD.
All Intel needs is to fix their drivers. Don't know if they're capable of that.
What The Register is forgetting to mention is the absolutely insane power consumption of NVidia's new cards. I don't see how you can afford one in Europe unless you're making money with it. And crypto mining is dead.
What are you talking about?
The 4090 uses less power than a 3080ti while being 63% faster.
https://www.techpowerup.com/review/nvidia-geforce-rtx-4090-f...
I mean, it's a lot of power, but you don't buy a lamborghini so you can pick up grandma from the shops.
> Even if the 12GB 4080 manages twice the performance of the 3070, at nearly twice the price, that doesn't make it better value.
I'm not a GPU expert: what other characteristics other than performance are there to consider when deciding value? Intuitively, it seems like if something is twice the performance and less than twice the price, it is a great value.
For instance, suppose some not-too-demanding game runs, at quality settings you find satisfactory, at 50fps with one GPU and at 100fps with another. You are not going to enjoy playing at 100fps twice as much as you enjoy playing at 50fps.
Suppose it's 20fps versus 40fps. Playing at 20fps might be a horrible enough experience that 40fps really is twice as enjoyable; but in that case you probably won't play at 20fps, you'll turn some settings down so that everything looks a little fuzzier or the reflections aren't as realistic or something. You'll get a worse experience with the worse-performing GPU, but again it's unlikely that you enjoy your games half as much.
Even if you are doing something, like training neural networks, for which twice the performance really does translate into twice as much benefit in some meaningful sense, that's not necessarily the same thing as "worth twice as much money". It might be worth more than that (now you can train the NN you're working on in time to get a paper into that important conference, and before you couldn't). It might be worth less (training neural networks isn't a large fraction of what you do and you don't really care much how fast it happens).
Analogies: A car that can drive at 200mph isn't worth 2x as much as one that can do "only" 100mph, to most people; to a racing driver it might be worth a lot more than 2x. An oven that can heat things up to 600 degC isn't worth 2x as much as one that can reach "only" 300 degC, to most people; to a pizza chef, it might be worth a lot more than 2x.
It's not a great value when it's nearly twice the price and double the performance two years later. It has been reasonably common to get 50% or more performance increase between generations for the same price, and rebranding a XX70 as a "4080 12GB" to increase the price doesn't change that. It's also unusual for a XX80 to be "nearly double the price" even a generation later.
The pricing is out of line with what would historically be expected, as is having a completely different die.
I don't think that's the case, specially with a TDP of 450W
> So where does AMD fit in this new GPU world order? We won't know for sure until the House of Zen rolls out its RDNA 3 GPUs early next month, but the company's Ryzen 7000 CPU pricing does offer some hints
The RX 6xxx XT gen do better than the newer intel GPU and with a LOWER TDP!! close to half the TDP in fact!
I don't understand why this journalist intentionally ignore and downplay AMD, maybe the article is sponsored? maybe they prepare the shutdown of AMD since it's Taiwanese, China?
The RX 6xxx XT gen do better than the newer intel GPU and with a LOWER TDP!! close to half the TDP in fact!
In rasterization. But the future of gaming is "DLSS" and ray-tracing. Both areas that AMD struggles so far. It's generally accepted, but muted in most reviews that both Intel and Nvidia are ahead in ray tracing and upscaling tech compared to AMD. It doesn't bode well for AMD's market has what appears to be a pretty competent opponent in that space now with Arc.
That's marketing BS, and today nobody is playing in the future
RX 6xxx gen is 2 years old already ;), and yet it does better for todays needs and for already released games
That's not the case for the Intel lineup
That's not what I meant. I meant DLSS and RTX are the technologies that GPU companies should be focused on. Intel and Nvidia are focused on those things. But DLSS/RTX are the present and future of gaming.
RX 6xxx gen is 2 years old already ;), and yet it does better for todays needs and for already released games
That's not the case for the Intel lineup
At some price points, in some games, yes. You could say that about all 3 vendors. There's no outright win, especially for AMD and Intel. They battle in the price/performance field. If anyone has an outright win on performance, it's only Nvidia.