Coding Mistake Made Intel GPUs 100X Slower in Ray Tracing
tomshardware.com
tomshardware.com
I mean it's all based on the state of the driver used with pre-release GPUs and often the worst GPU of the lineup.
You could say this are beta drivers, but that is somehow not something people mention.
I mean sure there was a lunch of arc mobile GPUs but only in some specific region which was neither the US nor EU, it's as far as I can tell bit like a closed beta release from the dynamics involved.
So shouldn't we wait until the "proper" market launch of the dedicated gpus in the US before taking it apart as a catastrophe?
And sure older Games might not run as well, may some will never (which doesn't mean they don't run good enough to be played nicleye). Maybe except on steam because of the emulation of older direct X versions being base on Vulcan, that will be interesting.
Intel seems to have rightly recognized that the driver advantage is a huge moat for those guys - they have to instead compete on price and focus on having good support for titles that will get them the the biggest chunk of the market.
That said, man, if they could have released these a year ago the wind would have been at their back way more than it is now with GPU prices trending back towards MSRP.
FWIW I've heard DXVK actually can run on Windows, but you can't use it in AppX containers. Perhaps Intel bundling it with Ark's drivers would be a better option moving forward?
It's okay that those cards are reviewed according to their current state. They are pretty bad. But that doesn't mean the second or third generation of those cards has to be bad, when they remove the resizable bar requirement (so they can be used in older systems) and improve the drivers. DXVK would be a great route, but I doubt Intel can push this on a driver level - it would be more of a Steam or Windows thing.
I assume those cards will be a great option for Linux systems soon, where DXVK is the default anyway. It's good to have an alternative to AMD.
They'd have to add support for changing the base address register. The easy default is just mapping all the memory directly. It's a lot more programming work to transfer data when you only have a 256MB window to work with. No use sticking with a limit that PCIe hasn't had since 32-bit.
So requiring the “easy default” to be available means prohibiting any one with a system from ~2020 or earlier from upgrading to these GPUs.
Case in point Baldur's Gate Dark Alliance 2 just saw a re-release where the devs didn't do much but port it to other systems and render the game at a higher resolution. It had a release day bug that caused the game to crash so [some review sites pan the game](https://xboxera.com/2022/07/21/review-baldurs-gate-dark-alli...) giving it low scores of sub-5-out-of-10. The devs fixed it the next day. Now is that particular site going to revise their review? No. It sits at a 4/10 with a small note at the top saying the glitch has been fixed.
Court of public opinion is obviously a thing with Intel, and there's a lot of long-established fanboyism with "team green" and "team red", so there's a lot of people looking to Intel to fail.
Also I hope it's only a matter of time until companies embrace abstraction layers like dxvk.
It was really amusing to see Apple hype that term (which has been around in the PC industry for decades and synonymous with low cost and performance) so heavily around the release of the M1.
The contrasting term is "DIS", i.e. discrete graphics.
Promising to price their GPUs based on their performance in "Tier 3"* games will certainly help them win a lot of consumer goodwill though, especially with Intel targeting the critically underserved low and midrange gaming GPU market.
* For those OOTL, Intel has grouped all games in to three tiers based on driver optimization and graphics API usage. Tier 1 is DX12/Vulkan titles they have specifically optimized for. Tier 2 includes all other DX12/Vulkan titles. Tier 3 is DX11/OpenGL/older DX. Nvidia and AMD have had a decade+ to optimize their drivers for higher level APIs like DX11, including hundreds of game/engine specific optimizations. The result is ARC GPUs performing notably worse in DX11 titles than equivalent hardware from Nvidia and AMD.
Intel's promised pricing structure means that Your $250 ARC GPU will perform about as well in most Tier 3 games as a $250 RTX or Radeon. Meanwhile, Tier 1 and 2 DX12/Vulkan software will likely outperform competing devices in the same price range.
Secondly, most people in the world know China can ship hand-sized objects to their country and aren't above reselling.
Thirdly, in many countries it seems that postage from China is cheaper than postage from across the country.
This seems to me to be an officially deniable launch. A beta-launch as it were. A way Intel can say "Well, we didn't release it to you because it wasn't ready for you..." while still actually having launched it and having the whole world tell them what they need to fix.
But a release is a release. So they should be judged on their product. Which so far seems lacking.
I have hopes for the future though as they work on cleaning this mess up.
I get what you mean but people seem to be forgetting that Ryzen launched in an absolutely abysmal state and it took them a few iterations to get to the industry leader that we have today. I think Intel taking the loss leader perspective on this and essentially using people willing to buy as beta testers will pan out for them in the long run. They have been more transparent about it than AMD was with Ryzen or than Nvidia has been with literally anything which I appreciate.
If you want good reviews at launch, then launch a finished product.
This mistake could easily have been in other vendors Linux GPU drivers, they in the end don't have nearly the same priority (and in turn resources) as the Windows GPU drivers. And it's a very easy mistake to find. And I don't know if anyone even cared about ray tracing with Intel integrated graphics on Linux desktops (and in turn no one profiled it deeply). I mean ray tracing is generally something you will do much less likely on a integrated GPU. And it's a really easy mistake to make.
And sure I'm pretty sure their software department(s?) have a lot of potential for improvement, I mean they probably have been hampered by the same internal structures which lead to Intel faceplanting somewhat hard recently.
It's not like the team who write the drivers are likely to know of the team working on optimizing compilers, profilers, or anything at all really.
My experience has been that especially in companies working in diverse disciplines across disparate codebases, very little is shared. A team of 8 in a tiny company is just as likely to make the same mistakes as the team of 8 in a bigger company. At large companies with more unified codebases and disciplines, maybe one person or team has added some process which helps identify egregious performance issues at some point in the past. But such shared process or tooling would be really hard at a company like Intel where one team makes open-source Linux drivers while another makes highly specialized RTL design software, for example.
Because a massive company has enough money around to put the processes in place and hire skilled people to do both deep[0] testing and system[1] testing.
[0] https://www.developsense.com/blog/2017/03/deeper-testing-1-v...
[1] The definition of "system testing" I'm using: "Testing to assess the value of the system to people who matter." Those include stakeholders, application developers, end users, etc.
[0] https://www.intel.com/content/www/us/en/support/articles/000...
Source: I work for a similar massive company. You would not believe the amount of issues similar to this. This one is gettin attention because it happened in open source code.
I don't know what your position or political standing in the company is, but I assume that with the tech job market the way it is, if you still work there you care about the company to some degree. So perhaps bringing this issue up with (more) senior management is the way to go.
And if they say there is no budget, or that it would take a bureaucratic nightmare to make space for it in the budget, ask them what the budget is for dealing with PR disasters such as this one.
So you would have to profile GPU-side code, which is probably really hard; and you'd have to find slow memory accesses, not slow code or slow algorithms, which is even harder. And those memory accesses may be spread out, so that each instruction which uses the slow memory won't stand out; the effect may only be noticeable in aggregate.
This being a mistake/bug, I think it's obviously underperforming rather than being optimized.
The article jumps right into that:
> This is something to be celebrated, of course. However, on the flip side, the driver was 100X slower than it should have been because of a memory allocation oversight.
Maybe this is where the phrase "It's a wash" comes from
Anyway, I'm hopeful things mature quickly - having worked with Intel people, I'm cautiously hopeful
Within six months I discovered a bug in the new algorithm whose removal delivered roughly another 10x improvement. (That’s the most dramatic single speed up of a previously well designed algorithm I have delivered in my career.)
Numerical algorithm bugs can be tricky to detect when sandwiched between dramatic improvements like that!
> This is why there shouldn't be such a bit. :) Default everything to local, have an opt-out bit!
yep. why was the default the slow path?
Only thing I can think of is that they reused to code from their integrated graphics drivers which likely don't support GPU memory. So the default is the option that works on the millions of "GPUs" they have already shipped.
That said, the 810/815 that came after did have a (to my knowledge little-used) "display cache" feature, but they were otherwise still a UMA design.
I use the Rend3->wgpu->Vulkan chain, and can't get any info about how much GPU memory is available or left. If I load too many textures into an NVidia 3070, they spill into main memory and rendering continues, slowly. If I load too many meshes, I get a fatal error from the Rend3->WGPU layer, because the allocation that failed happens long after I made the request to load a mesh and it's too late to just return an error for the request.
Can anyone explain why it's so hard to find out the memory allocation situation in Vulkan?
https://registry.khronos.org/vulkan/specs/1.3-extensions/man...
In the midst of the pandemic, with the GPU (and chip) shortage on one hand and Intel releasing a decent card with good enough driver reasonably priced would have really given them a great window of opportunity to catch Nvidia and AMD in the GPU blind spot.
It's been so confusing a GPU story to follow that it shows that there has been a lot of confusion and poor planning plus prioritisation in shipping this discrete GPU from Intel, ironically a story not different from their other product and probably indicates a systemic failure. The next year or so should tell us if Pat G can really turnaround Intel or not, most likely not.
It seems like a super easy oversight to make, and if it’s always been there you won’t see a regression in your tests as existing baseline by definition couldn’t exercise the feature.
A profile of the tests is only helpful if you know what should or should not be there. I've always been of the impression that intel is a company that outsources writing drivers, which would easily lead to devs who are reading documented interface but not actually involved in the hardware design.
I got intrigued though, and indeed you can read the thing both ways: an optimization gave a 100% boost, or an oversight missed out a 100% boost.
Curiously, Tom’s went with the scare tactic, and got me thinking: is it because clickbait, or do they consciously support the responsibility in software engineering movement?
To note; a 100× speed up is a 9900% boost. A 100% boost is a 2× speed up.
Does anyone believe they deserve another chance? After how they treated us, and all their notable competitors throughout their history?
No, I do not believe intel entering the GPU market again is a good thing for anyone but them
You may downvote now