Intel Xe-HP Graphics: Early Samples Offer 42 TFLOPs of FP32 Performance
anandtech.com
anandtech.com
Both AMD and Nvidia drivers are dumpster fires in terms of stability (and for Nvidia, I've tried both nouveau and the binary drivers)
I've had great experience with the mainline amdgpu driver, and find it ridiculous to group AMD and nvidia together in this context today. With amdgpu AMD has moved much closer to intel on the mainline linux gpu support front.
Edit: incorrect terminology
Don't give money to someone that shits on you.
I admit I haven't used amdgpu in this capacity, but just the fact that we can scrutinize amdgpu's history in mainline is a huge improvement over the past.
- Intel GPUs support mediated virtualization passthrough through GVT-G not SR-IOV
- Nvidia blocks passthrough as a feature gatekeep in the driver. You can work around this block by hiding some info from the guest (the hacks mentioned). Consumer Nvidia cards do not support SR-IOV, special models that do require the proprietary driver and a license to SR-IOV (no hacks to workaround this)
- AMD works fine under passthrough but requires specific models for SR-IOV (though this works on the official open driver). SR-IOV is still unstable.
- None of this is really relevant to the original topic
Did they ever ship a fix in mainline for their cards becoming stuck in unusable states without powering off the host? (I believe the fix involved a sequence of commands for a PSP reset)
At least I believe AMD engineers were being helpful in giving details on it though.
Mainline (let alone stable kernels) was almost unusable for like 4-5 months after release, you were completely out of luck unless you followed the right incantations and used a precise firmware binary that had been available at some point at an AMD developer's personal repo but you later had to get through someone else's Dropbox. After that period it got stable-ish for me. Occasionally Mesa/kernel updates made some games stop working or freeze the system but workarounds. Yesterday I swapped a monitor for a higher res one, and I triggered a bug where DPM essentially shits its pants and you have to switch to manual power management. Meaning a year after release the card can't handle basic multi-monitor stuff.
I've had AMD CPUs exclusively ever since Athlon II, always coupled with Nvidia cards. I wanted to support them also in GPUs given their open source work but for me it's been a complete mess. If Intel GPUs are actually any good and don't shit the bed I'll sell this card for whatever I can get.
https://gitlab.freedesktop.org/drm/amd/-/issues/929
https://gitlab.freedesktop.org/drm/amd/-/issues/892
I built a PC in June and bought a 5500XT and it's bloody unusable. It hangs if I plug a monitor, it hangs randomly, it hangs when idle, it's not a hardware defect since Windows is fine with it.
Sorry but the Linux community has been bashing NVIDIA and praising AMD because "opensource" while there are critical bugs like that one that have been open for a year. It makes my blood boil.
My opensource friendly 5500XT is on a shelf and I'm using an NVIDIA card in its place. And you better believe I will not be giving AMD my custom again.
However the 3D programming experience still beats Intel drivers.
It's a big disservice to compare them to nVidia about driver quality at this point.
Any attempt of comparison to Nvidia is nonsense, they are different universes.
Hope that their GPU driver quality won't degrade to e1000e levels, where cards freeze randomly or lose links just because they didn't find the passing packets as exciting as I do, get bored and stop working.
I once bought a pci card with Intel Gigabit. Slashdot comments were very positive about this being the best card with the best drivers. Oh wow, just looking at the kernel dmesg was infuriating. Reset upon reset, where the NIC would just be unavailable for seconds. Even a major update (I don't know, 4x to 5x or something) did not even change anything.
Currently my laptop updated to linux kernel 5.7, where Xfce and Xfwm just break down and get stuck. Going back to 5.6 makes everything work. Nope, I don't trust Intel to make good Linux drivers. Will not buy again :( I never had this in 20 years of using AMD/ATI.
On Windows, they publish WHQL certified broken drivers.
Intel has freezes in Mesa. It is a known issue.
I'm asking because Nvidia has been an absolute pain on my laptop (e.g. sleep is terribly broken), and I'm considering a switch over to the Ryzen 4700U, which has pretty powerful integrated Radeon graphics.
So I'm trying to decide if I should instead switch to i7-1065G7 instead, which has good integrated (Intel) graphics.
This is a really important question to me; I would appreciate any answers.
Switching to a text console and back clears up the issues for a few hours to a few days.
With Pop_OS! using the open-source amdgpu driver, everything worked beautifully. The only notable downside (which I don't really care about) was not having polished utilities for tweaking the GPU's behavior, and uncertainty about the status of FreeSync.
My experience on Windows 10 has been much worse. There seemed to be a 3-way argument between my monitor, GPU driver, and Windows regarding HDR settings. And unlike on Linux, an apparent bug in AMD's driver causes instability when using HDMI to stream audio to my monitor.
If Intel made a 8+ core Xeon with built-in graphics for under $1000 I would buy it over a $300 Ryzen and a $200 discrete GPU just to have stable graphics on Linux because I'm experiencing $500 worth of pain in my graphics not working on a near daily basis.
I use Nvidia with proprietary drivers, and I have only one minor problem (the windows of some open programs get corrupted on resume). I also have Intel on my laptop and have no problems at all (I had an issue in the past, but I've solved it by updating the drivers).
Nouveau is definitely a "dumpster fire", if one wants to put it that way, but that's pretty much openly caused by Nvidia.
To be clear, I don't use pretty much any 3D though, so I can only speak for "typical office usage".
Most issues fix themselves when I switch to a text console then back to X11, but every now and then I have to restart X11 to fix things.
All the problems I work on need benefit from memory bandwidth and cache latency than raw FLOPS. I imagine others are in the same boat.
I was hoping this would be the start of some more architecture diversity like apples tile based deferred rendering.
The types of workloads run on GPUs typically like very high memory bandwidth and are usually willing to live with higher memory latency to get it. Onboard GPU memory is usually built with this in mind (trade off capacity and latency for increases in bandwidth). This is generally speaking the opposite of what you want in a CPU where people often want very high memory capacity and lower latencies, but may not limited by memory bandwidth, so simply sticking a GPU on die and giving it access to a memory subsystem that was not designed to feed a GPU is not going to make anything better.
That is not the only bottleneck involved.
Historically, GPUs have used GDDR ram as opposed to general purpose DDR memory. One of the key differences between GDDR and DDR is the bus width, which can be as large as 1024 bits, compared to conventional ram with a 64 bit bus width (although dual channel is effectively 128 bits). This much wider bus results in much higher memory bandwidth which is generally necessary to feed the truly enormous number of functional units in a GPU.
I suppose you could ask: why doesn't everyone just standardize on GDDR?
1. This would dramatically increase cache line size. I don't have data, but I assume this would generally be bad.
2. My recollection (but I don't have a source for this) is that DDR has lower latency than GDDR ram, so for branchy code (which CPUs often have to deal with, but GPUs typically never have to deal with), DDR could actually be faster.
3. DDR is cheaper to manufacture. Aside from being higher volume, a lower bus width just makes is simpler to manufacture.
Why would it change cache line size? GPUs also use cache lines in the range of 32-128 byte?! I think that is independent of the bus system/width.
I just assumed that bus width = cache line size. I guess I was wrong.
Sorry.
Does this mean there are over a thousand traces between the GPU chip and the memory chips? It would be pretty clear why regular motherboards don't use it if that's the case, the sockets for the chips would be enormous! You're talking about roughly doubling the pincount vs. a 64 bit memory bus on a modern LGA socket.
But they are starved of memory bandwidth! And the lower latency memory CPUs prefer is not the same as the high bandwidth memory that GPUs like.
Also there is different kind of caching. Modern APUs often have two ways to access memory. Over there own cache, or over the CPUs cache. So shared memory for CPU<->GPU gets the full advantage of a cache, but still it is a trade off.
If you want some work done by the GPU part of an APU, by sharing a pointer, you can do that today. But from the point of view of the CPU there is no prediction beyond the "GPU do X" commands. And a very high latency until the job is done. So you need a minimum GPU job size for it to make sense.
https://fuse.wikichip.org/news/1634/hot-chips-30-intel-kaby-...
Some (not all) GPU memory types are not cache coherent with the CPU. Some of the cache-coherent cases have poorer performance relative to the non-coherent memory from the perspective of the GPU's memory bandwidth and access latency.
GDDR RAM wants to be accessed in bulk - large rows at once. Things are easy when each thread wants a subsequent byte, but if not things become much slower. Caching and other techniques can help mitigate this, and they're (imo) the place for a lot of creativity in architecture design. Having more potential FLOPS just means more ALUs.
Thank you for explaining it so clearly.
Memory bandwidth is typically the bottleneck in GPUs. Meanwhile, the access patterns are typically very predictable. So they're able to prefetch data, so latency is generally not a problem. So GPU memory is designed to have very high bandwidth, even if it means completely tanking the latency.
On the other hand, CPUs typically always need better memory latency, and most workloads do not saturate the memory bandwidth. Unlike in a GPU, many memory access patterns are unpredictable. The bottleneck of operations that are pointer heavy tends to be limited by latency. All tree operations and linked list iteration tend to be latency limited. Languages like C#, Java, Python and Javascript where all data lives behind a pointer tend to benefit significantly from improving latency. So while improvement memory bandwidth up to a point is important, there's much more attention given to latency.
Still a little pissed at how Apple handled IMG / PowerVR.
when around 2005-2006 Nvidia was hiring compiler people from Sun it was puzzling ...
Too bad it never found its niche - it was too slow for HPC, lacked virtualization for servers and was ludicrously expensive (and hot) for desktops.
Even then, it didn't work out and Intel discontinued it. It was fast enough for some tasks, but not enough popular ones to make it a market success.
- For the KNC generation, you had to compile with the Intel C/C++ compiler, which at the time cost $$$ and completely ruled out anyone who wasn't already a HPC developer or a student with access to free ICC.
- The KNC PCIe card actually made a lot of sense hardware-wise, but ran really hot and needed water cooling for a bearable desktop experience. Of course, water cooling was only available late in the KNC cycle from third parties. The cheap KNC cards Intel made available for developers were passively cooled?! and clearly an afterthought.
- The system software was abjectly horrible. The kernel driver required ridiculous patching if you weren't using the officially supported RedHat/CentOS variants, and the MPSS system software itself needed hacks and workarounds to install and configure.
- The KNC environment was just bizarre. It was Linux, so you could SSH to the card and do things, but it didn't run normal x86 binaries. The interaction between offload modes and card configuration/security was basically guaranteed to break if you didn't follow every step of several dozen in the manual exactly, which of course could be a problem when the instructions were for RedHat/CentOS and you were using a different distro.
The KNL generation was poised to solve a lot of these problems - bootable, properly supported by GCC - and then Intel gave up on the PCIe card and the only developer platform was like $20k and it was absolutely out of reach for anyone not already invested in the platform. I'm convinced that there were a large number of users who were interested in the platform, and would have been well-served by the programming model and performance, but the barrier to entry was just too high.
Nvidia got this right from basically day 1 - CUDA was available at a range of price points with (roughly) the same capabilities, so a developer could have an idea, try it out and see the benefits. Throughout almost all of CUDA's existence, for the vast majority of applications you can conclusively answer if it will work for your problem by testing on a high-end consumer GPU that costs less than $1000. Intel never understood that then, hopefully they do now.
Apparently they still don't get it.
> One Tile: 10588 GFLOPs (10.6 TF) of FP32
> NVIDIA RTX 2080: 10.07 TFLOPS - FP32
Intels iGPUs have always looked great on paper and then hit a wall in the real world.
[0] https://www.pcgamesn.com/amd/big-navi-rdna2-80-cu-rumour
Considering that, it means Intel is coming in to the dGPU market with a not stellar device, which has historically had driver issues and underdelivered in practice... Even more pressure to prove their worth then.
And that’s with one “tile”. The presentation from Intel discussed 2 and 4 tile configurations.
Wow.
I’d imagine these things are super expensive.
https://www.anandtech.com/show/15974/intels-xehpg-gpu-unveil...