The Rome is a nice chip nonetheless, I hope they can rehire a big enough Linux team to get over these humps it isn't that much work on the CPU/platform side.
The Rome is a nice chip nonetheless, I hope they can rehire a big enough Linux team to get over these humps it isn't that much work on the CPU/platform side.
On the GPU side though, at least with the modern APIs (Vulkan, D3D12 and Metal), AMD isn’t too far behind Nvidia. I actually prefer RenderDoc to Nvidias Nsight because I can capture multiple frames instead of pausing the application to look at one frame at a time. That being said though, the OpenGL tooling for AMD is abysmal and so is their OpenGL driver. All things being equal, just swapping an Nvidia card for an AMD card gives around 40-50% speed up when running OpenGL purely by getting rid of the driver overhead. Although with Vulkan AMD and Nvidia are actually on par, so at least that’s solving itself for us.
Good excuse to build another machine!
Have you tried the AMD mesa drivers? The open source driver performance doesn't look terrible, and in my own experience, it tends to work far better for OpenGL than AMD's proprietary GPU drivers.
https://www.phoronix.com/scan.php?page=article&item=rx-5700-...
I have been working with performance tools on Linux for over 10 years and occasionally debugged graphics issues.
I should note that the OpenGL stack on macOS isn't better than the Windows one either. Probably not surprising anybody, given the lackluster support that Apple has for it. We are seeing similar improvements with Metal as we do for Vulkan.
Windows OpenGL driver is so bad that running it through ANGLE with DX backend is faster.
I built a Threadripper 2 machine recently for home, but I don't know that it would have similar utility in a work environment where I could use something like Incredibuild to farm out compiles. Is that not an option for you in your work?
Mobile is also not Intel.
The only case to left are the Wintel PCs - and even that just makes Intel look bad since AMD CPUs run the games just fine. For the actual gaming experience it basically amounts to no difference at all.
I'll still need to keep some Intels around for VTune and optimizing deeply for AMD is going to be a chore... But given how the market is moving, I don't think I can avoid making AMD the primary development focus.
Ryzen < 2(!) has a much narrower AVX unit which could easily cause a 50% slowdown.
Branchy code like that of compilers tends to run somewhat slower on Ryzen, though closer to 15% than 50% (comparing AMD and Intel CPUs with otherwise same performance in mixed benchmarks). The latter really depends. Core and Ryzen have different branch predictors, both so complicated that they can't be fully described in documentation anymore. Ryzen 2's first level cache configuration also moved closer to Core's, presumably for software that has been fine-tuned for Core.
Specifically, in both a physics engine solver and some machine learning backwards passes I was working on, there are areas where a little false sharing is likely. It's relatively uncommon and a negligible concern on the smaller Intel chips I tested on, but Zen cores will sometimes have to synchronize with another CCX's caches and it seems to come with a penalty vastly larger than naively expected latency or IF bandwidth.
A fully loaded 2950x ended up getting beat by a 1700x in the solver despite having virtually the same architecture, twice the memory bandwidth, and an observed clocks advantage. Cutting the used thread count down to the same as the 1700x helped a little, but it looked like Windows was scheduling the threads on as many CCXs as possible and it ended up still being slower.
On the upside, the 2950x blasts through friendlier workloads like collision detection and inference with perfectly reasonable scaling.
I'm hoping that the dramatic redesign in Zen 2's memory architecture (unified IO die and whatnot) will help things a little. If not, I'll probably have to rework some stuff.
I use ECC on a (consumer) Ryzen chip/board and edac-util seems to give me the same information that it does on Intel - what's missing?
One the CPU/platform side of things, my biggest annoyance is how far behind k10temp is (Zen2 support not in mainline until 5.4?), and how bad sensors support is in general on the boards (requiring reverse engineered non-mainline modules for my Zen/Zen2 workstations).
While I agree that on the GPGPU-front Nvidia is still ahead, I'm very happy these days on my workstations with the state of AMDGPU in the mainline kernels and much prefer AMD GPUs for my workstations now vs Nvidia cards, which has led me to pay a bit of attention to ROCm - it looks like they are making very steady progress and TF and PyTorch support seems pretty good at this point (also, stuff like MIVisionX/OpenVX, CenterNet, BERT support seem to all be working relatively painlessly [3]), although it'd be useful if anyone has a resource that does continual benchmarking comparing on-prem/cloud perf of the various platforms, it'd be nice to get good $/perf and W/perf numbers over time. (I assume that anything running with Nvidia's tensor cores still completely blows away what AMD has to offer atm).
[1] https://developer.amd.com/amd-uprof/
[2] https://github.com/RadeonOpenCompute/ROCm/commits/master/REA...
> I use ECC on a (consumer) Ryzen chip/board and edac-util seems to give me the same information that it does on Intel - what's missing?
Event-based sampling on Intel is accurate to an instruction-level (while event-based sampling on AMD is less accurate. You're forced to use the more complicated IBS metrics if you want instruction-level accuracy of events).
Intel also has branch-history data stored. Super useful for some developer tools, but I forget which tools those were...
-----------
I think AMD uProf is certainly usable. And the price is good (free). But Intel vTune is just light-years ahead.
AMD vs CUDA on the other hand is... closer than I think most people realize. CUDA has a bunch of libraries (Thrust, TensorFlow support, etc. etc.) which helps. But if you're doing high-performance coding, you'll likely have to write your own specialized data-structures. At least, that's the approach I'm doing with some GPU hobby code I'm writing.
TensorFlow (due to Tensorcores) and BLAS are solidly NVidia advantages. But general purpose libraries (ex: Thrust) is more of a convenience.
AMD's main disadvantage is documentation. But the tools are actually quite usable. AMD documents the lowest level well (the ISA), but their HIP / HCC / etc. etc. documents are lacking and difficult for beginners to follow.
AMD should work on updating their beginner guides (their OpenCL guides) to their ROCm framework. Even if its ROCm OpenCL 2.0 stuff, its important to get beginners to use their platform. Or at least, update their beginner guides to reference GPUs that have come out within the past 5 years...
Sadly also the third-party hardware ecosystem is (at the moment) pretty bad.
For example you can choose among almost 300 different motherboards with an Intel 1151v2 Socket [1]. They come in all possible sorts and combinations of form factor, chipsets, ports, etc. In comparison there are less than 100 motherboards with an AMD AM4 Socket [2], most of which in the large ATX form factor and with basic I/O ports.
Let's hope that all these third-party companies felt the change of the tide and are hard at work on new AMD-based products.
[1] https://geizhals.de/?cat=mbp4_1151v2 [2] https://geizhals.de/?cat=mbam4
[0] - https://www.macrotrends.net/stocks/charts/AMD/amd/revenue
As an aside: why won't everyone migrate to Vulkan compute shaders? I hate these "special" compute stacks >_< Clearly I'm not alone in thinking this: Tencent's ncnn uses Vulkan as the only GPU option, some Googler is working on clspv (OpenCL to Vulkan SPIR-V compiler)..
1. No libraries. cudnn, cublas, cufft, are all the fastest available (except maybe magma sometimes), plus no one writing an actual application wants to reinvent a fast gemm. Also cutlass, cub, thrust, ...
2. No c++. The "standard" seems to be glsl, and a "prototype" opencl c -> spir-v compiler doesn't give me much confidence in that approach.
3. No one wants to use vulkan apis directly, and there are approximately 5 billion different "vulkan compute" wrapper utility libraries. i.e. No consistent platform.
Vulkan compute and OpenCL are not entirely compatible (both are backed by SPIR-V (but different flavors of it.)) Khronos has chosen to maintain OpenCL (and OpenCL-next) separate from Vulkan.