Ryzen Threadripper Pro 3995WX Spotted
guru3d.com
guru3d.com
Am I wrong about that being a temporary problem (that would feel frustrating)? Are these cores just that much faster than CUDA cores? What else am I missing here?
I often hear people talk about getting CPUs like this for deep learning research, but all the deep learning work I’ve done goes straight to CUDA and lands on my GPU.
I'd be interested to know what workloads truly require that; most simulation solvers that I can think of that require large amounts of RAM are working on a discretized spatial grid, and the physical laws they're simulating are fairly "local"; e.g. you can process one area then move on to the next, in a nice stencil operation. This lends itself well to both parallelization and a sort of data-locality, so you can stream the necessary pieces of data in to the cores in an orderly fashion. I don't know what kinds of massively parallelizable workloads require much more random access of memory.
See https://wiki.openjdk.java.net/display/HotSpot/CompressedOops
Thread divergence.
CUDA-cores are linked and take if-statements and for-loops together. This means that if one CUDA core takes a for-loop 1-million times, 31-other CUDA cores will take the 1-million loop with them. (32-cores per NVidia SM). EDIT: CUDA keeps things semantically correct per thread by "throwing away work" (execution mask is disabled, so those 31-other threads effectively execute NOPs), but the cores are wasted.
Worst case scenario, your CUDA cores accomplish 1/32th the work they could do, as they spin idle waiting for the last 1 thread to complete a for loop or its unique combination of if-statements.
Matrix multiplication doesn't have any thread divergence, because all threads loop the same amount of times on all workitems. This means that Tensors / Deep Learning simply ignores the problem, because they're written as matrix multiplication problems.
Avoiding thread-divergence, or mitigating its effects, is possible, but requires advanced programming skill that few have.
--------
Examples:
Chess -- Traditional multithreaded chess algorithms have every thread check out a different branch of the chess search tree. But because each branch has different positions and moves, its very difficult to code it in such a way that all threads are doing useful work.
Web Servers -- If every thread is handling a different request, then every thread probably traverses a different set of if-statements and loops. This is a thread-divergence nightmare.
Databases -- SQL Databases traditionally have each thread running a different, independent SQL query. Like web-servers, its difficult to imagine parallelism from a high-level.
-------
However, elements of Chess, Web-servers, and Databases can probably be efficiently parallelized onto GPUs.
Chess -- This github repo demonstrates that enumerating positions can be done GPU-parallel: https://github.com/ankan-ban/perft_gpu
Web-servers -- Text parsing, Regex, and many other problems common to Web-servers have been parallelized on GPUs. The main question for me, is whether or not the PCIe traversal would be worthwhile (lower bandwidth than RAM, high latency). Web-servers tend to be I/O heavy and less compute heavy... probably not dense enough to benefit from GPU compute.
Databases -- Merge-join and Hash-join probably can be parallelized to GPUs.
Figuring out the optimal CPU + GPU team is going to be a research problem over the next 10, maybe 20 years.
Many other replies speak of the difference between “GPU” and “CPU” without describing it—it’s thread divergence. Processing stream-like data (network streams; compilers and other parsers; human input devices) can branch often and unpredictably using normal algorithms. Think for example how you would tokenize JSON without branching. A single GPU core is quite slow for the general case anyway. They are mostly good at math and bitwise operations. They were able to get this way because of the assumption that there would not be thread divergence for the primary workload, i.e. matrix math on contiguous blocks of memory, which harkens back to their origin as actual Graphics Processing Units.
If thread divergence weren’t such a killer for branching programs, you could probably write anything in CUDA. It’s not very limited; you can write most anything that can be expressed in non-exotic C. A lot the work has to do with schlepping textures around, which is an artifact of the expectation that your code will run at a distance from the messy, branching, unpredictable main program state, so everything (textures, shaders, etc.) should be loaded, processed, and ready to go when it’s frame-buffer time.
These days, arithmetic is cheap and control flow is expensive. GPUs and SIMD achieve high throughput by tying a set of arithmetic units together with the same control flow. As soon as you get complicated data-dependent workflows - as soon as you put "if" in your code or a virtual method - the GPU starts getting less efficient.
It's very much a "train vs car" argument: a train can deliver a large number of execution units , provided they're all going to the same place, whereas a car can re-route far more easily.
Its quite possible that someone figures out an x86 emulator on GPUs that fuzzes register values, similar to AVX512 fuzzing. https://gamozolabs.github.io/fuzzing/2018/10/14/vectorized_e...
------
Fuzzing is unique in that 99% of code probably traverses the same set of if-statements or branches. Whenever code significantly diverges (usually defined by Fuzzers as a significantly different instruction-pointer path of some kind), you've successfully fuzzed a new result.
Fuzzers... strangely enough... are programs I'd expect to run pretty well on GPUs or SIMD. If enough R&D research effort were put into it.
Gentoo is actually a usable distro again and most of the packages (even things like mono) install decently fast.
I do not personally, but I work with another team that uses tens of thousands of cores for calculations. Just day to day business.
There's all kinds of other things I do that utilize multiple cores.
I also run a lot of VMs.
Also, year-on-year performance increases for nVidia GPUs are relatively boring, while Threadrippers routinely double in performance. I have just built a 3970X/RX5700XT workstation, and while I have not benchmarked it yet I expect the CPU to smash the GPU for Cycles raytracing.
Also worth noting that CUDA is a proprietary architecture from a single company - targeting it locks you in, and you need to run proprietary blobs to make it work.
All in all, CUDA works well now but I think it's ultimately a dead-end architecture - hugely parallel CPU architectures will eclipse it.
Raytracing is not an ideal match workload wise for a GPU, too much divergence, but there are plenty of workloads where GPUs will have an edge for along time to come.
AMD cards offer comparable compute performance and are sometimes even better. What AMD lacks, though, are ready-to-use software packages like CUDA, cuDNN, cuBLAS, and the like.
NVIDIA basically bought their way into the research community and solidified their strong foothold by keeping CUDA proprietary.
It's fascinating how other companies (Microsoft) get bashed on a regular basis for using the exact same tactics (see Direct3D vs OpenGL on Windows), while people are perfectly fine with NVIDIA doing it...
This is really important, and I don't think people quite realise the scale. I was doing a Physics MSc in 2012 and it seemed that almost everybody in computational physics was doing some variant of reimplementing Y with CUDA.
That's not to say it was the wrong tool for many of the jobs – though for some it apparently was, judging by the levels of stress – but I did find it a bit weird how this was an approach to research that pretty much completely ignored vendor lock-in.
I assure you, the affected scientists who have to fight tooth and nail for funding are less than impressed that they are dependent on the mercy of a single vendor.
On the other hand, wasn't GPU mode added much later than CPU mode? Maybe there has been less development time put into it.
The CPU is never faster than the GPU. CPU + CUDA is way faster, but when I use OptiX (RTX) it is unbelievable how fast it is.
The GPU Raytracer just shines here.
The Blender benchmarks: https://opendata.blender.org/
Maybe that time will come but I am not sure if that will happen soon.
There's quite considerable variation in the device ranking if you look at each scene render time - the GPU is faster on the quicker-to-render (sub 1 min) scenes, but slower on the slower-to-render scenes (1+ min).
I wonder why that is? I see the 3990X is hugely faster than anything at these slow scenes, which is pretty interesting.
I'm not familiar with rendering, but in the deep learning world it is moderately easy to get considerable speed-ups by using multiple GPUs. Is the same possible here?
So it all depends in what is being rendered and how a render core is optimized for it.
Also things like b-tree calculations and other algorithms might influence this.
Column stores like SAP HANA. Compressed bitmap indexes leverage SIMD but not floating point matrix math GPUs. The workstation vs server case is narrower though; benchmarking, dev testing, and local MS PowerPivot like analysis on large in-memory datasets.
* I'm splitting data sets into batches of around 50-300MB, but each individual batch must be processed sequentially. Each batch loops through the data (multiple times), and every single iteration of a loop depends on state from previous iterations.
* Altough each batch essentially does the same thing, it does not operate in lockstep. Lot's of ifs, different execution paths, differently sized sub-problems, etc.
I'd benefit from having 200 CPU cores and enough RAM to fit the batches, but I can't let individual GPU threads run through batches of 50MB, each.
It's the ads though, and the calibre of people who create ads...
Not only that, but better performance is what pushes you over the edge, not popups, connection slow downs, auto playing videos or tracking?
The high end renderers usually split up a large image into "buckets" and then perform path tracing on those regions of an image, and combine the results together. If you have more cores, you can increase the number of simultaneous buckets to evaluate.
To put this in perspective, a lot of the recent news around real-time raytracing is when a game engine does 1-4 samples per pixel and then aggressively denoises (in 2D) the result. That's why you often see splotches or errors. On the other hand, high end VFX renders will end up doing hundreds or even thousands of samples per pixel to get a clean, high quality, physically plausible result.
You can run on multiple nodes, but that brings its challenges, see Jepsen.
And with such a huge CPU core counts, a lot of the very difficult work of organizing an algorithm to run well on a GPU can be avoided because the CPU approaches the same benefits with a more logical architectural design.
A closer analogue of x86 cores is what Nvidia calls a "Streaming Multiprocessor" and AMD calls a "Compute Unit". There's generally only up to 64-80 of those on the highest end GPUs too. However the difference is that each of those cores is able to perform 40-64 SIMD operations at once, compared to x86's puny 8 or 16 or so.
Building software, for one. C compilers and python interpreters don't run on a GPU. Lots of stuff doesn't run on a GPU. In fact in practice the only things that run on a GPU are the tiny handful of known subproblems that the industry has collectively decided are "GPU problems".
Like it or not general purpose scalar software is, has always been, and will always remain the standard mechanism by which computing hardware is applied to new problems. Everything else is an optimization around the edges.
Edit: another reply to this top level comment: https://news.ycombinator.com/item?id=23788755
pattern like the one below will execute up to 10 times faster if array is sorted due to data locality and branch prediction
sum: float
a: float array[some_huge_number]
fill_array(a, random(1))
for f in a
if a[i]>0.5
sum=sum+fIMHO this is an area ripe for more exploration. If I were working on it, I might look to linking first before compilation, because the basic link task is more similar to what GPUs are good at (advanced stuff such as LTO is a different story, though).
This is incorrect.
The other issue is loading that memory - PCIe 4 is still the transfer time bottleneck between GPU and main memory.
My understanding is that the GPU memory models are different enough that what an OS traditionally calls "virtual memory" couldn't be implemented in the same way.
Well, the problem with that is that no make system can utilise so much core. Even linux kernel, probably the biggest C project in mainstream use, can not consistently load even 16 cores with mostly handwritten makefiles.
P.S.
I do not say that building software is not CPU parallelisable in principles, I'm saying that regular make systems have trouble handing big number of parallel tasks in practice, even in big, well run projects that use handwritten makefiles.
And my day job involves lots of iterations over the Zephyr test suite, which builds and runs hundreds of individual test apps X dozens of platforms. This is almost a pure build-time problem and scales very linearly indeed.
https://stanford.edu/~sadjad/gg-paper.pdf
Actual screencast of it in action compiling FFMPEG:
https://asciinema.org/a/257545
(Perhaps the issue is you are bottlenecking on I/O or some other resource?)
It is true that there are tasks where threading matters, but still require a CPU rather than a GPU. I wonder however if these tasks do need full SSE/AVX etc. Couldn't these extensions be removed of the CPU cores and instead have the necessary work performed by the GPU?
It would be interesting to produce statistics on how much these extensions are used in these scenario. Imagine how much space and complexity could be saved on a CPU die by making stripped down versions. That space could in turn be used for more cores!
I read a little about the Xeon PHI cpus, which iirc, is a multicore CPU with a very small ISA, but I wonder why x86 makers aren't trying to go in that direction: isn't there plenty of dedicated workloads which would happily run on these (eg, web servers), or is this just a (too) simplistic view?
SSE/AVX shares an L1 cache that's damn near instantaneous to access for the CPU core. Total L1 bandwidth is on the scale of TB/s.
PCIe -> GPU takes 1-microsecond to 10-microseconds per access, and operates only at 50GB/s (or 1/20th the speed of L1 bandwidths).
------------
Case in point: Memset is very commonly AVX'd to clear out L1 cache and initialize ~1kb to 32kb of data to 0 as quickly as possible.
There's no way for "memset" to move from CPU to GPU unless you feel like obliterating the entire point of L1, L2, and L3 cache. If you moved a "memset" to GPU, it'd operate only at 15GB/s (the speed of PCIe 3.0 x16 lanes), far, far slower than L1 cache AVX-loads/stores.
SIMD units, like SSE and AVX, are highly "local" and have huge advantages.
The main problem is, afaik, that there is not enough control about where the code will run in these languages. At some point, one will want to describe all the algorithms using a single language, and somehow describe how the workload will have to be distributed across all the processors, or at least that's what I've been thinking about for a while. Once you have that level of control, the need for a versatile CPU is less clear. Note that nowadays people seems happy with hybrid solutions where the code is scattered across several languages (eg, one for the main program and one for the shaders, or for the client side UI), so my position is maybe not very strong.
HW-wise, is it possible that integrated GPUs are the first steps toward an architecture where CPU and GPU have better interconnections (ie, larger communication bandwidth and smaller latency) to the point where SIMD becomes moot? There is also the SWAR approach, where one doesn't rely on intrinsic SIMD instructions, but instead emulate them (though it's probably not very realistic for floating point computation).
Some other ideas:
- Apple has this neural engine in their latest chips, which is basically dedicated HW for neural networks
- In the wild, people are getting more and more interested in building their custom ASICs to cut software's middle-man cost: for them, the CPU solution is not good enough
- Intel recently introduced a new matrix ops extension in their CPUs: maybe at some point they'll introduce full GPU capabilities directly baked in the CPU? I am a little worried about the resulting ISA.
Anyway, I am not an HW engineer, nor a very good software one. I only have a limited view of the difficulties in writing good, CPU or GPU efficient code. My first post was prompted by remembering the first "large scale" multicores CPUs 15 years ago (specifically the Ultrasparc T1) which wheren't SIMD heavy. The direction naturally shifted as progress was made on SIMD to try to compete with GPUs, when it seems to me that originally CPUs and GPUs were complementary.
I tend to support modular solutions, but I don't know how costly that would be in term of efficiency at the HW level.
Xeon PHI on the other hand was the first host of AVX-512 instruction set. Sorry.
With a GPU in the loop, there is an extremely painful hop across PCIe before you can experience any degree of cache coherency. There are protocols, buffering and driver stacks involved. Latency is much higher. You have to plan ahead and think about what memory the GPU needs to mutate vs what memory the CPU needs to mutate. You have to split your application, language, and frameworks across 2 computational domains. How we have tolerated this for so long is beyond me. I think most people who get caught up in it are chasing shiny marketing a lot of the time. This criticism aside, I do think GPUs are still very useful for many applications and users.
With just a CPU, you are looking at L3 as the typical worst case for cache coherency domain. And L3 is getting to be ridiculously large. If your high-performance application's binary image cannot entirely fit within L3 of a modern x86 CPU, you are probably doing something very wrong. The other big advantage with just a CPU is you have 1 cache-coherent memory domain and a single instruction set to answer to. This means you can write 100% of your software in a single language/framework. Well-architected x86 applications can push instructions many times faster than their core clock speed would seem to indicate possible. Using special frameworks that have sympathy with this memory model mean that just a single x86 core can be made to push tens of millions of logical business transactions per second. And then you have 63 other cores to play with. Extremely deep pipelining and OoO execution are what allow for x86 to chew through general purpose computation so easily. Anything with a lot of recursive depth (e.g. raytracing) benefits massively from this style of processor.
I think the next big revolution is developers realizing that they can start porting GPU applications to CPU. Things like raytracing are much better suited for the memory model offered by a pure x86 domain. Sure, a GPU can accelerate some aspects of raytracing, but then you then need to shuttle that information back and forth over a higher latency link, and also fight with 2 completely different technical stacks. Needing to know how to write code for both GPU and CPU is easily the biggest problem with all of this. Especially when you consider how fast GPU APIs move relative to x86. Performance is not the only constraint when dealing with software, especially software as complex as AAA game engines and 3d modeling software. Someone still has to reason with and maintain this stuff over time.
The next UK supercomputer is based around AMD's 64 CPU chips for this reason. One concern is that we may now be limited by the shared memory bandwidth. https://www.nextplatform.com/2019/10/18/amd-cpus-will-power-...
Servers serving multiple independent requests, VMs or Docker containers.
In other words, the same thing people were doing with four 16-core servers, just now it takes up less space :)
Not only that, huge amounts of software could use a lot more concurrency if it was structured differently. Even a web browser (to my knowledge) doesn't split up DOM layout, parsing individual files, decompressing images, decompressing video etc. into separate threads for each file.
Comparing GPU/CUDA/TPU Cores to CPU Cores is like comparing a riding lawnmower to an automobile. A GPU/CUDA/TPU Core is optimized to handle a different type of workload than a CPU.
Just because your riding lawnmower can cut your grass faster than your car, doesn't mean your lawnmower is the best thing for highway driving.
There are other area where big general purpose X64 or ARM cores excel too, like bit twiddling, memory random access, etc., but branching logic is the fundamental one.
Think of two agents that run for awhile and end up in different parts of a virtual world, that require very different code to execute. This might be pretty difficult to parallelize on a GPU.
After collecting a bunch of experience, you update whatever function you’re optimizing. This could be something like a neural network that is the brain of your agents. The update step can happen on a GPU.
But then we need to go collect a lot more experience again, so we’re back in CPU land.
Experience collection often dominates overall compute time in reinforcement learning. Which means you want a lot of CPUs.
I'm a bit behind on GPU architecture, but presumably workloads which are branch-heavy, lack coherent execution within a warp (see dragontamer's comment), and don't make use of floating-point.
GPUs are not the equivalent of manycore CPUs.
This indicates 3995WX supports not only 8ch but also RDIMM/LRDIMM that's not supported on current Threadripper. It should need a new Socket rather than current TRX40 (TRX80 was rumored a year ago). Threadripper gets closer to EPYC.
I expect PCIe lanes limitation still remains for make difference (It's many even limited).
We've got a weird use case, we're using (black box) software that's sensitive to clock speed, scaling only to about 8 threads until it starts levelling off aggressively (to the point where a 3970X benchmarks as faster than a 3990X). So what we do is we run 4 instances, each with their own GPU and 8 cores assigned to it. I'd be surprised if this makes sense anywhere else (as EPYC has all sorts of other advantages in HPC).
For others, I believe this is the board in question: https://www.asrock.com/asrockrack/general/productdetail.asp?...
But if you want to simulate production workloads on your machine say with Apache Spark or Flink, you are in the prime. You couple that CPU with 256+ GB memory, and you can start exploring the bottlenecks in your batch algorithms using as many cores as you want.
[1] https://www.pugetsystems.com/labs/articles/After-Effects-CPU...
Massively parallel operations take a different focus on your MPI/bus/queue architecture than we've ever needed for desktop software. BUT I think what we've already seen is that in general purpose computing there will always be someone ready to consume resources as they are available.
Both feet don't move at the same time, one takes a step and then the other.
It would really be interesting to see if it's feasible to Overclock the best numa node to a degree that gives this a couple of high single thread performance cores comparable to the 3900X or 3950X while retaining stability overall.
Do you mean the best ccx? Or best chiplet? There aren't any numa nodes here, and the cache & memory architecture are already the same as a 3900x.
Per-core & per-ccx overclocking does already exist on ryzen though, should work the same on threadripper.
It would be interesting to turn the 3990/3995 into a does-it-all chip with best in class single-core performance on certain cores and just lots of a tad slower cores in general.
These are the cores that can turbo the highest.
You can attempt to manually OC higher, but that by & large doesn't work without extreme cooling. Typically instead it's about achieving higher all-core frequencies than the built in all-core turbo, typically via higher voltages & power limits (and of course much better-than-stock cooling)
Indeed if you look at single-thread cinebench numbers you'll find the 3960x & 3970x right in the middle of the 3700x, 3800x, and 3900x pack: https://www.guru3d.com/articles_pages/amd_ryzen_threadripper...
They all hit around that same 4.5ghz single core turbo mark without any overclocking. And Zen2 by & large doesn't really go much beyond 4.5ghz anyway, so there's not all that much manual overclocking you can do without going sub-ambient cooling anyway.
The only real downside to Threadripper 3rd gen is the price. It's already a pretty killer jack-of-all-trades CPUs otherwise. It's a very competent gaming CPU without any tweaks at all right out of the box. It doesn't at all have the cons of the 1st & 2nd gen Threadrippers, which were actually NUMA and therefore came with huge gaming downsides.
Your 4 channels with 2 slots each have the 4 channels daisy chained to 2 slots, but they are just 4 channels. With 8 channels and 2 slots per channel you can have 16 slots, with 32 GB unbuffered (regular) DIMMs you are limited to 512 GB of RAM, for 2 TB you need registered or load reduced RAM.
Nearly every consumer CPU for the last 20 years is dual channel, meaning "2 drive RAID 0". If you've seen things like recommendations to get paired DDR memory sticks, this is why. You only get this RAID-like benefit with multiple RAM sticks (just like RAID 0 of a single drive doesn't do anything). It's also why some consumer products have unexpectedly bad performance for the CPU specs - cheaper laptops may skimp here and just run a single stick of memory instead of 2. Which means half the memory bandwidth.
8 channel then means the CPU can do this up to 8 sticks for 8x the bandwidth instead of the more common 2x.
that is like have 32 cores on a normal 2 channel desktop, throughput would often be a limiter there.
The real reason for adding more memory channels to a general purpose CPU is actually to increase the maximum amount of memory per node. A few years ago, workstation and server platforms often allowed 3 or more DIMMs per channel with performance declining sharply with memory density. Since then both AMD and Intel have limited each channel to 2 DIMMs and increase the number of memory channels from 4 to 6/8 to make up for the loss.
While more memory channels do enable better performance scaling, it has also made the CPU socket more complex and fragile than ever: The TRX40 socket has 4094 pins and literally requires the CPU to be bolted down using a torque wrench to ensure good contact.
AMD use this term themselves.
I'd expect closer to $4000 or more.
> is this for the consumer line or server line of products?
Neither or both. Threadripper has high-clocks (like consumer) and high-core counts (like servers). It'd be for "workstations".
And yeah there is a bit of a gap there. You could always get a low end EPYC but then you miss out on the turbos.
It's somewhere between the extreme enthusiast market and the workstation market. It's not _truly_ a workstation product as those will instead typically use Xeon or Epyc for things like full official ECC support & other such features. Threadripper muddies this line a bit as it does actually have official ECC support, but motherboard support is not as reliable as you'll find in Epyc boards.
Intel also has products in this category like the i9-10980XE - which is an 18-core product that's essentially the consumer version of a Xeon chip with ECC removed. Intel calls this their X-series line. And then since Intel removes ECC support, you'll again not really find this in workstations.
It's basically "take the server product, remove (some/most) of the server features, jack the power budget to the fucking roof because you can assume water cooling, and allow overclocking" line of products.