Cores of Rendering Madness: The AMD Threadripper Pro 3995WX Review
anandtech.com
anandtech.com
I'm amazed how Intel has lost their crown. They don't have the best servers, HEDT and consumer desktop cpus. And if they are competitive they consume significantly more electricity
Plus on laptops AMD has more cores and consumes less power whilst Intel touts how their CPUs beats AMD on AI photo upscaling.
Not to mention how Apple is in the process of ditching them
I wonder if the exec/upper management levels still pay themselves well? (of course they do)
Its hard to stress just how useless AMD have been for most of the last decade barring the Zens.
Companies like VMware moved to per core licensing. And most have established infra in intel. Cycling that out is much easier said than done. Especially with the capex squeezes coming.
Some of the AMD chips don't support SMT, so maybe this statement doesn't tell the whole story. I recently looked at the Lenovo ideaPad Flex 5 that was offered in an AMD flavor with Ryzen™ 5 4500U or an Intel flavor with Core™ i5-1135G7. The AMD chip is 6 core / 6 thread, and the Intel chip is 4 core / 8 thread.
I'm supposed to care about the number of threads rather than the number of cores, n'est-ce pas?
In my case it turned out to be a moot point because I couldn't get the AMD version anyway; it was available on Amazon from a third-party retailer but that's always sketchy and it was more expensive. But I think the "AMD laptops have more cores" argument deserves an asterisk because that's not really the whole story. (And to your point, having more cores doesn't necessarily guarantee superior execution if there are other aspects of the computation structure that bottleneck your workload.)
https://www.youtube.com/watch?v=uKzJNIs8Pnw
I'm sure I could find a limit if I tried made-up tasks meant to do that, but it's plenty for all the real-world things I use it for.
It happens. One process mistake cascaded and derailed their whole lineup.
AMD doesn't have to bet their future on process and can, if needed, move to a different supplier. If, however, they make an architecture mistake and end up in a dead-end, Intel will be able to catch up in one or two generations.
Intel still has a significant lead in several important segments and these deep supply chains take a long time (and a lot of money) to shift. AMD has been showing a clear leadership in cost/performance for some time now and it's still not easy to fill a rack with AMD boxes.
Conversely if you counted a GPUs independent parallel threads of execution like on CPUs, you'd max out at a very comparable 64 "cores".
If you count ALUs, you see that the CPU has many of them, but you don't see how difficult it is to chunk up data to keep those fed.
If you count independent threads, you see that the GPU has few of them, but you don't see how it conceptually has many threads, which simplifies the programming of each thread while gracefully degrading only in proportion to how much branching you actually use.
I realize that under the covers this is more of a compiler/language thing than a compute model thing, but for whatever reason I just don't see much SIMT code targeting CPUs, so again, human factor.
Yes, "CUDA Core" is (was?) analogous to "FP32 ALU" - (it may have transitioned to FP16 nowadays to double the count?)
Yes, EPYC has 16 FP32 ALU per Core.
But, single GPU "Compute Units" often have 2-4x that (up-to 64 FP32 ALU), and a lot of other differences which when added up at the whole GPU level don't mean they max out anywhere near '64 "cores"' "if you counted a GPUs independent parallel threads of execution like on CPUs"
That's where they couldn't be more different!
Also, for accuracy, vis-a-vis "independent parallel threads of execution like on CPUs", even the 64-core Threadripper has 128, not 64.
And many GPUs with 64 physical "Compute Units" would have at least 640 by the same measure (since many have 10 independent programs which can be resident in hardware at once, compared to SMT2 on Threadripper).
But when you break it down further, while a model GPU Compute Unit from a few years ago might have 64x FP32 ALUs in hardware, that CU has the capacity to schedule and execute single instructions on 10*64 logical threads (= 640 logical threaded instructions in-flight on a single CU at once). On a 64CU chip, that's 40,960 logical floating point instructions in flight at once, each one of these coming from a different logical thread of execution in a SIMT model.
The superscalar CPU cores can also have lots of instructions in flight, but they are more deeply pipelining instructions for fewer threads of execution (focused on single-threaded performance).
This is all a long way of saying that a.) you may not have been so far off when comparing raw ALU counts; but, b.) you couldn't be misrepresenting the facts more when comparing the differences in "parallel threads of execution."
This is a core architectural difference between CPUs and GPUs, and while yes, there are similarities in ALU count, the way the transistors and ALUs are utilized by software is quite different.
To oversimplify, and ignoring performance of applying these different programming models to different compute architectures:
Using an example 64-CU GPU compared to 64-Core EPYC:
64 CU GPU as SIMD: 640 Threads (SMT10) and 4096 FP32 ALUs (SIMD64 - actually 4xSIMD16)
64 Core CPU as SIMD: 128 Threads (SMT2) and 1024 FP32 ALUs (SIMD16 - actually 2xSIMD8)
64 CU GPU as SIMT: 40,960 Threads (10:1 Threads:ALUs) 64 Core CPU as SIMT: 2,048 Threads (2:1 Threads:ALUs)
So, this model 64-CU GPU can schedule 20x more logical hardware threads in a SIMT programming model than the 64-Core EPYC CPU, despite having only 4x the FP32 ALUs.
My final caveat would be that GPUs are over time increasing single-threaded performance, and CPUs are becoming wider. So, in a way, there is some architectural convergence - but it's not to be overstated.
One last note: Niagara was a bit GPU-like, with round-robin thread scheduling, and POWER now has SMT8. But differences remain.
> Yes, "CUDA Core" is (was?) analogous to "FP32 ALU" - (it may have transitioned to FP16 nowadays to double the count?)
It's still FP32 ALUs (NVidia disables FP16 in the driver for gaming cards). The doubling between Turing and Ampere is due to the combined int/fp units being counted as CUDA cores. This also means that for int instructions, the expexted performance is not actually increased, another reason the "core" term is unhelpful.
> But, single GPU "Compute Units" often have 2-4x that (up-to 64 FP32 ALU)
Yep, guess I should have left that AVX2048 joke in there for you. It just wasn't very important IMO.
> Also, for accuracy, vis-a-vis "independent parallel threads of execution like on CPUs", even the 64-core Threadripper has 128, not 64.
I did not consider SMT here because while relevant for keeping the ALUs fed with instructions, they don't increase the amount of raw peak compute possible. Those 128 threads can still only issue 64 avx instructions at any one time, which is what really matters here.
I'm willing to pay $4000+ for my profession's tools. But when it comes to a hobby, cheap or free is better.
I like Krita (https://krita.org) as a painting tool. Darktable (https://www.darktable.org) does photo editing, although I haven't used it myself.
Edit: Guarantee not a single person downvoting this is an audio video or photography professional. This attitude of OSS above all, damn the users is likely part of the reason there aren't more viable creative apps for Linux, and more viable OSS creative apps overall.
And as mentioned, Affinity/Adobe software is not available on Linux, which immediately disqualifies it for me personally, even though I'd like to upgrade past GIMP and I'm very much used to Photoshop after years of using it.
At a certain point the old HP (before they where shit) Joke if you have to ask the price you cant afford it still applies.
Lets say I'm a programmer who needs to spin up multiple VMs / Docker / etc. etc. to properly emulate some server setup. Multiple cores and tons of RAM (more than 256GB) needed on a regular basis, with a few NVMe drives in RAID0 so that those VMs can be snapshot in blazing fast speeds.
Paying $100 to $4000 for VMWare workstation, Visual Studio subscriptions, SQLServer non-express, Red Hat Virtual Machine License, Visio, and other such top-class tools would be 100% reasonable, along with the $10,000+ Threadripper pro computer with 512GB RAM or whatever.
But image editing wouldn't be part of that user's workflow. Suddenly, the professional developer is asked to make a web-banner on the top of his Github document with the logo of the company on it.
Guess what? Said developer probably is downloading GIMP. The hassle of going to corporate to be compensated for lol $25 on another image editor is just not worth the time.
As for the wasteland of usable OSS media applications, I think that's hyperbolic, but I agree the interfaces are often surprisingly poor in relation to the technical capabilities.
In terms of the available apps, I think you may be blind to just how great production and editing apps are outside the OSS space for everything from audio production to vector design. Literally in terms of great OSS design / media apps there's VLC for consumption, and Blender. That's it.
There is no usable OSS alternative with comparable functionality (let alone usability) in most of the categories used by professional (or even enthusiastic amateur) for most creative work. Those apps that do exist are decades out of date in terms of features and usability. It's as though the evolution of design language and user interface that took place after the release of the iphone never happened.
Here's a simple analogy. Imagine an operating system where the only IDE was notepad or its equivalent. Writers could point at that and say, whats the issue it meets all my needs, you can absolutely code in it, its uses almost no memory, supports unicode etc etc. Programmers would rightfully find it completely inadequate because it's not suited to their much more specific, rich, evolved usecase.
Similarly there's a complete blindness in OSS circles about how much better design / creative apps are in the rest of the computing world, and how little evolution there has been on these platforms in the OSS world, and how completely the available options fail to offer any of the functionality necessitated by a modern creative workflow for image creation, audio creation, video creation etc.
For example? As a Linux user I would love to pay for an intuitive image editor. Unlike what you suggest the problem is not the OSS attitude, it's that there's no alternative.
Either way Gimp works well enough for my purposes. The UI could be a bit more robust, but it's not nearly as bad as some people make it out to be. For the sake of my blood pressure I'd rather use Gimp a day than MS Office for half an hour. At least Gimp's native file format is deterministic.
In terms of commercial alternatives available on Linux, your options are limited for the same reason that so many professional level graphics programmes are not available on Android. Linux users spend enormously less per desk on software, the balkanisation of the platform makes support difficult or impossible. Worse, unlike Android in the mobile space, Linux has so small a percentage of the desktop market that it likely isn't worth developing for (for commercial creative applications).
If you think this situation is the fault of commercial app developers, you're exhibiting exactly the OSS snob attitude I was criticising.
Kind of wondering what kind of image would be 30GB... :)
* https://openslide.org/formats/aperio/
* https://web.archive.org/web/20120420105738/http://www.aperio...
Pages 14 onward give the info. Seems like it's a variation of TIFF, in tiled format.
Main tile (image) size is 120k x 80k. That's huge!
29GB when saved uncompressed. I'd hope most applications would use the optional compression features. :)
Maybe Gimp has just rotted my brain over the years with its UI...
So I can see going from Gimp to a more standard interface being difficult, if Gimp is what you're used to. Conversely so many people hate using Gimp because of the same reason if they're used to Photoshop
On top of that, I downloaded Krita again, and the UI seems more approachable than I remember. Don't know what's up with that, maybe there was some freak misconfiguration last time I tried. Ah well.
The max RAM for a TR is 256GB right now, with 8x32GB - only unbuffered DIMMs are supported.
TR Pro can take up to 2TB RAM.
It also seems likely you will be able to drop a Zen3 Threadripper into the P620's sTRX4 socket.
The visual purity of seeing all those plugins loading, and having it take a long time seems to be part of the mystique they are looking for. If it just started right up, it wouldn't send the same message.
But the test is with a very big XCF file, loading lots of high quality layers. I would assume that takes all the time. Unless it is the MS Windows filesystem that is terribly slow, that might also be an option :)
warm-start (that is, launching after having previously killed the process but stuff's probably in a cache somewhere) Photoshop: 4s GIMP: 3s
I don't know if that's still the case.
People use their workstations for office/web browsing/etc tasks, so a large part of the workload continues to be single threaded. Which is why the Threadripper exists at all (its a epyc with faster single threaded perf).
So a good workstation is one that does well with both single threaded as well as multhreaded workloads. So to ignore the other half of the desktop workload performance doesn't give a complete picture.
There's no indication or suggestion that this is a good choice for a generic desktop system (note that desktop and workstation are different segments). The type of workload that you'd be buying a $1K+ CPU (this includes Threadripper and excludes desktop Ryzen) is the sort of thing that needs some form of specialized hardware. Maybe that's lots of IO (via 64+ PCIe gen4 lanes), or maybe lots of cores, or maybe lots of RAM (via 4-8 channels and support for registered DIMMs).
If you have one of those needs, then it's probably worth a tradeoff of a few hundred MHz of low-utilization peak clock frequency. That said, with low core counts, the highlighted CPU clocks >4GHz. That's on Zen2 cores, so it is largely similar to what you'd see at similar clocks in Ryzen desktop CPUs. If you want lots of general purpose workloads, go check out benchmarks for Ryzen 3000-series. There is ample coverage of these processors.
Many of the vfx, engineering, etc workloads can (and are) also run as batch jobs on server farms. So, frequently what your looking for is a desktop that can do quick single user trial runs, and then the whole thing is farmed out to the cluster for the final product.
But there are also segments (like building code) where the workload actually may have long sections of single threaded work. AKA, the machine burns through a bunch of cores in parallel, then it sits around with a single core linking/packaging/etc (yes I know some build systems can parallelize that too).
So again my point is that if all you need are a pile of cores your not looking for a workstation, your looking for a server which is tuned for max throughput across a lot of cores. Frequently this means running more cores in a more energy efficient configuration (lower clocks/etc).
Thanks mods, as always.
Sounds strange.
(can't find it for sale yet; hopefully the price is as low as the performance)
https://www.phoronix.com/scan.php?page=article&item=asus-50-...