AMD Threadripper 3990X 64-Core beats dual Xeon Platinum 8280 in benchmark leak
hothardware.com
hothardware.com
A few key wins with the three cloud giants, and Intel will feel it. And the cloud giants can start with things like managed messaging, databases, etc, where it's not directly exposed to end users. Similar for the various high end AARCH64 CPUs. Intel is in for a ride.
And Amazon is full steam ahead with their own ARM CPU. I wonder if Google and Microsoft paused and think they will need to act now if they haven't already.
[1] https://www.anandtech.com/show/15433/intel-q4-fy-2019-result...
The difference is that AMD is starting to do the same, and it seems the process and architecture lets them make more chips than Intel.
And this forces Intel to lower prices on the more powerful chips, something that had not happened in about a decade.
Right now the winners are the customers, having such great chips at good prices.
I think Graviton has really big potential. Maybe the biggest challenge is for AWS to market it well.
Just get the biggest honking Noctua fan+heatsink you can fit in your case and call it good. It's silent, simple, reliable, and if the fan does fail the heatsink will save your CPU.
Furthermore, you are supposed to use distilled water or non-water based coolants with high electric resistance.
https://www.leonroy.com/oldsite/img/IMG_1802%20(Large).JPG
Regarding distilled water, it is indeed not very conductive at all but it is also quite reactive and if it touches electrical traces and PCBs starts to leach ions from the surrounding until it starts to conduct at which point the machine will short circuit (speaking from first hand experience alas).
Seems like it would be easier to make a "never disassemble" system not leak.
Also not sure if you're counting AIO coolers in with water cooling but they are very reliable.
On the other hand, a heatsink provides a decent amount of passive cooling all by itself, so when the fan fails there's a buffer that slows the rate of heating to a safe margin.
If one of the two fans fail i’d guess temps may raise a few degrees and that’s it. I could even manually test it with one fan and light loads to check one fan performance.
In my experience air cooling with a sufficiently large heatsink is simple and very reliable. Water cooling is more expensive, trickier and can fail catastrophically.
Source: Been rocking Corsair AIO coolers for around a decade now. I did have have the pump fail in one of them after around seven years of use.
Perhaps if you don't forget to activate the BIOS CPU warning at high temp, you shouldn't worry about this .
I recently fixed a PC where the pump died without anyone noticed, and the only evidence, was that the BIOS was powering off the computer when the CPU gets too hot.
Bought the thiccest Noctua I could find on Amazon and haven't looked back.
I've got a custom loop with my Ryzen system (420 + 280 rads, mostly EK kit including a D5 Glass XRes) and it's been enjoyable putting it together, and if you're not totally incompetent easy enough to make leak-free with the latest generation of fittings.
As for a pump failing, as others have said the CPU will thermally throttle before damage is done, but for added peace of mind I'm running my setup with an Aquaero which will shut the computer down if the coolant temp breaches a set threshold, or it senses the pump has failed.
The biggest thing for me is that my computer is EERILY quiet unless I'm pegging all 8 cores, and the GPU, at 100% ... which makes for a nice peaceful environment.
Also modern CPUs have internal temperature sensors of course and don't 'burn up'.
This sounds like someone who is afraid of water cooling, not someone familiar with it.
The first rule of plumbing is that it's guaranteed to eventually leak.
(I don't know if this is actually a rule of plumbing, let alone the first one, but it sounds like it should be)
Also, FWIW: integrated liquid cooling solutions are if anything simpler and more reliable devices than fans: smaller motors running at lower RPM last longer.
Do I need to panic?
Although you can buy the CPUs in TRAY format without the cooler, the prices are almost identical to the BOXED version and your warranty is lower since they're designed to be sold to businesses not consumers.
I think this is exaggeration. I have 3600 with stock cooler and at full load clocks doesn't seem to fall below 3.8GHz, 86°C (well, I tested on linux make -j12, maybe on some AVX loads it is worse).
Funnily enough, siege has recently released a vulkan based build that runs at a cool 74 degrees.
The wraith stealth is just an incredibly anemic and disappointing heat sink. It's "functional"but only at average loads. This is in comparison to the previous generation's cooler for the midrange CPUs that was about twice as large. It's not a deal breaker, my 3600 is still miles above anything I've ever run previously, and it isn't dying under the heat, but it might just run one year less than it was initially designed for at those temperatures
My own Wraith Prism is just sitting in the basement, and I don't see myself ever using it, unless something happens to my high-end cooling system and I need to install something while getting it fixed, at which point it'll go back to the basement again.
It would be interesting to see statistics on how many Wraith Prism coolers sold together with 3900X CPU:s actually end up being used. I expect the number to be quite low.
It can, just fine. A better cooler keeps it cooler, and you might get an extra 0.1 mhz from it. With the Prism, if you have had/having trouble with it keep in mind two things:
1. there is a switch on it for a higher or lower fan setting 2. you have to crank that little lever HARD all the way over to properly clamp it down. Lot of people miss that.
It can't handle it in the sense that it's very loud in doing so, and the CPU will boost more with a better cooler, even with no other changes.
I bought a 3900x and planned to use the stock cooler initially, but that lasted about 3 weeks before I replaced it due to noise primarily (I got a be quiet dark rock pro 4).
At idle with default fan curves, it was on the knife-edge of max-speed not just loud but continuously revving up and down so incredibly distracting. It would continually grab my attention even with headphones on.
I barely lasted a day before ordering a Dark Pro 4 and now it's silent.
A friend of mine recently asked if my PC is okay in a very worried tone, I said that’s how I build a PC and he now makes remarks of his annoyance every time the afterburner kicks in, all while it’s just a cruise flight to me.
Maybe you just got lucky... or perhaps double-check that you aren't accidentally under-volting :)
As the other comment says tho it is subjective.
I have been feeling like there's an issue with the board/BIOS/AGESA's PID control algorithm.
But AMD value is tremendeous, especially comparing with Intel CPU supporting ECC. I'm thinking about buying Xeon W-2265 workstation and it's just 12 cores for $1200, AMD is much cheaper and for that price I can get Threadripper monster.
AMD CPUs are used on servers so there is no reason they won't make all it's possible to have a stable system, are you referring to AMD GPUs? I read it takes up to 1 year after release until the graphics driver gets stable on Linux.
It’s true that AMD products has yield(as anyone else) as well as parameter tuning/compatibility issues especially early on. I don’t think their issues are of reliability kinds(like early failure in industrial Intel Atom) though.
I never had, or even heard of stability problem with AMD's CPU Core, but during the early days AMD surely had worse chipset especially in I/Os. Where it had compatibility or performance issues. But these days most of them are third party IPs, from USB Controller to PCI-E Express, and they are the same IPs used by millions if not billions on Smartphone, Tablet and many other use cases. So you can be sure they are pretty damn well tested.
Never had a single issue.
FWIW I used a 1st Gen Ryzen at home and it would constantly BSOD under Windows but ran completely stable with linux (this was when Ryzen was brand new)
I doubt ECC would help prevent it(other than maybe by being slower?)
Come on, AMD announced it would be $3990 at CES!
Is this now a factor? I know memory has gotten faster, and multi-channel of course helps, but is that all really enough for typical workloads?
In 2005, Opterons used a 940-pin socket and had dual-channel DDR, for 6.4GB/s total bandwidth to feed two cores. In 2009, Xeons the LGA1366 socket to provide triple-channel DDR3, 32GB/s for four (later six) cores.
Now, we're up to 4094 pin sockets providing 8 channels of DDR4 for a total of 204.8GB/s of memory bandwidth for 64 cores, which is exactly the same per-core bandwidth we had in 2005.
(And for comparison, a $200 graphics card will also have around 200GB/s of memory bandwidth.)
Modern CPU cores are way faster than ones in 15 years old CPUs.
That 2015 processor has AMD K8 cores. I wasn’t able to find exact info, but slightly newer AMD K10 can do 8 single-precision FLOPs/cycle, at 2.8 GHz.
Modern Zen2 cores can do 32 FLOPs/cycle at approximately same frequency, i.e. per-core performance improved by a factor of 4.
Both numbers are theoretical limits only achieved with very specific workloads and heavy use of manual vectorization. Still, real-life general-purpose code scales quite similarly between them. CPU benchmark says their single-threaded performance improved by a factor of 3.5.
And God said, "Let there be Cache." And there was Cache.
https://www.nextplatform.com/2017/02/23/promises-challenges-...
"AMD mentioned a latency increase of 7-8 ns when memory encryption is enabled, which results in a 1.5% performance hit in SPECInt"
Here we would be compensating with smaller data transfers. Wouldnt have to be an amazing algorithm, even RLE would probably give latency improvements.
As you can imagine, processor performance has grown a lot more than 6x over the years. The amount of cores in a package alone has increased more than 6x! So how do we get around this?
1) Larger caches: the most efficient optimization to any slow task is to simply not perform said task. Having a larger L3 cache means sending requesting fewer pages from DRAM.
2) “Hiding” latency with bandwidth: if I have to incur a (relatively) large time penalty tp go to DRAM, I better make sure to make it worth my time. Instead of always getting data in 4KiB pages, often times modern processors support “superpages” all the way up to the gigabytes. If you’re going to go drive out to the store, it makes sense to buy groceries for the week instead of just the candy bar you’re craving right now, right? That way when you actually have to cook the work has already been done.
I am probably missing some nuance here, especially when it comes to memory controllers and putting caches on DRAM, but this is the gist.
The specific implication details make a difference, but nothing is going to provide a 4x improvement in latency so other considerations take precedence. The real change is ever growing cache sizes allowing CPU’s to minimize random RAM reads.
“For a completely unknown memory access (AKA Random access), the relevant latency is the time to close any open row, plus the time to open the desired row, followed by the CAS latency to read data from it. Due to spatial locality, however, it is common to access several words in the same row. In this case, the CAS latency alone determines the elapsed time.”
btw 10cm of track on pcb is about 2 clock cycles at 3GHz, ~0.6ns. Now you will love this one - in the last 20 years we went from ~15ns CAS down to ... 7ns, twice as fast! Even better - we got down to 7ns in 2006 https://www.newegg.com/g-skill-2gb-240-pin-ddr2-sdram/p/N82E... and stayed there for the last 14 years.
What we got instead is higher density and susceptibility to rowhammer.
Further, MB traces are just part of the story. It’s the physical distance traveled from a CPU’s memory controller to the furthest physical cell location on a chip which is significantly more than 10cm and it’s a round trip timing. Even just the physical DIMM’s are 13 CM wide.
A more accurate ~25 cm each way = 50 cm total ~= 3ns out of 6.56ns timing for a DDR3-2133 & CAS 7 ~47%. CAS 18 DDR4-4600 is 7.82 ns or ~40%.
PS: As to progress, DDR2-1066 at CAS 4 was 7.5n where CAS 18 DDR4-4600 is 7.82 ns that’s some progress.
Further, every layer of cache adds latency and cost as you need to check it before going to main memory. In the 486 and early pentium days External Cache was common, but it’s no longer useful.
Page size is irrelevant for purposes of getting data from RAM to CPU / cache. This is done by cacheline size, which is usually 64 B. Higher page size helps with better TLB efficiency and also for disk to RAM copy (if that is strictly page-based).
The saving grace is not that there is a lot of memory bandwidth to go around; there isn't; the saving grace is that a lot of processing either doesn't use too much data (making caches effective) or is complex enough to not be limited by memory bandwidth. Rarely using all cores and threads at the same time helps a lot as well.
Wow wow wow.
Your references are TIGHT.
----------
The main RAM / CPU issue is latency: it takes over 100-clock cycles to communicate with DDR4, which a CPU could process ~400 instructions if they were lined up just right (modern CPUs are super-scalar, executing multiple instructions per clock tick).
Only once you have sizable caches + data locality will you solve the latency problem.
Without locality, I'm not sure if the latency problem can be solved at all. Fortunately, many problems have an element of locality that can be taken advantage of.
(Anyone have pointers to current EPYC vs Xeon STREAM results?)
My experience is that it takes well structured memory access to actually max out memory bandwidth even if the numbers seem like the CPU would be starved.
Things like adding two huge images with SIMD and writing out a new image could be like this. It would be much better of course if the image is tiled and more operations are done per tile, which would use the cache.
Good use of the prefetcher but poor use of the cache could still run through memory bandwdith.
They still have some performance advantages over all competitors.
But if you are not enterprise, I see no point in paying their prices.
As you say, free products have caught up and are just good enough.
The VMware thing was that they figured out how to make virtualization on i386 compatible hardware hobble along despite the hardware architecture being technically unvirtualizable, using clever binary translation style tricks. Virtualization support came to x86 hardware many years later.
If Xen has a show stopping bug I get to patch it myself.
Oh you mean technical reasons? None.
Also, I'm pretty sure if there's a Xen bug you can throw Citrix under the bus (assuming you're paying them for enterprise support).
Alternative exist, but VMWare has a lot of mindshare and their products are pretty good, so they're hard to beat.
Didn't Citrix buy xen a while back (or some part of it)? They had something called xenapp, now apparently "Citrix Virtual Apps and Desktops (XenApp & XenDesktop)"?
Not really the same as "a supported xen distro" I guess.
I agree about linking. Also Fedora (and therefore RHEL 9) is about to move to LTO which makes everything much worse.
It takes years for the big companies to get interested in such a big of a platform switch. For now, they've been sticking with Intel and observing how the situation develops. If Intel keeps screwing up, which they do, eventually the big migration will happen.
With Intel, I get advanced warranties, presales support, next day parts, details marketing and product information alongside access to privileged information and roadmaps.
With AMD, I was waiting the best part of 3 months for a call back, and their "partner" site which is open to the world is severely out of date and doesn't even have all products listed.
Intel is frictionless when it comes to doing business, AMD put up roadblocks.
I'm still running a i7 2700k and haven't even felt the need to upgrade. I guess i didn't want to when Intel only released $600-$800 scam CPUs, now were seeing reasonable competition is it worth another look?
[0]: https://www.pc-kombo.com/us/benchmark
[1]: https://www.pc-kombo.com/us/benchmark/games/cpu/compare?ids%...
There are many gamers using 144Hz displays now, so 60 FPS is just one of multiple possible targets :) Witcher 3 is also really not the heaviest of games.
> Of course processors HAVE become massively faster but the question remains whether one has a workload that's sufficiently parallelizable to warrant the upgrade.
Well, sure. But all common workloads are parallelizable now. Whether it's games or running your browser, the time of strictly single threaded workloads is over.
> Individual core speed improvements are still lacking.
Interestingly, that really depends on your definition, of what one expects. The difference in single threaded workloads is smaller than in multithreaded workloads - obviously, there are more cores - but the cores are also ~30% faster in this case, and Ryzen 3000 got another round of singlethread performance improvements. For example https://www.computerbase.de/2019-07/amd-ryzen-3200g-3400g-te... shows that, a benchmark of single core application workloads.
Every day work just feels more snappy and compile times are down by a lot. Single thread is a lot faster and of course having 12 threads instead of 4 makes a huge difference (of course that effect might be less for you since you already have SMT).
Paired with a 1060 I even notice quite a difference in gaming. Even though I didn't necessarily have 100% usage on the CPU, I was 10-20FPS below benchmarks for my GPU. This didn't really make sense to me, but after I upgraded I indeed got those extra frames.
I haven't done any scientific on the performance improvement, but according to PassMark, I am getting roughly a 120% performance increase on parallelizable workloads, even when taking my OC into account.
That can happen when the game performance relies on the workload of one thread (or at least less then cores/thread count) that couldn't be properly offloaded to the idle cores. The 3600 has not only more but also stronger cores than the 2500K, and thus even without reaching 100% processor usage before the game can run faster now, because the critical thread runs faster.
Therefore I really thought it couldn't have been a CPU bottleneck (since one would assume it could have scaled at least another 10% on every core, but in the end it was.
Before that I had a Radeon 6970, that I bought with the 2500k, but I upgraded that in 2016.
In hindsight upgrading a week or so before an exam might not have been the smartest decision, however.
The AMD Ryzen 9 3900X 12-core was what I was looking at. This new 3990X looks amazing, but its probably overkill for what I need.
I would recommend getting a different cooler though. The stock cooler, while adequate, runs noisy and doesn't feel all that sturdy.
In my mind this makes it preferable to opt for a stronger GPU, if you have to decide between a stronger CPU or a stronger GPU (e.g. you have 300 $ left in your budget and want to decide where to put them).
Apple is already using AMD graphics, so they already have a business relationship.
They don't have to. Most of the Mac users don't care. They know that the latest Mac is the best Mac they can get.
The figure from 2017 /2018 were 80% Laptop and 20% Desktop. And since Apple is kind of stuck with AMD for Graphics, it sort of make sense if Apple unify their GPU stack on Laptop as well. Assuming they get the Thunderbolt issue sorted.
I feel like Apple could have easily prototyped the Mac Pro on older TR designs and evolved the design as AMD got deeper into sampling the current gen.
I think what eliminated the risk of switching to AMD for them was the ability to test their thermal design on already shipping parts from Intel. The chipsets haven't changed at all so they've been working on a stable platform since day 0. Or maybe they got an offer they couldn't refuse from Intel.
CPU compatibility with macOS is also a total non-issue, considering how well the AMDOSX community is doing.
And there is Thunderbolt, its certification is still not released from Intel. Before than Apple could not leave Intel just yet.
I have one computer with Gigabyte X399 Aorus Xtreme and it is an utter crap, especially the firmware. It randomly resets its settings, even after you just saved them. Sometimes it takes few reboots to get it to acknowledge, that yes, I did enable AMD SVM and want to use it.
Of course, it makes no sense if you only game. TRX40 platform can be better used as a workstation but is indeed expensive, I would say too expensive, for what you get. Only 4 DDR channels, only 256GB RAM at max, only expensive unbuffered DDR4. Most TRX40 boards have only one 10Gbit NIC. Only 4 PCIe slots... This is 2010 level tech.
The only real selling point of this is TR CPU performance and performance per cpu cost, but otherwise this is a mediocre platform. It does have PCIe 4, yes, but that has little benefits for now. "Pro, Extreme, Creator" tags do not make this on par with the established Xeon workstation platforms. This is how a high-end workstation board looks like:
https://www.servethehome.com/supermicro-x11spa-t-motherboard...
You can get it for... wait for it... $500. Clearly TRX40 mobo manufacturers are taking advantage of gullible consumers.
And that's Intel fault, by basically blocking consumers a few years ago from using and buying Xeon processors...
You meant 4x PCIe slots? That must be a joke. Even x299 with only 28 PCIe lanes has 7... TRX40 could happily support 8x PCIe 3 slots, but nobody makes such board.
But no, I meant basically everything else. The amount of ram, USB, SATA and M.2 slots as well as USB headers is nice, compared to regular consumer board at the very least.
Now imagine having an 8-slot motherboard for 3990X that could be configured to e.g. 4x PCIe 4 x16 or 8x PCIe 3 x16. Wouldn't you like such motherboard?
For Deep Learning I don't really care about PCIe 4.0, but I do care about being able to fit as many GPUs in as I can (they talk to each other outside PCIe anyway). With TRX40 I am limited to 4 GPUs. Even PCIe 3 x8 is good enough for DL.
Intel do servers at all levels including low core count/high core frequency, where as AMD have a very decent core count, but, performance between the models is largely just an increase in core count (there is very little difference in frequency - e.g. NO high frequency/low core)
So many jobs still require Windows licenses unfortunately and since the shift to core count, it's ludicrously expensive to license a lot of the new AMD machines on SPLA licensing.
Not that I'm not trying though...
Meanwhile if all you need is single thread performance for lower core counts, that's Ryzen rather than Epyc, but it's not like they don't make that.
For example, I7/I9 is much cheaper and more powerful than entry level Xeons, but, you VERY rarely see them in a rackmount chassis/server.
Same goes for Epyc and Ryzen.
Whether or not it's common, it's available if you want it.
So relevant? Probably just as important now as then, but a few P3 designs not melting two decades ago shouldn't give you confidence to do this with a 9900KS.
P3s melted, Throttling was a brand new P4 feature, probably a must considering how hot they ran
That's what I meant by "a few P3 designs"; some, not all of them.
This was at the height of Intel meddling. They directly bribed vendors, ugh Im sorry, I meant offered rebates under MCP (Meet Comp Program) in exchange for strict no AMD commitments. DELL, HP all took Intel $, $6 billion in kick-backs.
Another Intel tactic was manufacturing facts and positive press stories, something we now call fake news. Ubisoft loves money, so it will come as no surprise to learn they took a brib^^ promotional marketing funds to plaster huge "Designed for Intel MMX" on game BOX https://www.mobygames.com/images/covers/l/51358-pod-windows-... Hint: MMX is Integer math only, unsuitable for accelerating 3D games. Whats worse MMX has no dedicated registers, and instead reuses/shares FPU ones, this means you cant use MMX and FPU (all 3D code pre Direct3D 7 Hardware T&L) at the same time. In POD its used for one specific sound effect (audio filter) and has zero influence on game speed. BTW this wouldnt be the 1st time Ubisoft sold its clients to a highest bidder. Nvidia copied this technique with "the way its meant to be played" campaign https://news.ycombinator.com/item?id=22090413.
Intel SSE also got a big push with fake 3D acceleration claims https://www.vogons.org/viewtopic.php?f=46&t=65247&start=20#p...
Intel version: "At the time, I was working for Intel and was involved in the launch of the Pentium 3, aka Katmai.
We _engaged_ a number of games manufacturers to provide demos showcasing not only Screaming Sindy's Extensions, but the arcane and mysterious Katmai New Instructions.
One such outfit was Rage Software, now sadly deceased. Rage provided demos of Incoming and an early prototype of a game called Dispatched, which as far as I know never actually saw the light of day. Dispatched featured a strangely-arousing cat riding a jet powered motorcycle. The first version I saw was running on a 400MHz Katmai and was still in wireframe. It was bloody impressive."
Reality, according to hardware.fr: "Let's start with Dispatched first. This is actually a Rage Software game that should come out late 99, which Intel showed the demo at Comdex Fall to highlight the benefits of the SSE. Big interest, it is possible to enable or disable the use of SSE instructions at any time.
Nothing to say in terms of speed, it goes squarely faster once the SSE activated, + 50% to + 100% depending on the scenes! But looking closely at the demo, we notice - as you can see on the screenshots - that the _SSE version is less detailed_ than the non-SSE version (see the ground). Intel would you try to roll the journalists in the flour?"
SSE version is less detailed? How convenient! Rage Software Dispatched never came out. The only outfit, other than Intel, in possession of this software was Anandtech. They used this exclusive press access to pimp out Pentium 3 benchmarks manufacturing fiction like this https://images.anandtech.com/old/cpu/intel-pentium3/Image94....
Without the advent of Quantum computers, solving P=NP for our computational benefit, and/or some other unworldly mathematics, the PC industry is pretty fucked in ~10-15 years save those who transition into services.