AMD EPYC 7713 'Milan' Zen 3 CPU Benchmarked
wccftech.com
wccftech.com
I suppose it's more surreal than comparisons I made back then because there's very little difference, really, in the operating system and popular languages.
I'm pretty sure double the core count at 90% of the speed will get you better performance in _a lot_ of scenarios.
And we'd still have be able to make processors with lots of slow threads of course for applications where that's cost or power effective, that's comparatively very easy.
These days however, the difference between fastest core performance on low-core processors and fastest core performance on high-core processors is rather a lot less than 50% (Comparing like for like, eg AMD to AMD and Intel to Intel).
This kind of machine is nice to look at from afar but for most apps trying to get anywhere near a 64 fold speedup on your application on this would mean scrapping the code base and doing 1 failed rewrite, followed by 1 marginally successful one, taking up ruinous amounts of calendar time and engineering resources.
But of course it's diferent now than in 2005, because today we can't do any better.
I was skeptical so I downloaded some SPECint benchmark results ( https://www.spec.org/benchmarks.html ) of T1s, Power and Xeon, compared them and thought "mh, probably a bad idea" and then I ended up having to invest quite some time to convince my management to keep using "normal" servers.
On the other hand, later, I had a bit of fun hearing stories from other colleagues telling me how slow those T1-machines were once they started running on them their normal DB-workloads => after months of everybody complaining about bad performance, everybody went back to normal servers => a lot of time & money (& nerves) spent for nothing.
Normal IT politics, I guess.
Lunch driven procurement.
In all seriousness, someone is trying to make an app to benchmark software[0]. You might be interested in that thread. Thanks for the link.
(famously used in the Cray 2)
Similarly expect similar prices since nothing there really changed either. Yields are a very known quantity at this point.
But the problem isn't so much with the Core Count but the TDP per Core. Imagine it is 3W per Core, 128 Core gets to 384W. And that is excluding the IOD.
I did some back of the napkin math and realized that was roughly equivalent to my Zen 2 box.
So of course my failure to replicate the author's success was not the fault of my hardware.
In the end it seems GA/GP was an "overoptimizing" for very non-convex/non-linear search spaces (much more than real problems) and that we're better off with NN/SGD representations of problems
I have had a lot of fun this year building a neural network from scratch as well as genetic programming. For real tasks, it does seem that the former has won out.
I have a layman's suspicion that GP will find itself useful in some way and have a resurgence in the future, but it might just be naïve wishful thinking.
I once worked for a server manufacturer that did a setup with hundreds of 1U size, dual socket Pentium 3, 1.0 GHz CPUs. In the era when the only way to have more than one core was a dual socket motherboard.
[1] https://www.gene-expression-programming.com [2] https://en.wikipedia.org/wiki/Gene_expression_programming [3] https://ccc.inaoep.mx/archivos/CCC-17-009.pdf
Hopefully this changes with Zen 3 looking like an awesome consumer chip.
EPYC = Brand
7 = 7000 Series
25-74 = Dual Digit Number indicative of stack positioning / performance (non-linear)
1/2 = Generation
P = Single Socket, not present in Dual Socket
When scanning the product stack or test results it's really annoying to look for the last digit though. But I guess Intel does that too with the whole vX thing at the endIn their other markets they have very clear generation naming, even going out of their way to give naming-room to their laptop chips.
AMD really want to be in server space because that's the safe spot (in revenue) for the next time Intel comes backs, and that's the one they missed last time they were ahead.
It is too hard to buy these awesome chips and have them NOT run the serious stuff that our group wants to. I wish AMD would come up with a solution that makes it magically work with code compiled for AVX-512 :-(
I remember there being also an issue of linked MKL libraries shipped with various software (e.g. Matlab) not using the appropriate AVX2 codepaths on 'unsupported' AMD processors and isn't was falling back on slow non vectorized code.
Do we need all those cores in a single machine? Do we know how to use them? Beside some scientific workload, how are they used?
I guess the standard answer is virtualized them and sell them in chunk as "cloud" but then why prices are still so high?
I am not complaining, I just feel that something is off, and I don't understand what it is.
AMD knows they have some serious advantages there, especially on the efficiency front, with Zen3, and they're going pedal to the metal.
The prices felt off mostly because CPU are not the only cost, and in many cases it isn't even the most expensive cost on Server. Memory is, especially ECC Memory. Memory price per GB hasn't fall at all. And in many cases they are even more expensive. And it is only in recent few years you start seeing "High CPU Instances" because it is cheaper to include a Higher Core Count in the plan than to increase the Memory.
And because of the increase in VM / Core Count and Tenant, the server are now equipped with multiple High Speed Network connection. And if you look at it from a Cost Per Network Card / Port / Bandwidth / Server / Tenant. Price hasn't fallen much either. Adding in your cloud / VPS instance are paying Electricity, Co-Location and Networking charges, none of these have an Moore's Law and are part of the TCO. And if you include the continuous raise of wages, things doesn't look all so good.
But Zen 2 is barely deployed so its effect on the industry is not here yet. Linode already use Zen 2 and they are already offering better price plan than DO or other Medium Cloud vendors. I expect Zen 3 to be big with many of the Cloud Vendors. And you should see those reflected on pricing.
If this wasn't needed or wasn't useful then we also wouldn't ever see companies buying more than a single rack (still want redundancy and hot spares ofc), but they do.
So yes we know how to use all these cores. You take what's currently 4 machines and consolidate it to a single machine. Or similarly 40 machines to 10 machines, etc...
This processor has 64 PCIe lanes, and a drive uses 4. If you use half the lanes for disks, you’d need to multiply my numbers by 8. So, in addition to running application logic and talking to the network, each core needs to parse 312MB or dispatch 62K IOs per second.
Those numbers are certainly achievable, but hitting them requires some care. So, this processor is fairly well-balanced for moderately compute intensive workloads, such as query processing, or maybe frontend logic (which would probably want a machine with more network and less disk).
Note that the 128 cores are just to keep up with storage that easily fits in 1U or a fairly compact desktop tower case.
If you measure speed as the IO to compute density, these CPUs are actually slow and bulky by historical standards.
Also, there are sites like Source Hut or Travis that run thousands upon thousands of CI jobs every minute. I worked at a place where we were resizing thousands of photos and hundreds of incoming videos ever minute. One of these could easily replace 10~20 nodes in a transcoding farm for companies that do these types of workloads on physical servers.
There are tons of commercial applications that would keep these things busy during business hours, and some that would also keep them busy 24/7.
https://youtu.be/72AHENDeTEI?t=892
I guess it's actually pretty good for any CPU class.