7nm AMD EPYC “Rome” CPU with 64C/128T to Cost $8K (56 Core Intel Xeon: $25K-50K)
techquila.co.in
techquila.co.in
It's not like I cheaped out either, I have the same cheese-grater like anodised dark grey aluminium case with impeccable internal structure for easy upgrades. (Link to case: https://www.youtube.com/watch?v=gcNsHS2U8RM)
The technical background for this is that Metal and OpenGL acceleration requires the GPU driver bundle to be loaded into a process's address space. (effectively a dynamic library) With 10.14, Apple has enabled the "library-validation" codesigning flag just about everywhere, which means that only libraries signed by the same developer as the process's main executable or by Apple can be loaded into a process. Hence, 3rd party GPU drivers won't load in WindowServer for example, which makes them dead in the water.
Apple could of course implement an exception or sign Nvidia's driver, but so far it looks like they're not interested. Perhaps there will be a solution to coincide with release of the new Mac Pro which once again contains PCIe slots you conceivably might want to fill with Nvidia cards, but I wouldn't count on it.
(FWIW, my company maintains the macOS driver for one of 2(?) manufacturers of USB graphics adapter chips, so I have a pretty good idea of how the macOS graphics stack works.)
https://www.macworld.com/article/1154456/nvidiagpu-macbookpr...
This left the major computer vendors (Dell/HP/Apple etc) with very large bills due to the high failure rate of these Nvidia GPUs in laptops. It was right after this that Apple stopped using Nvidia, I've always assumed bad blood over this was a major factor.
The cost to Nvidia ended up in the hundreds of millions (based on charges one can see in quarterly filings from this time), the total cost of the failures industry wide is pretty hard to establish given laptop manufacturers weren't exactly in a rush to let customers know the extent of the issues...
One of their highest priorities at the moment appears to be to remove any 3rd party code running in the context of Apple processes. For examples, you don't need to look far: kext -> DriverKit/System Extension transition, hardened runtime, etc. The Nvidia driver breaking is just a consequence of that increasing lockdown. The simplest solution would be to allow nvidia's code to run in the context of core system and user processes again (kernel, WindowServer, just about any GUI app) which presumably is not palatable to them or they wouldn't have made the change in the first place. The alternative would be to farm this type of driver out into a sandboxed user process. (Nowadays this means DriverKit.) Aside from potential performance issues, this would be a lot of work, so it's probably never going to happen unless literally millions of end users start complaining.
> Threadripper with many NVidia GPUs as well?
Yes, it works with Threadripper, but not with NVidia GPUs, but that's because regular Macs don't work with NVidia GPUs as well. My understanding is that it's a standoff between NVidia and Apple largely due to NVidia's poor behaviour the last time NVidia was included in Macbook Pros around 2012.
Yes, I have one of those famed 15” rMBPs that many devs say “over my dead body” when you suggest upgrading to a newer model. Zero GPU complaints after the initial bugs were ironed out (and the bad batches replaced).
Note that AMD offered virtually no competition to nVidia in either the mobile or the workstation configurations (power or efficiency) up until at least the availability of the R9 Fury a couple of years ago.
I do realise that might be an untenable position for someone in Hollywood since FCPX is not a match for Premiere / After Effects at that level, however, there are plenty of other things that do the job of Photoshop and Illustrator, and they are not alternatives, Sketch is genuinely vastly better than Illustrator, and for Photoshop, there have been quite a few alternatives that the whole class of image manipulation apps got commoditised.
If I had an advice to give to my college self, I would advise not even starting with Adobe: if you're a designer and use their tools, the tools you need to make money with isn't even available for sale, you can only rent. So you might end up in a place where Adobe jacks up the price of Creative Suite by 10-100x per month after becoming dominant like Oracle, and you would have to pay your dues to their feudal fiefdom just to be able to continue your profession - else you'll starve. Not a great place to be for anyone.
Even if Microsoft is gauging companies with licenses.
Do you have power management (and sleep) working on amd?
> Do you have power management (and sleep) working on amd?
Yes (null-power-mgmt.kext) and yes. Though mind that this is a desktop, so power mgmt. / sleep needs are much less demanding than a laptop.
Specifically, the kext works by disabling power management from within OS X. However I have a kill-a-watt and the the power from the plug varies based on the CPU load, so it seems to be behaving correctly.
That said, I would not do this on a laptop. Just buy a MacBook Pro if you need that. My hackintosh needs are strictly because Apple (until a week or two ago) just didn’t make a computer as powerful as I needed. I’ll probably just go buy a Mac Pro in the first chance.
With me being a quiet computing / low power fanatic, null power management isn't good enough I'm afraid. Not to mention that amd-osx says "sleep will either work or not".
Using a hackintosh desktop for the same reason as you and yes, considering getting a Mac Pro. But it's good to be aware of your alternatives.
Making it with AMD is very tempting, but the fear is less compatible and will break even more than intel hackintosh?
Or perhaps they are trying to hurt Intel by selling it at a price where Intel can't make a profit if they lower their prices to be competitive?
Or, maybe it has to do with trying to quickly grab market share. Maybe AWS and other purchases of servers have a somewhat fixed budget, so cheaper chips translates directly to more chips sold?
Both those explanations seem unlikely to me.
AMD is working around this by using 8 separate CPU chips, wired together with one interconnect chip.
Intel is reeling from major design flaws and process stagnation right now. It makes sense for AMD to punch hard. Now is the time.
Also note that the eternally predicted ARM64 wave into servers, workstations, and cloud is not materializing. So far nobody has been willing to build such high performance chips and price them aggressively enough. All things considered it's an amazing window for AMD to take the market lead.
Personally I think Epyc this cheap really harms ARM's chances in the data center.
If Apple goes ARM64 for Mac it could indirectly help ARM get into the DC by showing that ARM is not just for small stuff.
Longer term you also have the RISC-V wildcard. A mostly free core and ISA could in theory move the whole game to the foundry and lead to a race to the bottom on price/performance.
It really depends. As long as performance is there and the apps work as expected, I'll deploy my workloads on whatever runs them for the best cost per transaction. Additionally, we don't have a say on what architecture our hosted services are on - if AWS decides my RDS databases are to move to ARM64, or Google decides my CloudSQL will run on POWER9 - I probably won't notice as long as the performance is right.
In the past I have deployed production Python-based workloads to amd64 and SPARC (I did POWER too, but for fun) without change. I'm most certain I can do the same with ARM64 or RISCV.
It’s going to be a huge factor in the DC. Never before have people been able to customize silicon. And now (ish) they can for figures that make sense to more than just the usual suspects.
Edit-I’m seeing some articles that make it look like arm might already be usable on Windows. There’s still legacy software issues though.
A serverless environment cares even less.
There is no oxygen down at 8k for Intel, they will have to move to chiplets to get the yields up, to drop the prices that far. AMDs costs for these parts is probably double what the consumer versions are, so we are probably looking at 7k+ of profit per chip.
They are forcing Intel to move to chiplets or be really really wounded. Clearly AMD is not participating in a duopoly game, which would be the expected anticompetative behavior.
If I were AMD I would give away the mobos. This is going cause a lot of Intel parts to be EOLd early.
If they priced at intel-$1k, for instance, many organizations would rule it out because the operational and additional hardware costs to switch would easily dwarf the $1k discount.
Sort of like people still buy gas cars even though the long-term economics of electric make so much more sense. The math has to be obviously beneficial (e.g. same sticker price, same range) for the switch to happen en masse.
AMD is not trying to get as much money out of it, they're trying to get market share, in a segment of the market where "single product line all identical all across" is a big thing.
Last time they were on top they didn't, and Opteron failed to take enough market share (a story that is not as well known because the bigger one was the customer chips and Intel abuses with oem). They don't want a repeat.
Its obvious that moore's law is stagnant, and clock speed is dead, etc, etc.
But are we at least getting the same core size at a lower cost ?
It wouldn't matter. You can't look at cost / transistor without factoring in Die Size and yield. Not to mention the cost of Higher Performance and Low Power Transistor are different. And the wafer price ( Cost / Transistor ) also exclude all the design cost and tooling around the node.
And Note: That 5K CPU has 256MB of L3 Cache. May be I could Run the OS not in RAM, but in Cache.
So there's not really any way to run system out of only cache, outside the tiny examples of early boot where all CPUs do this.
I wonder if one day we put even more cache in the I/O Die. Or even Stacked DRAM directly on top of it.
[0] https://en.wikichip.org/wiki/intel/microarchitectures/broadw...
[1] https://en.wikipedia.org/wiki/Skylake_(microarchitecture)#Mo...
[2] https://en.wikipedia.org/wiki/Kaby_Lake#Mobile_processors
Could the 8 core Mac Pro have been $3650 with a $650 Rome CPU, vs. the reality of $6k with a $3k Xeon?
So until USB4 is out, it is not Open "yet". And even if USB4 is out, there is no guarantee USB won't mess up with only USB4 4x4 that support Thunderbolt 3 40Gbps.
Here's the Verge article[1] that also speaks to this:
> Although USB 4 will integrate Thunderbolt 3’s features, Intel says that the two standards will coexist. While USB 4 is open, Thunderbolt 3 is not, and Intel requires manufacturers to be certified to use it. It also offers these manufacturers more support with reference designs and technical support. USB 4 might have the same specs, but Intel provides other Thunderbolt 3 services that go beyond the hardware itself.
[1]: https://www.theverge.com/2017/5/24/15685096/intel-thunderbol...
Most programmers don't care whether a CPU has AVX512 support. Most programmers will use a library or an OS service that'll abstract that away and pick a certain code path appropriate to the running CPU.
I don't think a human being can get a full understanding of a modern x86 without major brain surgery that is not yet available.
Huh? That Xeon in the base-spec MP is a $750 part: https://ark.intel.com/content/www/us/en/ark/products/193739/...
What value?
Is it worth spending $1-3k on a Mac if you just want to browse the web? Probably not unless you're really into their branding. Is it worth it if you have an iPhone, iPad, and Apple Watch and use many of Apple's shared services? Probably, and you get even more value out if it if you use popular software that targets macOS like Adobe products. If you're buying the ecosystem, the price is a lot more justifiable.
Personally, I don't own any Apple devices and I primarily use Linux, so my only one case for macOS is the iOS simulator and XCode so I can build and do basic tests for apps on that platform. And honestly, since that's the only extra value I'd get from them, their hardware seems a bit overpriced (I'm already paying for the developer license and sharing profits, hardware costs and lack of selection are salt in the wound compared to Android).
The mac pro uses workstation parts (Xeon W) not workstation parts. It doesn't use a $3k xeon for the base configuration (the 8c16t is a $750 chip) and probably wouldn't use an EPYC either (it'd use Threadripper, which hasn't been announced yet for Zen 2 but the 12c24t Zen+ is $650).
I think you mean *server parts
While not addressing your points specifically, which I agree with, and more for anyone else reading as I assume you already are aware...
It is important to remember that the (relatively) high-core-count Zen2 Ryzen chips do differ from Threadripper by more than just core count.
TR has a lot more PCIe lanes and it's quad channel memory (as opposed to Ryzen's dual channel).
Looks like clickbait to me.
Question: for single threaded performance (*n where threads are independent) do I want higher watts per core or higher number of cores? By "performance" I mean throughput.
Typically, games won't take advantage of new instruction sets until it's ready to be a minimum requirement, as otherwise you need to maintain two code paths, the benefit of which is only to be speeding up execution on what are already the more powerful CPUs.
What games can benefit from AVX512?
If the consoles go ARM before AVX-512 has been standard on new PCs for years, it may just not be worth it.
Gotta love the price segmentation.
This requires good planning of the game, i.e. this reduces inlining so one needs to be careful with putting the important bits in the JITted code
It might work in theory, but it increases the cost of testing too much for it to be worth it.
I saw it the other day and was amazed I hadn't seen it before.
FreeBSD kernel uses ifuncs for SMAP, XSAVE, ERMS (Enhanced REP MOVSB/STOSB), RDSEED.
Edit: I checked and they're not ifuncs
AVX512 though, even if you don't use the full width registers (eg to avoid throttling), the new instructions are very useful and not at all present on Zen 2. I wouldn't be surprised if Zen 3 kept 256 bit vector units, but supported the AVX512 instruction set.
Intel CPUs do downclock for both avx and avx-512. Many motherboards let you configure that in the bios.
I bought both a 9940X and a (delidded) 7980XE for avx-512 intensive workloads. I use a water cooler with a 360mm radiator. The bios for these (overclockable) chips contains "AVX Offset" and "AVX512 Offset" parameters. I haven't really tested avx loads, but the avx512 downclock is necesary. They run at 70-80C when running all cores at 3.5GHz in avx512 heavy loads. I don't want to push the temperatures further than that. I don't think there's any practical way to avoid having to downclock. It does gives a sizeable speed boost overall (for those workloads).
Re: this thread I'd bet those avx512 workloads will be faster as avx2 workloads running on two 64 core CPUs with avx2 than one 56 core CPU with avx512, all else equal. But it sounds like, instead of all else being equal, things like IPC favor Zen2.
EDIT: If folks happen to be interested, I compared a bunch of different "^" (aka, "pow") functions in Julia, running on my 9940X here: https://discourse.julialang.org/t/slow-arbitrary-base-expone...
Someone else shared results with their Ryzen 2950X here: https://discourse.julialang.org/t/workstation-advice-for-mos...
The vectorized versions were those with "sleef" or "xsimd" in their name. They tended to be 1.5 to 3 times faster, while the nonvectorized versions were 1.25 to 1.35 times faster on the 9940X.
Some of my other code is likely to show a much bigger difference. For example, many small matrix multiplication operations get to take advantage of avx512's masks to vectorize efficiently, as well as the fact avx512 has 32 instead of 16 floating point registers to hold larger matrix blocks in registers, increasing the vfma to vmov ratio.
I suspect the 3.2x difference in speed in the "jsleefpowcob!" benchmark is because of the register counts. I suspect with avx512 the compiler was able to avoid register spills, while with avx2 it had to reload a lot of data on each loop iteration.
The biggest problem with avx512 IMO is that compilers seem bad at taking advantage of it (eg, they never use masks) unless you babysit them / write code with vectorization constantly in mind.
gcc for example will not use 512 bit vectors by default. You must explicitly specify "-mprefer-vector-width=512". My tests (mostly just the Polyhedron Fortran benchmarks, as a set of numerical code) seemed to confirm that gcc (gfortran) was doing the right thing.
Meaning unless you intend to go low level and use it yourself (which can be a rewarding hobby!), or have workloads where optimized libraries exist, you won't see any benefit from avx512.
My use case is solving large LP/MIP problems for power markets, do those algorithms benefit?
Try to use perf-tools to determine how you're currently using these execution units, and then you can look how much it might help. Rule of thumb: double the bitwidth of vector instructions gets you 80% more speed, instead of the theoretical 100%.
At least this is a move that doesn't require a full recompile.
Intel Xeon 8280 QS is $1,800 USD each, a pair of those (56 cores) on a supermicro mb with 12 x 16G RAM is about $5k USD. It has an impressive Cinebench R15 score of 7,800+
https://ark.intel.com/content/www/us/en/ark/products/192478/...
That's a lot of caveats to hide behind two letters.
You voiding your contract with Intel does not make me a criminal. At best this is a civil tort.
Trafficking in stolen property is illegal in all 50 states, although the severity and specifics vary. If you buy them across state lines, 18 U.S. Code § 2315 also applies, making it a federal crime.
And taking and selling the QS/ES chips is not "voiding your contract". It is theft, and given the value of the chips, would qualify as a felony.
It's exactly the same as if I were to lend you my bike and you were to sell it to someone else. That someone else is buying stolen merchandise.
> [...] are the sole property of Intel [...] Are not for sale or resale.
So, whoever sells ES/QS chips is selling the property of someone else without authorization to do so.
Incidentally, the cost of the Xeon Platinum 9200 are also given in the article.