AMD 2nd Gen Ryzen Deep Dive
anandtech.com
anandtech.com
https://www.anandtech.com/show/11859/the-anandtech-coffee-la...
https://www.anandtech.com/show/12625/amd-second-generation-r...
The fact that the i7-8700K shows up as slower than the i7-8700 now is a bit weird too.
It's impossible. The 8700K and the 8700 are the same chip, just that the 8700K has an unlocked multiplier and a higher default clock[0] - better binning. The 8700 will never be faster in any application.
[0]: http://www.cpu-world.com/Compare_CPUs/Intel_CM8068403358220,...
I think it is more likely to just be testing margin of error
When the demand for the i7 920 was way up, there were a bunch of better chips sold as such.
Savvy overclockers knew which batch numbers they were and sought the particular runs.
Finally, even with identical parts, the assembly can affect performance. The same heat-sink installed slightly more effectively in terms of the mechanical interface (no contamination, just the right amount of heat-sink compound, no air bubbles) will allow the chip to run at higher power without any thermal throttling.
With many systems, turbo is only a temporary burst benefit for a lightly loaded system. Many server or HPC systems would intentionally disable turbo to have consistent sustained throughput.
What did they do differently ? They didn't ask AMD for permission or sign an NDA for exclusive "early" access, and just sourced the parts on ebay from one of the numerous samples that were given to ODMs/OEMs.
They just had a big head start because they acted independently and didn't wait for vendor approval to publish their tests.
Edit: here's an english source with the magazine contents: http://www.guru3d.com/news-story/cpc-hardware-amd-ryzen-7-27...
Regarding performance impact, short answer is - yes. Basically, it removes most of the advantage of the cache for Intel.
So, even without Spectre & friends, these benchmarks are essentially meaningless?
Reviewers have generally settled on the former option; OS and driver updates tend to increase performance, so testing with contemporaneous software might artificially inflate the apparent performance deltas in favour of newer components. Nobody wants to buy a new processor thinking that it's 40% faster than their old one, only to discover that most of the difference was just OS optimisations. You'd rather have your readers be pleasantly surprised than bitterly disappointed, so it makes sense to choose a benchmarking methodology that errs on the side of favouring older parts.
Spectre is a weird exception to business as usual, because we saw a huge performance decrease in a single update that affects one particular optimisation in one particular processor generation. In this case, it probably makes sense for reviewers to re-run all of their old benchmarks and set a new baseline.
I thought the OS update was only part of the patch and that OEMs also needed to update their BIOS, which means users have to manually update their BIOS.
New software doesn't affect existing benchmarks.
I thought it was mainly workloads relating to virtualization etc... that were effected significantly
January 29th Microsoft had to revert the Intel patch discribed:
https://www.zdnet.com/article/windows-emergency-patch-micros...
There's also implementations in March for the Spectre variant 2:
https://www.digitaltrends.com/computing/microsoft-windows-pa...
There's much more in here.
The majority of DirectX stuff in the innermost-loops automatically DMAs to the video card without even needing a kernel call. Video games care a lot about performance and did their best to never touch kernel-space anyway.
Compilers, Web Servers, Databases, etc. etc. were affected very strongly. Compilers open files. Web Servers open Sockets. Databases communicate with mutexes / semaphores. All of these are OS-level features and require a kernel-mode switch. Meltdown means that every kernel-mode switch clears out the TLB-cache.
So each time a compiler opens a file (MMap or otherwise), it basically loses its entire TLB cache. Every. Single. System call. That's why its a big deal for servers, but not a big deal for video games.
The amount of speculative-execution that games rely upon for many of their components is insane. With the patch applied, I went from a solid 60FPS at 1440x900 in GTAV down to about 48 FPS (as measured by the Steam FPS counter) on a 7th-gen i3.
Simple files are more straightforward. You do one MMap and then there wouldn't be any other system calls made. I'd bet that compilers are "slow" because they have to make many, many, many mmaps to do anything. Each #include turns into another MMap, which Meltdown mitigations cause a full TLB flush each time.
I don't know if SQLite would be faster or slower in the context suggested here. But I'd like to suggest that perhaps it is worth running the experiment before reaching a conclusion.
I got word on reddit that pc-per said they did apply the spectre/meltdown patches for their benchmarks. https://www.pcper.com/reviews/Processors/Ryzen-7-2700X-and-R..., the results are as one would expect.
It's very suspicious when Ryzen 2 numbers are that far off and whilst they keep getting all this amazing access to AMD for trips and CEO interviews and new toys.
Edit: If the comments on the AT article can be believed, then Toms Hardware's Intel benchmarks didn't include the latest Meltdown/Spectre firmware patches because they didn't know that the X470 platform had them to begin with.
AMD's officially supported memory is 2933 MT/s. Intel's officially supported memory is 2600 MT/s.
Other review sites overclock memory and run both systems at 3200 MT/s RAM (closer to reality: you may have to spend $20 on RAM but the faster speeds seem to be worth it these days). It seems like Anandtech prefers to run the system on default settings however.
Its definitely a discrepancy. But I don't necessarily think Anandtech is in the wrong. It just goes to show how much these little decisions matter in benchmarking, and why its good to have multiple benchmarks from different sites.
So, do you go for the "obvious overclock" with 3200 MT/s RAM, or do you actually follow the official recommendations from the manufacturers? Regardless, its good to see that different websites have tested the two different cases and leaves it up to the reader.
See this line: https://www.anandtech.com/show/12625/amd-second-generation-r...
> As per our processor testing policy, we take a premium category motherboard suitable for the socket, and equip the system with a suitable amount of memory running at the manufacturer's maximum supported frequency
Personally speaking, I think running RAM at an overclock at 3200 MT/s is more fair to both systems. Both systems achieve 3200 MT/s easily, with Intel overclockers often hitting 4000MT/s or more (so really, Intel has the better memory controller). So its just... odd... that AMD would be equipped with better RAM.
But Anandtech's approach is reasonable (although they probably should list the timings of their RAM in their setup page). After all, most computer builders probably don't know the finer overclocking details between the systems, and may instead just go for the officially supported numbers.
at that point you're comparing a system that AMD doesn't sell to a system Intel doesn't sell
I build fast machines but I don't think I've ever overclocked one, I certainly can't remember.
I value stability and longevity, my 2500K lasted til end of last year (it still works fine, it's just my laptop is faster).
Its why Intel has XMP: so that your sticks of RAM will talk to the motherborad and automatically overclock themselves to factory-tested values these days.
These "easy overclocks" I'm talking about are very simple! The factory gives you an XMP profile of the RAM under specified conditions, and then the motherboard enables a checkbox for XMP-profiles in the BIOS.
No need to tweak timings or anything. These days, overclocking RAM is a question of "Do you want to run your RAM at the speed tested at the factory" ?? Or would you rather run RAM at speeds specified by JEDEC from 2014?
Every RAM manufacturer worth a darn has an XMP setting, where they've thoroughly tested the RAM at these higher speeds. You might as well spend a minute, click the button in the BIOS/UEFI prompt and accept the free performance benefit.
But DDR4 RAM has support for 16 banks per stick! A CPU with two memory controllers can have 64 outstanding requests to DDR4 RAM simultaneously (16 per stick, x2 from the chip select, x2 for the 2nd memory controller).
And since modern cores are highly multithreaded, and also have prefetchers and whatnot, I'm sure the modern DDR4 bus is quite active.
Perhaps it doesn't matter too much for single-threaded situations (single-threads only benefit from prefetch and out-of-order executions). But a Ryzen 2700x is a 8-core / 16-threaded beast, which will surely eat up all of those outstanding DDR4 requests. Intel's 8700k is 6-core/12-threads as well, so they'll rack up those memory requests and keep the bus busy.
And if each of your Assembly commands are AVX (256-bit requests per instruction)... well... that's going to be a busy bus.
This is kind of like the old days when CPUs used a Front-side Bus (FSB). If you increased the FSB frequency it would increase the CPU frequency, the RAM frequency, and even PCIe depending on how it was wired up. It made things simpler to have the FSB be the sole clock-generator.
So despite the increased bandwidth moving from 2600 to 3200mhz RAM not having much impact on performance (in most applications), the increased frequency of the bus between the cores allows for greater performance beyond that.
IIRC Intel uses a ring-style bus with it's own clock, hence why it this effect only applies to AMD.
> the old days when CPUs used a Front-side Bus (FSB)
They still kinda do. It's not the bus, but there is one clock called BCLK (base clock) now and almost everything is derived from it. It's not adjustable on cheaper boards, but with a good board you can absolutely do the same adjustment :)
Sandy Bridge and newer get extremely unstable when you toy with the base clock more than a few MHz though. You're basically stuck shelling out an extra few bucks for a "K" chip if you want to overclock on Intel now.
https://overclocking.guide/category/intel-oc-guides/skylake-...
Of course on Kaby and later they closed the loophole...
Edit: I missed that the benchmark methodology is on a different page from the results.
Memory latency / bandwidth is a piece of any benchmark. Some benchmarks are purely CPU (Linpack), but others are better overall tests of the system (ex: Cinebench or Blender)
But if you have a particular workload in mind some of these aspects might not important to you and you'd be best served by just benchmarking your particular application on a given chip.
https://www.pugetsystems.com/labs/articles/After-Effects-CC-...
I assume these tests are unpatched Intels though. Since my use case for CPU is video rendering and not gaming, I still don't have a reason to switch to Ryzen, as much as I want one. Threadrippers are still way out of my price range (~$1K) and are about the same performance as top tier i7 8th gen.
As to why After Effects CC seems to prefer Intel so much would indicate some specific optimizations Adobe performs, optimizations that seem to favor Intel for now, not general raw multi-core performance of the Intel CPU.
For single core performance the situation is reversed, the higher clocked (but lower core count) Intel CPUs will obviously have much better single core performance which is important for games as even today they tend to be coded relying on a few threads.
I bet its AVX-512. Which is "Intel Specific" but that's the entire point of AVX-512, a new Intel specific instruction set that has high performance benefits.
Overall, AMD is betting on general purpose cores with at most 128-bits per operation. AMD supports the AVX 256-bit instruction set, but its "emulated", as the ALU pipelines are 128-bit (granted: AMD can gang together two ALUs in a 2x128-bit configuration to execute the 256-bit instructions)
Intel supports 3x256-bits per operation on their CPUs, with the 512-bit AVX512 instructions using two execution ports at a time.
So for any code that's written to be SIMD-heavy, Intel machines will likely be faster. Not only because of AVX512, but also because Intel's AVX-256 bit instructions are better supported compared to AMD (3x256-bit for Intel, vs 2x128-bit for AMD)
------------
With that being said: the benefits of SIMD approach Ahmdal's law just like anything else. In particular, AVX512 uses so much power that Intel cores are forced to downclock.
So AMD can run more cores, at higher frequencies. But Intel's cores do more work per clock tick and per unit energy. It seems like AMD's 16-cores (each doing 2x128-bit) are only slightly slower than Intel's 10-cores (each doing 3x256-bits) in those SIMD-heavy tests.
It's not quite that harsh for workloads that aren't entirely FMA. Zen has 2x128-bit adders plus 2x128-bit multipliers.
IMO, the more important bottleneck is the load/store unit. Intel's system has 2x 256-bit loads + 1x 256-bit writes to the L1 cache.
While AMD's system has 2x128-bit load+store units for its L1 cache.
Thus my numbers above: 3x256-bit for Intel (2 reads + 1 write), while 2x128-bits for AMD (2x read-or-write).
The questions for me are: Is it worthwhile to upgrade? And will my old memory work in with the new CPU? (my guess would be "no")
Your old memory will work with your new cpu, to 99%. AMD improved memory compatibility with software updates, and the memory support in the 2000 line is supposed to be better again. If your ram is slower than DDR4-2666 you lose some performance compared to what you could have, but that was already true for the 1700 - DDR4-3000 is a good minimal ram speed target, as Ryzen profits a lot from being combined with faster ram.
Are the uarch people at AMD running around with their hair on fire trying to fix all their Spectre/et al. bugs, and skipped a release to focus?
AMD's plate is rather full. They're juggling a lot of cool projects despite having way less money than Intel or NVidia. Zen+ was always going to be a "low effort" minor update.
I'm frankly surprised that this minor update got much of a boost at all. Its a good sign for Zen2 next year.
Edit: and what is ECC status? Still not officially supported?
I was going to do that with AM3 but then they dropped the socket. and the upgrade path is quite big here if it's all on the same socket.
In the same time Intel went through 4 sockets: LGA 1156 -> 1155 -> 1150 -> 1151