AMD Announces 7950X3D, 7900X3D Upto 128MB L3 Cache
anandtech.com
anandtech.com
Note how the 7800X3D, a single CCD SKU, has an advertised maximum clock of 5 GHz, compared to the 5.6/5.7 of the two dual CCD SKUs (the 7700X below it advertises 5.4 GHz). Don't expect the 3D cached-cores to exceed 5 GHz on those other parts.
This means that the performance profile of the two CCDs is very different. Not "P vs E core" different, but still significant. If you run (most) games, you want to put all their threads on the first CCD. If you run (most other) lightly threaded workloads, you'd often want the other, non-3D CCD.
This seems like a rather significant scheduling nightmare to me.
Intel did similar work with E/P cores.
Indeed, that's Lakefield as seen in the Intel Core i5-L16G7.
The Scheduler changes (in Windows 11 at least) make it extremely pleasant to use on the Lenovo X1 Fold, even if there're some drivers problems (fortunately easy to fix: https://csdvrx.github.io/ )
That seems to answer my (sibling to yours) comments question. Perhaps the 7800x3d is the only sku unbothered by older OSs, then.
It would be annoying for kernal developers to make an exception for individual processors like this, but it's totally doable to make the scheduler handle this case.
With X3D vs non-X3D cores which one is going to be faster depends on the workload. Most games benefit (sometimes drastically so), a lot of other stuff doesn't / is slower roughly in proportion to the clock reduction - just look at the 5800X and 5800X3D of the prior generation to see how split the benchmarks are; virtually every synthetic benchmark and most intensive things - compiling, encoding, running JS and other interpeters - suffer, games win.
I suppose a good heuristic for desktop users would be to pin processes using the GPU to the X3D cores.
Would be interesting to see at comparable clocks what the performance uplift of hitting the huge cache more often on the other ccd vs hitting main memory.
Now, thinking about it... you can't evict to the distant L3, though. So, it really depends on the workload and whether the remote big L3 is "warmed" suitably for you.
but also for most daily desktop workloads the difference is irrelevant tbh.
games already should be pinned to one CCD for max performance
compilation and similar will optimally spawn across both CCDs making it not matter (through you want to only split parallel computation units across CCDs but parallel code gen for the same unit should be on the same CCD)
So I don't thinks it matters outside of artificial "challenges" to try to maximize performance beyond what is normally reasonable for a desktop system (due to work/time involved in doing so).
I know there has been such designs in the past but I don't know how it works in the Ryzen CPUs.
At the same time, that latency is still peanuts compared to hitting main RAM.
The die with the cache probably has better latency (provided the cache doesn’t connect through the IO die), but lower clocks making it better with memory limited workloads.
The other die will be better at non memory bound work, but should still be much better than normal at memory bound tasks too. I suppose it remains to be seen if lower latency and lower clocks beats higher latency and higher clocks, but I suspect 10% higher clocks won’t compensate enough for cache hits being several times faster.
All in all, this sounds more like a technically really interesting challenge, rather than a high-stakes bet. Even a "bad" scheduler probably won't do much harm.
Realistic worst case, quite a few workloads that might have benefited from the new cache don't... but probably aren't hurt much either.
On linux you can use 'taskset 0x1 ./bench' to force a process onto a specific mask of cores. Sounds like this sort of thing will be increasingly necessary if you want to compare code variants on these CPUs with precision.
Regardless though if you want reliable numbers, even on a CPU with uniform cores, you should be doing multiple runs and comparing average and variance. This will always tell the actual story regardless how the platform underneath behaves that day.
A 5GHz CCD with +20% would then be only barely faster than a 5.7GHz CCD with no 3D, at which point the pairing you suggest makes little sense.
Besides, 5800X3D attained 5.6GHz so there’s no big reason to think 7000X3D would be capped at 5GHz.
It's way easier then big.LITTLE or current Intel CPUs with E and P cores, so it should be a non issue.
It's only a relevant significance when you want to maximize performance but most desktop tasks, including many games, do already perform to a degree where you are unlikely to notice the difference without profiling or being very very sensitive.
But then for gaming having more then 8 performance cores is currently useless for most games (assuming you don't do anything "costly" in parallel). It's to a point that for many gaming benchmarks disabling one CCD can lead to slightly better results on a 7950X (due to more thermal and power budged for the other CCD).
Or in other words if you only care about gaming don't go for a 2 CCD CPU it's not worth the money (similar for Intel go for 8 P-cores but going for more E-cores has most times dimishing returns AFIK)
And if you do other things:
On windows AMD will work with windows to make games run nicely.
On Linux it anyway can be a good idea to pin you game to one CCD manually (even for the 7950X; using cgroups, e.g. through systemd). As this should implicitly lead to other services being run on the other CCD (simple due to the first already being utilized highly).
Now highly scalar applications which use both CCDs fully could have some interesting dynamics I guess. But again it should be much more predictable then doing similar across P & E cores.
But there is one fun use-case where this looks awesome for: Embedding a gaming VM in you system. (Dedicate all of CCD1 to the VM, use IOMMU for GPU and NVME pass through).
If we don't do core pinning, would the "whole CPU passthrough" be sufficient to preserve the scheduler logic in Windows (and any such logic in Linux, if it is or will be available in the near future)?
Alternatively, as you suggested, one could feasibly allocate one CCD entirely to Windows and another to Linux, to optimize each one for performance in gaming and productivity correspondingly. Though for heavy multi-core tasks, passing the whole CPU to Linux would still be preferable, I think; same for Windows if playing strategy games. So this approach might not suit all cases equally well.
I look at this as Zen 5 test bed. Microsoft already said scheduler tweaks will come in Win 11 to better accommodate Zen 4 X3D, and I assume Linux kernel devs have been prepping for a bit.
I agree that Intel is already having to play with optimizing BIG.little with Alder Lake and Raptor Lake. Functionally this is kind of the same for AMD, where you want that cache hungry application all on the same core complex. Not exactly apples to apples, but it's at least still fruit.
This is a good opportunity to get a year or more of battle testing this stuff in the field to ensure its at least up to snuff, as well as plan future iterative improvements.
I believe the 7950X3D will have its voltage locked, but in terms of perf/W should still improve over the 5950X. And the 5950X's power efficiency is fantastic.
The Ryzen Pro processors are also fantastic. The 5750GE was a decent clocked 8-core that peaked out of the box at only around 39W, supports ECC and DASH, other nice features like Memory Guard, Shadow Stack, etc. Those are great in 1L form factors where you can build capable physical clusters where nodes can idle at only 9W total power draw.
Some Raspberry PI SoCs also had the RAM soldered on top.
AllWinner V3s, S3. Theres a SAMA5D2 SiP. Bouffalo BL808 (featured on the Pine Ox64). There's a lot a lot more. I think there's a couple with even more memory too.
Intel's Lakefield, with Foveros stacking, was an amazing chip with 1+4 cores and on chip ram. High speed too, 4266MHz, back in 2020 when that was pretty fast. This is more for MID/ultrabooks, but wow what a chip, just epic: add power and away you go. Ok not really but not dealing with routing (and procuring!) highspeed ram is very nice.
Intels been doing such a good pushing interesting nice things in embedded, but the adoption has been not great. The Quark chips, powering the awesome Edison module, had nice oomph & Edison was so well integrated, such an easy to use & so featureful small Linux system... wifi & bt well well well before RPi.
It would be fun to see DRAM-less computers but I more imagined them being big systems with a couple GB of sram. There's definitely potential for low end too though!
I have been lurking on the Ox64 for while but I need a few more green lights:
- Is the boot rom enough to fully init the SOC? Aka, I don't need to run extra code I would need to include in my keyboard firmware on the sdcard.
- The hardware programming manual misses the USB2 controller with its DMA programming. Even with some SDK example, you would need the hardware programming manual to understand properly how all that works.
- I want to run my keyboard firmware directly from the sdcard slot, and that directly on the 64bits risc-v core, possible? (no 32bits risc-v core).
I'm not the most familiar with it, but I believe all hardware init (setting clock source, initialing USB, GPIO, etc.) is handled by the flashable firmware of which there are open source SDKs for.
I guess this is a "standard" flashing protocol over usb, enabled by the right button pressed at power on (plugging the USB cable). Would I need to including the flashing support code into my keyboard/SDcard loader firmware or is it handled separately by a different piece of hardware?
Any specs on the format of the firmware image, to know which core will run the real boot code?
Erk... soooo many questions in the wrong news :(
Edit: to clarify, there's a bug in the bootrom that prevents the initialization of the USB device. Newer revisions of the Ox64 may fix this.
- how do you run anything with that board if you cannot flash anything, I don't understand?
- I cannot use it as usb keyboard controller because of a bootrom bug? (power/data via usb-c)
Now I am confused.
The bootrom bug only prevents you from flashing via the on-board USB-c port. You can use a separate USB <-> UART device plugged into GPIO pins to boot/flash.
This is weird.
If resellers with stocks know that, they will send back the boards and ask for a refund.
meh.
The main distinction between application processors that can run Linux and microcontrollers that use onboard RAM (and often Flash) is that the former have an MMU. It's attractive to imagine that your SBC might only need something as simple as a DIP-packaged Atmega for an Arduino, and I can imagine a system-on-module - actually, saying that, I think several exist, ex. this i.MX6 device with a 148-pin quad-flat "SOM" with 512 MB of DDR3L and 512 MB of Flash:
https://www.seeedstudio.com/NPi-i-MX6ULL-Dev-Board-Industria...
Whether you consider that Seeed branded metallic QFP (which obviously contains discrete DRAM, Flash, and an iMX6) to be a single package, while a comparably-sized piece of FR4 with a BGA package for each of the application processor, DRAM, and Flash on mezzanine or Compute-module style SODIMM edge connectors would not satisfy your desire for an embedded Linux processor with less routing complexity, I don't know. They build SOMs for people who don't want to pay for 8 layers and BGA fanout all the time.
I don't think there are enough applications for embedded systems that need 128M of onboard SRAM that won't support the power budget, size, complexity, and cost of a few GB of DRAM.
L3 cache is orders of magnitude faster than using RAM.
You're talking a maximum of 50GB/s for DDR5, versus 1500GB/s for L3 cache
https://en.wikipedia.org/wiki/List_of_interface_bit_rates#Dy...
https://meterpreter.org/amd-ryzen-9-7900x-benchmark-zen-4-im...
It's a paradigm-shifting increase in processing speed when you don't need to hit RAM.
There is a use case when you can improve performance by keeping compressed (LZ4) data in RAM and decompressing by small blocks that fit in cache. This is demonstrated by ClickHouse[1][2] - the whole data processing after decompression fits in cache, and compression saves the RAM bandwidth.
[1] https://presentations.clickhouse.com/meetup53/optimizations/ [2] https://github.com/ClickHouse/ClickHouse
Size? But then vendor could just ship the CPU+RAM stacked on top of eachother.
But how much die real estate for 8GB/16GB of sram in such fantasy world?
Then 8GB sram with a modern CPU, Zen4 for instance, is a die of ~ 9 top-of-the-line GPUs dies.
And now, with 3D? ... mmmmmh...
What is the size of the apple M2 die already?
It's also extremely niche to have a workload that requires such high CPU performance, but that it would fit including a linux OS in 128MB. Usually something like that is FPGA or DSP territory.
I think what you want is a cheap ARM CPU with DRAM stacked on top of it on the same package (which exists).
I dunno the only direction I care for at the moment is TDP
This is with clamping down TDP on the 7950x to roughly 110W. It's an absolute beast.
[1] https://old.reddit.com/r/Amd/comments/xzj69v/is_anyone_runni...
Since I was scared off 128 GB by all those complaints, I built a 7950X system with 2*32GB memory, and it runs well/stable at its rated 6000 MT/s with EXPO on (also, AMD says 6000 MT/s is the sweet spot for AM5 memory interface, so I am happy).
Motherboard: ASRock X670E Steel Legend
RAM: G.Skill Trident Z5 Neo 64 GB (2 x 32 GB) DDR5-6000 CL30
Here is another thread I remember reading: https://old.reddit.com/r/hardware/comments/za0x0q/level1tech...
Here's my stream results (with the cpu in 105watt "eco" mode):
$ wget https://www.cs.virginia.edu/stream/FTP/Code/stream.c
$ gcc -O3 -fopenmp stream.c
$ a.out
Number of Threads counted = 32
Copy: 45571.7
Scale: 40672.0
Add: 45317.3
Triad: 42759.0
With array size 8,000,000 Copy: 31551.3
Scale: 30983.7
Add: 34403.6
Triad: 34484.17950x X670E Steel Legend, standard mode, 2*32 GB = 64 GB RAM at 6000, CL 30.
Array size = 1,000,0000
Copy: 61488.8
Scale: 55905.4
Add: 61320.2
Triad: 57706.5
Array size = 8,000,0000
Copy: 43631.0
Scale: 43062.7
Add: 46782.6
Triad: 46966.8All of the web apps I build inside of Docker with WSL 2 reload in a few dozen milliseconds at most. I can edit 1080p video without any delay and raw rendering speed doesn't matter because for batch jobs I do them overnight.
Writing to disk is fast and my internet is fast. Things feel super snappy. I've been building computers from parts since about 1998, I think this is the longest I ever went without an upgrade.
I did like you: rocking a Core i7-6700 from, what, 2015 up until early 2022. 16 GB of RAM, NVMe PCI 3.0 x4 SSD, the first Samsung ones. I basically build that machine around one of the first Asus mobo to offer a NVMe M.2 slot.
It was and still is an amazing machine. I gave it to my wife and I'm now using an AMD 3700X since about a year and... I'll be changing it for a 7700X in the coming weeks (hopefully).
The 3700X is definitely faster than my trusty old 6th gen core i7 but Zen 4 is too good to be true so I'm upgrading.
All this to say: you can stay with the same system for seven years then upgrade twice in less than 12 months!
Agreed, but not for testing, please. Too much stuff out there already seems to be designed for or tested on what might as well be supercomputers like the above, and then get shipped out to run on Grandma's 7 year old <Misc Manufacturer> laptop.
Maybe some group should get consensus about modern monitor 'paper white' so at least everybody has a daily monitor in same setting no matter how good/bad his monitor is.
I think for general use most desktop and workstation users are going to get more benefits from faster random access and won't notice very high linear access speeds. I have two recent SSDs and one of them has a median access time of 12µs and the other has a median of 30µs. Even though the latter has gaudy benchmark numbers, can stream at many GB/s and can ultimately hit almost a million IOPS, higher-level application benchmarks lead me to prefer the former because random access waiting time is more important.
Something more general might be game asset load speed. Those are often sequential reads. They put a ton of engineering effort into the latest consoles simply to improve that one thing.
Most of that effort was to ensure the processors could actually ingest data at the speeds that off the shelf SSDs could deliver it. The Xbox Series X shipped with what was a low-end NVMe SSD at the time, and the PS5 used a custom SSD controller primarily so they could hit their performance targets using older, slower flash memory rather than being constrained by the supply of the faster flash that was just reaching the market at that time.
And there are still hardly any games that even make a serious attempt to use the available storage performance.
I've been holding out for a long time.
Can't wait for my code to compile 10x faster and game at 10x the fps.
Kind of disappointed an Intel NUC from 2015 handled the displays better.
7950x is more efficient than 13900k, so that's your answer. Idle usage ryzen is often worse than intel though, so pick your poison.
Its a fantastic CPU.
Also throw the guy a few bucks if you're on MacOS
Going to need to wait for bench marks.
The 7950X3D makes for a phenomenal server chip (perf/cost).
Do any cloud providers beyond OVH/Hetzner offer these "desktop" class chips available for hosting use cases?
It's worth noting that AMD also uses FinFET transistors, which are physically 3-D [0, 1] (compared to MOSFET transistors, which are planar / 2-D).
[0]: https://en.wikipedia.org/wiki/Fin_field-effect_transistor
[1]: https://www.amd.com/en/press-releases/amd-demonstrates-2016j...
By pulling the chips back to 120w, it looks like!
Of course, 4/5s of what I play are indies that would run on a potato, so it's probably not worth upgrading. I nevertheless probably will build a new computer later on in the year, and a 4080 / 5800x3D combo is awfully tempting ...
But for games cpu can improve game performance, especially when the single core perf is upgraded
I don't expect them to make this kind of change anytime soon.
Good thing UPS is a unique acronym that isn't used elsewhere when discussing computers :)
AAAAA
(anti-"ambiguous acronym" advocates anonymous)
I fully expected my data management software to be much faster on the new hardware since it uses multi-threading to make use of all those wonderful cores, but it blew me away just how fast it is. https://www.youtube.com/watch?v=OVICKCkWMZE
I might have to bite the bullet later this year if the 7950X3D benchmarks look really good once it releases.
I'm on a 3900x and 2070 Super. I mostly reduced using the desktop when I got the M1 and M1 Pro devices. I now use it remotely from another room in the house. I'll hold out despite the temptations, and only update in 2-3 more years from now.
holy fuck it's horribly expensive these days.
I'd love for me to be able to just run all my builds on full blast on a 24 core beast at home, but if your interests also include games, you're looking at multiple thousands just to upgrade. Needless to say, my apartment, family, vacations and a dozen other things are well above in the priority list.
* $500 for a CPU (7800X starts at $450)
* $60 for a cooler because the Wraith Prism is shit.
* $200 for a motherboard, maybe $100 if you want to try to cheap out and get something that might have unstable voltages, shit PCIe bandwidth or no NVMe slots, etc.
* $200 for 32GB of quality DDR5-4800 (or 5600 if you want to be fancy), and that's easily used up these days.
* $200 for a quality 750W power supply
* $100 for a case with good enough airflow, assuming you don't already have one.
* $200+ for at least 1TB of NVMe storage, easily much more
So, assuming a new build that didn't get incremental upgrades in the past, building a new, powerful PC these days is going to run you $1500. Without even picking any top of the line stuff. Guess what didn't get included in there ? GPUs with their bloody ridiculous prices. If you're going with NVidia, this 4000 generation is a waste if you're buying anything but the 4090. You could absolutely buy a 4080 (or rather a 4070Ti), or a 4070, but they're such a horrible deal in terms of price/performance. And that's going to cost you at the absolute least $800 (for a 4070, which is a dogshit card). Or you can try to find a series 3000, but that's also going to run you $1000+. If you're going with AMD, your problems are similar, for cards that are really subpar. As for Intel, well, let's just say an A770 with your high end CPU might cause a few bottlenecks here and there.
So, yes, if you're lucky enough to find deals _and_ to have stuff that you can still use from an old, recent rig, sure, building a new PC isn't _that_ expensive. If you have to do major upgrades, you're looking at multiple thousands. Pulling that much money out in one go for something that is ultimately not extremely necessary is something that can only be afforded by a very small percentage of people.
And if you already have an older rig… just bring it over from that.
So, no, any upgrade not done in the last 5 years means pretty much a full rebuild now. New sockets + DDR5 means that your old build is going to hit a brick wall.
Your old PSU might be insufficient for the new build, but it's not like 750W PSUs are a new thing (or particularly required if you're not doing a high-end build, even today). Power efficiency tends to yoyo around over the years.
No, you absolutely do not "need" to upgrade your GPU, and certainly not immediately. Your old GPU will not be doing a worse job in the new computer than it was in the old one. And you can always replace it later on, should the need actually arise.
I think a core part of the misunderstanding here is around expectations. The build plan for "a top-spec gaming computer" is going to look pretty different from "a top-spec dev computer that I can also game on". But, if anything, gaming has hit seriously diminishing returns in the last couple of years. It's quite hilarious and sad to see NVidia try to make real-time path tracing and 4K gaming into things, in a desperate bid to make GPU performance a relevant factor again.
The only compromise compared to the last gen is unfortunately the RAM - 64GB this gen compared to 128GB last, until they sort out DDR5 4-DIMM dual-rank configs..
So you need to be paid ~$30/h to recoup the cost in a year.
Of course if you're contractor it means you will... earn less unless you bump the price but that's the "motivation" of being paid by hour...
Pretty motherboards with 110% of desired featureset: Insert Fresh Kidney and/or Lung
I just built an i9-13900K replaced the LGA1366 2P board in it. It's an astounding performance bump, to mildly put it lol. I plan on using this chassis with this setup for probably another 5-10 years before upgrading again.
refs?
Case I’m using is a SC743TQ-1200B-SQ
I love it. I have 6 more of them that I use for various things (3 are dedicated to my plex cluster).
The case having a transparent side panel, on the other hand...
So I tend to end up buying something on the higher end even if I'm not using all the board's features. My current tower uses an ASUS ProArt X570 Creator.
My last build: MSI MAG X570 $200, Ryzen 9 5900X $400, G.Skill 2x16GB DDR4 3600Mhz CL16 $165, Samsung 980 Pro 1TB $150 = $915
That's not a budget build by any means, but it's also not crazy when compared to previous builds in that tier over the past 15 years.
That wasn't the first computer I owned, just the first one I built. I am sure some really old timers remember the ones that cost more that $5K.
For the base model.
I built about half a dozen machines from parts since 1998 and it was almost a constant to spend $650-800 for a mid-range machine for about 20 of those years. This includes everything but a monitor, but it does include a decent video card.
His part list for $915 doesn't include a video card. With today's market a mid-range card will put you at a grand total of about $1,250-1,300. That's approaching 2x the cost of a solid mid-range machine that you could build 8 years ago.
Here's a couple of line items from my last part purchase in 2014:
- Intel Core i5-4460 LGA 1150 CPU $182
- Crucial MX100 256GB SSD $105
- Cooler Master Hyper 212 EVO $35
The crazy thing is, here we are 8+ years later and that same CPU cooler is $45, that's ~23% more expensive than almost a decade ago.The exact CPU and SSD aren't worth price comparing because no one would buy them today since we have way better hardware at similar price points. For ~$180 you can grab a Ryzen 7 5700G and a 1TB SSD is about the same price as 256GB from back then. That feels like it's on a higher end of mid-range for today. It's really video cards where you get killed.
Spent $290 in April 2020 for a Radeon RX 5600 XT - a few steps below the top end $400 5700 XT!
The 6600 XT is actually about that much now, which isn't awful. But it's also increasingly far from the top end (behind the 6700 XT, 6800, 6900 XT, 7900 XT, 7900 XTX) and those range in price from $350 to $1000!
I lost my invoice for my video card but it's a GeForce GTX 750 Ti. I'm pretty sure it was around $150 went I got it. It's not the best thing in the world and today it's good enough to casually play some older games. It has no problem powering multiple 1440p displays which is why I continue to use it. For its time when I got it in 2015 it was maybe middle of the pack as a mid-range card. Playing those era games at 60 FPS at 1080p is fine.
Nowadays, forget about getting anything near that price point even though every other piece of hardware has huge upgrades at similar price points as back then (as seen with my previous post on CPUs and SSDs). It's especially bad too because almost 10 years has passed. Expectations have rose. Nowadays you would hope to be able to have a 1440p display running at 120+ hz and be able to play games at that resolution at 120+ fps.
Today you're looking at something like $300 to get a comparable card relative to what's available today vs back then. I haven't done the research on what's the best card in that range but a GTX 2060 is about $300 nowadays but it also retailed at about that price in 2019. In theory it should have dropped by now since there's been multiple new generations since then.
I've seen a bit of apologia for this, claiming that costs have risen so of course prices would increase as well. But prices are going up so much faster than inflation that I'm not sure that passes the sniff test. We should see 4080s for $850-$900, not $1200.
All of this just in time for 30-series supplies to dwindle ... new 3080s are back up to the $1k mark. Sigh.
4.35 a day isn't cheap, but I was spending more than that on starbucks everyday. So I stopped spending 10 dollars a day at Starbucks and opted for the rolls royce. I have a nicer rig, I'm healthier not drinking starbucks, and I'm saving a little money even if we factor in electricity costs.