Ryzen Z1's Tiny iGPU
chipsandcheese.com
chipsandcheese.com
I get it, it’s a CPU/SoC. But the end users care about FPS. Dedicate more of your die area to GPUs, and you’d crown the competition.
Steam Deck, and the PS5/Xbox chips are the sole exceptions; I wish those chips were available to consumers to buy.
Give me a quad core, or hexa core at most, with top of the line integrated graphics. The extra 4 cores are wasted die space for gamers.
The fundamental problem here is that GPUs require more memory bandwidth than CPUs, the platforms are created for CPUs, resulting in a pretty hard cap on APU GPU performance using the same platform. Generational GPU performance gains on APUs happen when the platforms update to faster memory.
The big break in this will be Strix Halo, coming from AMD either next year or late this year, as it will support 256-bit LPDDR5. Even using the fastest LPDDR5 available on the market, this will still just barely match the memory performance of Radeon 7600, the weakest current-gen discrete GPU AMD has on the market.
You can do better than that simply by putting some HBM or GDDR6 on the APU itself. Then you don't need to change the socket or have to add memory channels, and you don't need that much of it because you still have the ordinary memory slots to provide some DDR5 where the OS can keep all its dreck bloat. 8-16GB of the stuff would be enough to rival any of the current midrange discrete GPUs.
Wow, I had no idea integrated GPUs had come this far. This is fantastic news.
Also Apple laptops do have large memory busses. The Apple M1 Max chip for example has a 512bit memory bus.
I do know that overclocking VRAM (or at least GDDR6) barely increases performance, but that’s mostly because it’s usually already been pushed to the edge at stock. My old 3070Ti could easily overclock 10-15% higher, but the memory had to self-correct significantly more errors so the actual improvement was a wash.
I’m also curious if gaming handhelds would benefit from switching their entire memory architecture over to VRAM, like both consoles have done in this generation.
GPU performance is proportional to both bandwidth and compute.
Is that true? How are you measuring it?
A mini pc can be a good compromise but not the best at either.
This is just not true—I do the vast majority of my gaming on my laptop and only switch to the dedicated GPU for super graphics-intensive games like city skylines.
As usual, this attitude conflates "FPS-obsessed entitled divas" with "gamers"
Historically yes. You got 16gb or whatever on the card near the compute core. So what if you put 16gb of the same memory on the package with the APU instead?
Shout out to PCIe for unhelpful latency characteristics and barely adequate atomics support as my primary objection to discrete GPUs.
It's also plausible that games will become less demanding of the state of the art in GPU. If a playstation APU can play them beautifully, maybe a higher spec PC one will manage the same despite the OS and libraries overhead.
I think dedicated GPUs are probably on their way out.
I would be interested also because AMD has better support in Linux and their mobile GPUs are hard to find in laptops.
There's not that much demand for a mediocre-to-poor GPU that's more powerful than iGPUs of current era.
Ultimately we will see if/when the Strix Halo gets released, how priced it will be and how successful. I would likely buy a laptop with it.
- AM4 motherboard
- 5800X3D
- Aftermarket low profile CPU cooler
- 3600MHz 16GB RAM
- RTX 4060 Low Profile
- 2TB NVMe SSD
And it’d be smaller and much more powerful than a PS5. If you get the parts secondhand, it wouldn’t be that much more expensive either, especially once you factor in PS+ and how cheap Steam games are.
The only hairs in the soup are SteamOS and Nvidia drivers.
SteamOS doesn’t yet have a standalone install. You can work around that by installing ChimeraOS and having it launch Steam Big Picture on passwordless boot.
Nvidia currently has relatively poor open source drivers, but nvVLK is rapidly improving, with probably good performance in a year or so. Your other option is hoping AMD or a board partner release a low profile card, but I don’t see that happening. Or the dark GPU horse that is Intel.
You could also go for a larger case, which would allow you to fit a bigger GPU. Something like an AMD 6800 XT would deliver a 120FPS experience at 1440P and high/ultra settings, and do so quietly with an undervolt.
The Ryzen Z1 Extreme uses AMD’s high-end Zen 4 APU configuration, with eight Zen 4 cores and six RDNA 3 WGPs.
The ROG Ally uses it and you can dedicate 50% of the memory to the GPU and it performs quite well.
The RX 7600, AMD's lowest-end current-gen graphics card, has several times more CUs and memory bandwidth, so I think there's some room for improvement without undercutting their dGPU business.
The Z1 Extreme is essentially the same chip as the 7840U, AMD's highest tier U-series (15~30W) chip. In other words, it's a general-purpose design, not specifically gaming oriented.
so it is better to think that only a part of the cpu is dedicated for running the game, so you'd need more cores than otherwise.
Honestly, the whole Z1 lineup puzzles me a bit, since it feels like just a re-badge of the regular mobile chips with some slight tuning. If AMD was going to go through the effort of making a gaming focused chip, they should lean into that with more emphasis on the GPU side, and perhaps some extra cache or more memory bandwidth. Maybe even drop a couple of CPU cores. What we got instead just feels like a cheap marketing ploy.
The issue is that iGPU performance has been hardly changing for probably close to a decade, meanwhile dGPU performance has significantly risen (despite using older process/nodes at some times). What we’re seeing now is a correction. The meteor lake iGPUs apparently trade blows/are very similar to the 780m.
More specific to the 780m, its performance is apparently close to a gtx1060… a midrange gpu from 2016, about 7 years ago. It’s very nice to run a decade old game like gta5 at high, but modern titles like genshin can’t even hit 30fps on medium at 1x scaling. Also, not to mention that DDR is much slower than GDDR.
So: is the gpu better than previous gens? Definitely. Is it overpowered? No, previous gens have been underpowered.
this is what I mean by the unconscious pro-AMD bias that people regularly engage in. not only is that not true at all (Iris Pro 6200 and Iris Pro 580 both are significantly faster), but actually Crystal Well is a very interesting/prescient design in hindsight, it is the type of "playing with multi-chip modules and advanced packaging" thing people get super excited about when it's AMD.
https://www.anandtech.com/show/6993/intel-iris-pro-5200-grap...
https://www.anandtech.com/show/9320/intel-broadwell-review-i...
https://www.anandtech.com/show/10281/intel-adds-crystal-well...
https://www.anandtech.com/show/10343/the-intel-skull-canyon-...
https://www.anandtech.com/show/10361/intel-announces-xeon-e3...
> allowing the APUs to work together with external GPUs
nobody does that because it sucks. best-case scenario is when your dGPU is roughly similar to your iGPU performance so you get "normal" SLI scaling... ie like 50% improvement. And in any scenario where your dGPU isn't utter trash, it completely ruins framepacing/latency etc.
> and with the 5700g (which imho was a milestone) we finally had 30fps+ for all but the most brain-dead games (alan wake i'm looking at you!)
that's actually because the 5700G is still using the 2017-era vega design, which was not that advanced technologically when it was introduced. AMD dropped support a while ago but the writing was on the wall for a long time before that.
AMD should have moved 5000 series APUs to at least RDNA1 if not an early RDNA2. The feature deficit is significant, a 2017-era architecture being sold in 2023 and 2024 (already unsupported) is just not competent to handle the basic DX12 technologies involved.
Alan Wake didn't do anything wrong, AMD cheaped out on re-using a block and it doesn't have the features. Doesn't have HDMI 2.1 or 10-bit or HDMI VRR support either... and the media block is antiquated.
and if you'll remember all the way back to the heady days of 2017... AMD bet on Primitive Shaders and could never get that to work right on Vega. PS5 actually has a primitive-to-mesh translation engine that does work. RDNA1 still does better than Vega but RDNA2 is where AMD moved into feature parity with some fairly important architectural stuff that NVIDIA introduced in 2016 and 2018.
https://www-4gamer-net.translate.goog/games/660/G066019/2023...
Again, people love to jerk about Vega (gosh it scales down so well!) but honestly Vega is a perfect Brutalist symbol of the decline and decay of Radeon in that era. There is no question that RDNA1 and RDNA2 are incredibly, vastly better architectures with much better DX12 feature support, much better IO and codecs and encoders, etc. Vega kinda fucking sucks and it's absurd that people still simp for it.
The "HBM means you don't have to care about VRAM size!" and other insanely, blatantly false technical marketing just sealed the deal.
But when AMD rakes-in-the-face with Vega or Fury X or RDNA3 it's "a learning experience" and maybe actually just evidence of how far ahead they are...
Most of the various controls are sitting right there in Linux (or other OS) for power control of CPU & GPU. If cooled, could we just crank a Z1 Extreme up to a 100W core and have it be like a 7940HS? Or is there really some power binning differences?
I don't know what AMD charges for their cores. With Intel, there's been a decade of the MSRP of ~$279 for a chip, but the chips coming in a variety of different sizes across the power budgets; you'd pay the same for a tiny ultra-mobile core as you would for a desktop core. What we have now makes that look semi-ridiculous. It's the same chip. Different power budgets.
I think it shares the die with the 8700G which runs at 65W.
And that only begins to dig in. I was asking about overclocking, which is going beyond the base TDPs. Here's the 8700G's 780M gpu hitting 156 watts on overclock, for the same die: https://www.tomshardware.com/pc-components/gpus/amds-radeon-...
I highly highly highly encourage blowing away any thinking that a chip says it's TDP is X watts so that's what we'd get. This chip has been seen in the wild drinking vastly vastly more power if given the chance & settings to. I think my question still stands, is there any binning or real difference that would keep a Z1 Extreme from doing similarly? Or are we just bound by how much power we can put in and how much heat we can take away from the Z1 Extremes out there?
I have a handheld with a 7840U (GPD Win Mini), and I love it. I do use it mostly for games, and I suppose if it were labeled a Z1 Extreme instead of a 7840U, I'd be just as happy with it. So I can somewhat see where you're coming from. But also I think it's becoming more common to want to run "real" (non-gaming) workloads that can leverage a GPU on devices without a discrete GPU, so I still think it makes sense as a general-purpose part. (Also, I think the Z1 was an Asus-exclusive part, at least initially, so if there wasn't a non-exclusive variant, then I'd be stuck with something inferior.)
> ...and while they quote a 28W TDP, the actual max-power draw is stratospheric for a U series (well over 50 watts).
The ideal TDP for that chip is around 18W, with diminishing returns after that. (I usually run mine at 7-13W depending on the game.) Beyond 25-30W, you get only marginal performance gains relative to the amount of additional power, so while it technically can use over 50 watts, it's clearly not designed for that and the extra performance isn't worth it when you're running on battery.
Yeah, I know the 8x00G exists, but it's kinda too little too late.
Let's say that the CPU + powerful iGPU cost 95% of what a discrete CPU and GPU cost - but now you can't buy them separately, can't upgrade them separately, etc. You're less likely to get the mix of CPU and GPU that you're looking for since you can't select them independently. Why not just package the RAM with the CPU too? Apple's done that, but I think most people don't love that because it means they can't upgrade their RAM independently.
It also places constraints on how good something could be. Let's say that you produce new GPUs every 18 months and new CPUs every 12 months. Well, now you need to synchronize them. If the new CPU is ready to go, but the new GPU is 3 or 6 or 9 months out, what should your product releases be?
By having them separate, someone can buy the latest AMD CPU even though the next-gen GPU is 6 months out. When the next GPU comes out, they can buy that and upgrade the graphics and CPU on different cycles. Syncing up different product cycles isn't always easy.
I think the reason why is that they don't think there's likely a market. With things like a PlayStation or Xbox, it's going to (pretty much) have one set of capabilities for its 7-year lifecycle. You can integrate the CPU and GPU because there's only one buyer and because the CPU and GPU release have to be synced anyway for the console's release. With PCs, the release doesn't have to be synced and there are many buyers with different priorities.
The main advantage of APUs would be the costs (theoretically). If it ends up being 95% of the cost of a more traditional architecture, what would be the point?
To work around memory issues, these CPUs would need some onboard memory which would increase costs a bit, but the tradeoff is that it’d make for simpler, cheaper low-end motherboards. One can imagine a mini-ITX board with nothing but a CPU socket and a couple of M.2 slots that’d cost significantly less than current entry-level ITX boards. A full system upgrade could be performed by simply swapping out the CPU which would be great for non-enthusiasts; without a power hungry discrete GPU, power requirements are unlikely to increase meaningfully (and in fact are likely to decrease with upgrades), so upgrading wouldn’t necessitate a PSU change. These hypothetical boxes could easily stay relevant for a decade or more.
Higher end SKUs of motherboards for this type of CPU could have the usual RAM slots (acting as a second tier of slower RAM in place of swap), PCI slots, etc.
Intel's Lunar Lake mobile due out this year is supposedly using 16GB or 32GB on-package LPDDR5X RAM. Rumor has it it's 8533MHz. That'd be 68.2GB/s. https://www.tomshardware.com/tech-industry/manufacturing/int...
Still feels like we need more ram bandwidth. In general I think throughput is the new moat, the new market segmentation; consumer cores now have gobs of CPU and GPU (albiet we seem to have plateau'ed), but limited PCIe and ram bandwidth. We kind of started seeing USB4 compelling some more bandwidth, but I think even Intel is no longer offering USB4 on chip in many upcoming mobile chips, so that's kind of been defeated too.
Anyhow, AMD's next Strix Point is due end of year or there-after, but then their big Strix, Strix Halo has quad channels. That'll be exciting as heck. Folks may finally get their console crushers, perhaps.
Under the desk there's a linux box made with a Ryzen 5600G, that one does 20W at idle.
My daily driver is an M1 Max MBP which I’m happy with, but if I were to build a Linux productivity box (no gaming, that’s handled by a different machine), something like a 7840U on an ITX board with a cooler just big enough to practically never be audible would sound pretty great.
[edit] The 7900 delievers ~89% of the performance of the 7900X at 170W. 7700 & 7600 90% compared to the X counterparts.
I'd expect that you could probably get okay yields at okay costs if you are running a process that's a rev or two behind and making smaller chiplets that are then wired together after testing -- like the pentium pro's cache but for main memory to get 2 / 4 / 8gb ram all on "chip"
It'd probably cost more than a normal CPU, but the trade off is much more speed and the actual computer / motherboard at that point would just be a couple USB devices. You (the CPU maker) would grab a lot more of the per-unit profit.
Ah -- that's why they don't do it; no vendor would want their milkshake drunk...
Because it's not just the memory chip, also the interposer it's stacked on top of, which now needs to be bigger, which means you need to find more room in the PCB (today) or a bigger silicon interposer (likely in the near future) which reduces yields, and so on and so forth. If you wanted to have more than the very limited SoC RAM, you'd also then be looking at having multiple DRAM controllers, which also adds to surface area, and so on and so forth.
Generally the keyword you want to search for is “industrial motherboard”, some server motherboards have them too. Asrock Rack and Asrock Industrial and Supermicro are great.
https://www.asrockrack.com/general/productdetail.asp?Model=A...
https://www.asrockind.com/en-gb/IMB-X1231
https://www.asrock.com/mb/AMD/X300TM-ITX/index.asp
(I fucking love asrock rack’s design team, absolute madlads and some seriously impressive density etc. they’re out there making the designs people don’t know they want, romed8-2T and genoad8x-2T are fantastic.)
Mini-box themselves sell dc-dc converters and “picoPSUs” that do this as well, although idk if they sell one with pcie power plugs. The M350 is a surprisingly high quality case for the price, and while my first picoPSU failed almost immediately (first shutdown iirc) they warrantied it no problem. A very funny trip through some oracle branded ordering/invoicing framework. Good people, good products.
My one criticism is that there is very obviously a lot of psu noise. I had a dell laptop and my pc desktop and the picopsu on an audio push-button switch and I kept getting a ton of ground loop and couldn’t figure it out etc and finally noticed it stopped and then came back as I plugged the M350 from the audio switch. Iirc it also showed through usb dac as well. I had the DIN plug brick and I think it has a lot of noise.
Intel NUCs are also extremely high quality implementations and have a great aftermarket with brands like akasa and hdplex etc.
You're right that they likely won't make 2 different dies. Desktop AM5 chips will just get a package with some of the memory controller pins unconnected. The big question is whether they'll also package the full width laptop chip with on-package memory in a package for desktop that's incompatible with AM5.
If they don't, somebody is going to solder that monster laptop chip into an ITX motherboard. People will grumble about a motherboard that can't upgrade either the CPU or the memory, but if the performance is there they'll still buy it.
I don't think putting the memory in the package helps much with practically achievable clock speed or bus widths compared to just soldering the memory nearby. (Consoles and discrete GPUs aren't doing on-package memory despite running GDDR at significantly higher frequencies than LPDDR.) And given that, there's even less reason to expect a messy hybrid configuration with half the memory controller connected to on-package LPDDR5x and half routed to DIMM slots, whether or not it uses the AM5 socket.
What I would suggest (and my suggestions are worth about as much as I'm getting paid to write this):
AMD should create an I/O Die (IOD) with 512 bit DRAM memory width. Then create an AM5 variant with 128 of those connected to the socket and the other 384 tied off. Then create variants with 256 bit and 512 bits connected to on-package memory and no off-package memory pins. Sell those to laptop manufacturers.
A big motivation on laptops is that slow & wide memory uses a lot less power than fast & narrow; on-package also uses a lot less power than driving the signal between chips.
Then in early 2026 introduce AM6 with 4 channel 256 bit memory support. Sell AM5 & AM6 in parallel for a while.
Why not? They already do the equivalent for SRAM. The big cost of a wider memory bus is routing it through the socket and the system board, which you're not doing since only the narrower bus goes through there. The wider bus is solely within the APU package.
You could also take advantage of the additional channels -- have e.g. 8 memory channels within the package and two more routed through the socket, for a total of 10. Now if you have 8GB on the package and 16GB off of it, you have 10GB striped across all 10 channels and another 14GB striped across two.
Continuing to have memory slots also allows you to sell chips with and without on-package DRAM and use the same socket. High bandwidth memory only makes much sense if you have a strong iGPU, since CPUs are rarely memory bandwidth bound. But low end systems and high end systems wouldn't have that, only mid-tier ones would. The low end system has a small iGPU where HBM is both unnecessary and too expensive. The high end system has the same small iGPU, or none at all, because it uses a big discrete GPU.
HBM is also more expensive than ordinary memory, but Windows will sit there eating several GB of RAM while doing nothing, so having some HBM and some DDR should lower costs compared with having the combined total in HBM.
You use HBM or GDDR if you want very high bandwidth, like top-end discrete GPUs would get. But then the memory itself is more expensive and you want the external channels to reduce cost, so the OS bloat can go in the cheap memory and preserve the limited amount of high cost memory for what needs it.
This is notably not what Apple does -- they're just using ordinary LPDDR5 with a wide bus, equivalent to having a lot of memory channels. It gets them several hundred GB/s worth of bandwidth, similar to a midrange discrete GPU. If you were going to do that, you could put most of the channels within the package and still have two of them outside of it.
That sort of configuration would allow some flexibility. The on-package memory might have lower latency (if they're both just ordinary DDR this isn't going to be much difference if any), but if you configured the system to only interleave between the on-package memory channels then the "close" memory could achieve that lower latency. Interleaving the external channels into the same pool would have a small latency hit but increase bandwidth by e.g. 25%. Which could be configured in UEFI based on your expected workload.
1) physical space. There isn't a ton of leftover room once you've added compute and IO.
2) segmentation. A sufficiently powerful iGPU would cannibalize sales of AMD's own discrete GPUs.
3) board partners. A sufficiently powerful iGPU would rob AMD's GPU board partners of sales.
4) heat. Keeping the two hot things apart from each other makes both of them easier to cool.
5) memory bandwidth. Even the dinky iGPU in my Ryzen 2200G is heavily constrained by RAM clocks.
FWIW, I have a Ryzen 2200G hooked up to my TV, and it's totally adequate for casual gaming.
However, on laptops which aren't constrained by backwards compatibility, Strix Point Halo appears to have both a beefy GPU and a 256 bit memory bus.
As others have pointed out, you can't fit enough memory bandwidth though the AM5 socket to feed a powerful iGPU.
Fixing this problem kind of requires abandoning the current desktop form factor and switching to a unified module with both the CPU and soldered-on memory. Though at that point, the motherboard is doing little more than power regulation and breaking out all the IO to various ports.
At that point, does it still count as a desktop form factor?
Apple seems to think so
Does it though? What stops you from putting some HBM onto the APU package and still installing it into the AM5 socket? It wouldn't even preclude you from continuing to use the memory slots, that memory would just be slower than the on-package memory.
HBM memory is expensive, it requires a huge amount of extra IO on the die and an expensive silicon interposer. And you still need to keep around the old IO for the DDR5 memory. All that drives up costs, for what would still be a mid-range GPU.
Also, current software and games doesn't know how to deal with two pools of memory that have different performance characteristics, so the hardware would be underutilised.
The design which abandons the AM5 socket and switches to using much simpler and cheaper soldered-on gddr6 memory just ends up being cheaper and avoids the thermal and area limitations of the AM5 socket, so it could probably compete with high-end GPUs. It's just a better product direction.
AMD will either stick with their current strategy of APUs with their low-end GPUs because they would rather sell both a CPU and a dedicated GPU, or they will skip straight to a new form factor. The middle ground of trying to add more memory bandwidth to an AM5 style package just doesn't make any sense.
It's a product that replaces both a midrange CPU and a midrange GPU, which together have not only all of those costs but also the cost of needing two separate packages -- a CPU package for e.g. AM5 and a PCIe package for the GPU. Putting them together costs less.
> Also, current software and games doesn't know how to deal with two pools of memory that have different performance characteristics, so the hardware would be underutilised.
Except that's exactly what they know how to do, and they do it already. They expect there to be a slower pool of memory on the CPU and a separate faster one on a GPU. The system could expose them to existing applications in the same way -- the fast memory via the iGPU and the slow memory via the CPU. That's just software.
You could even do better if you have e.g. 16GB of fast memory and your game only needs 8GB of VRAM, because then the other 8GB can be used as L4 cache for the CPU and could plausibly fit the entire working set of the game in it while the only thing in DDR5 is the OS and idle background apps.
Meanwhile newer code which is aware of the configuration could make better use of it, e.g. by not having to worry about the cost of "copying" between GPU memory and CPU memory via PCIe since they're actually both directly connected.
> The design which abandons the AM5 socket and switches to using much simpler and cheaper soldered-on gddr6 memory just ends up being cheaper and avoids the thermal and area limitations of the AM5 socket, so it could probably compete with high-end GPUs.
That isn't cheaper, because then 100% of your system memory would have to be GDDR6. In a mid to high end system that's going to be dramatically more expensive than continuing to have a socket with DDR5 memory channels, because you'd have to replace e.g. 64GB of DDR5 and 16GB of GDDR6 with 80GB of GDDR6.
Meanwhile AM5 supports TDPs up to 170W and sTR5 up to 350W. 170W is reasonably sufficient for the combination of a midrange CPU and midrange GPU -- a midrange CPU is typically ~65W and a midrange GPU ~150W, which together hypothetically exceeds 170W, but a workload that simultaneously maxes out the GPU and all cores of the CPU is uncommon. In that rare case you would simply clock them slightly lower. Knocking the TDP of a 65W desktop CPU down to that of a laptop in that rare circumstance would have a relatively minor performance impact:
https://www.anandtech.com/bench/product/2685?vs=2665
350W would be sufficient at the high end for the same reason.
And if they were going to design a new socket (as happens from time to time anyway), they could give "AM6" a larger footprint and higher TDP without omitting the valuable external memory channels.
But creating a new CPU interface is expensive -- all the OEMs have to design new boards. So if your concern is cost then it makes more sense to wait until the next time you were going to do it anyway, e.g. when DDR6 becomes a thing, and do something else in the meantime.
Also the RAM bandwidth just isn’t there, and special mainboards with more memory channels eat up the cost advantage. And they’re hard to cool.
Does seem like they finally plan to do this with the AMD Strix Halo which looks to hit somewhere late this year or early next.