In addition there's a interest in having a lot of memory for LLM acceleration. I expect both CPUs to get more LLM acceleration capabilities and desktop pc memory bandwidth to increase from its current rather slow dual channel 64bit DDR5-6000 status quo.
We're already hearing the first rumors for Medusa Halo coming in 2026 with 50% more bandwidth than Strix Halo.
GPU:s have existed about 30 years. Embedded ones for 20 years or so? Why are the embedded GPU:s always so stunted?
Why are the embedded GPU:s always so stunted?
Memory bandwidth. Besides LLMs, gaming on an iGPU will always be more expensive for the same performance as dedicated GPUs due to memory bandwidth.Before someone tells me consoles using iGPUs, keep in mind that consoles use GDDR as its main system memory which has slow access times for the CPU. In a non-console, CPU performance is important. GDDR is also power hungry so they can't be used as the main system RAM in a laptop form.
It is the thermal envelope that defines pretty much everything nowadays. Without active management of it chips would die a heat death very fast. Which also means chips are designed with a certain chip external heat management in mind. The more heat you can get out of a system and away from a chip, the more powerful you can design these things. And game consoles do have active cooling, i.e. they sit between desktop PCs and thin laptops, probably sharing the thermal handling capacity with larger gaming laptops, if anything.
On Desktop, upgradability is very popular and obviously the returns from the cooling on discrete GPUs are immense. With GPU dies costing so much, due to their size and dependency on TSMC, pushing the faster but hotter is probably a cost effecient compromise.
On Laptops with APUs, you currently ususally give up upgradeable memory - the fastest LPDDR is only soldered on (today), and the fastest solution would be on-die memory for bandwith gains that only really Apple is doing.
Marketing wise, low core count Laptops appear to be hard to sell. Gaming laptops seem to ship with more cores than the desktop you would build - the CPU appears out-specced. I think this is because CPUs are cheaper, but that means a high-end APU would also need large CPU to compete. Now you've got a relatively unbalanced APU, with expensive hot CPU and relatively hot iGPU crammed in a small space - cooling is now tricky.
This is going to be compared with cheap RTX 4060 laptops - and generally look bad by comparison. I think what's changing now to narrow the gap is Handhelds, and questionable practices from Nvidia.
The Steam Deck kicked big OEMs into requesting AMD for large APUs.
Nvidia seems to have influence on OEM AMD Laptops - Intel CPU and Nvidia GPU for years now seem to ship first, in larger quantities, and get marketing push despite CPU arguably being worse.
Intel despite their issues seem to raising the iGPU bar too - their Desktop GPU investment seems to be paying off, and might be pressuring AMD to react.
If you want a good value for your money you need modularity and competition for the modules. If it's a one package deal the companies will charge so much that it curbs secondary markets that could be created, which could add value to the product.
Because gpu want a lot of silicon. 5080 is 300 mm^2. Meanwhile ryzen 9xxx is 50 mm^2.
Meanwhile CPU wants to use that wafer space for themselves. And even if you use 100% of the wafer space for GPU you will have a small gpu and no cpu
Not really closer. igpus got good enough to kill the low the discreet market basically entirely, but they haven't been "getting closer" to discreet cards. Both CPU SoCs and discreet GPUs have access to the same manufacturing nodes and memory technologies, and the simple physical reality is that they can just be bigger and use more power as a separate physical entity, along with memory better optimized for its workloads.
In a next iteration AMD could look into doubling or quadrupling the memory channels and GPU die area like as Apple has done. AMD is already a pioneer in the chiplet technology Apple is also using to scale up. So there's lots of room to grow for even higher costs.
Historically those cards had narrow memory bus and about a quarter or less video memory of high end (not even halo) cards from the same generation.
That narrow memory bus puts their max memory bandwidth at a comparable level to desktop DDR5 with 2 DIMMs. At the same time quarter of high end is just 4GB VRAM which is not enough for low details for many games and prevents upscaling/frame gen from working.
From manufacturing standpoint low end GPUs aren't great either - memory controllers, video output and a bunch of other non-compute components don't scale with process node.
At the same time unified memory and bypassing PCIE benefits igpus greatly. You don't have to build an entire card, power delivery, cooler - you just slightly beef up existing ones.
tl;dr; sub-200 dollas GPUs are dead and will be replaced by APUs. I won't be surprised if they will start nibbling at lower mid-range market too in the near future.
that already happened like 5+ years ago. The GT 1030 never got an update, so Nvidia hasn't made an entry level GPU since. Intel kinda did with ARC, but that was almost more a dev board
Probably, not. Because it need dedicate channel on hardware level.
- GPU are mostly for streaming applications with large data blocks, so usually, CPU cache architecture is too different from GPU to simply copy (move) data, plus, they are on different chiplets, and dedicate channel means additional interface pins on imposer which are definitely very expensive.
So, when it is possible to make SoC with dedicated channel CPU<->GPU (or between chiplets), but usually it used only on very expensive architectures, like Xeon, or IBM Power, and not used on consumer products.
For example on older AMD products with APU, usually, graphics core have priority over the CPU to access unified RAM, but CPU cache don't have any additions to handle shared with GPU memory.
On latest IBM Power and similarly on Xeon, invented shared L4 cache architecture, where blocks of extremely huge L4 (near to 1Gb per socket on Power, as I remember somewhere about 128Gb on Xeon), could be assigned programmatically to exact core(s) and could give extremely high performance gain for applications running on these cores (usually these things very beneficial for DB or something like zip compressing).
Added: example difference CPU cache to GPU, for CPU usual size of transaction is less than 64bits, may be current 128..256bits but this is not common on consumer hardware (could be on server SoC), just because many consumer applications are not optimal to use large blocks, but for GPU normal to use 256..1024bits bus, so their cache definitely also have 256bits and larger blocks.
Mean, in CPU could cut any core and it will work completely separated without other cores.
In GPU, typical, have blocks for example 6x CUs, which have one pipeline for all, and this is how they achieve thousand CUs or more. So, all CUs basically run same program, in some architectures could make limited independent branching with huge penalties on speed, but mostly just one execution path for all CUs.
Very similar to SIMD CPU, even some GPUs was basically SIMD CPUs with extremely wide data bus (or even just VLIW). So, GPU cache sure optimized for such usage, it provide buffer wide enough for all CUs on same time.
So, when CPU access GPU memory, CPU just directly access RAM via system bus, but not trying to check GPU cache. And yes, this mean, could be large delay between GPU write cache and data actually delivered to RAM and seen by CPU, but probably, smaller than on discrete GPU on PCIe.
I sure hope so. We could use a new board form factor that omits the GPU slot. Although my case puts the power connector and button over that slot on the back so it's not completely wasted, but the board area is. This has seemed like a good idea for a long time to me.
This can also be a play against nVidia. When mainstream systems use "good enough" integrated GPUs and get rid of that slot, there is no place for nVidia except in high-end systems.
Below the mini-ITX format with a GPU slot, there are 3 standard form factors that are big enough for a full-featured personal computer that is more powerful than most laptops: nano-ITX (120 mm x 120 mm, for 5" by 5" cases; half the area of mini-ITX), 3.5" (from the size of the 3.5 inch HDDs, approximately the same area with nano-ITX, but rectangular instead of square) and the 4" x 4" NUC format introduced by Intel.
With a nano-ITX or 3.5" board you can make a computer not bigger than 1 liter that can ensure a low noise even for a 65 W power dissipation for the CPU+iGPU and that can have a generous amount of peripheral ports, to cover all needs.
Keeping the low noise condition, one could increase the maximum power-dissipation to 150 W for the CPU+iGPU in a somewhat bigger case, but certainly still smaller than 2.5 liter.
I expect that we will see such mini-PCs with Strix Halo, the only question is whether their price would be low enough to make them worthwhile.
The fabrication cost for Strix Halo must be lower than for a combo of CPU with discrete GPU, but the initial offerings with it attempt to make the customer pay more for the benefit of having a more compact system, which for many people will not be enough motivation to accept a higher price.
One of (many) factors that were holding back this form factor was the gap in iGPU/GPU performance. However with the frankly total collapse of the low end GPU market in the last 3-4 years, there's a much larger opening for iGPUs.
I also think that within the gaming space specifically, a lot of the chatter around the Steam Deck helped reset expectations. Like if everyone else is having fun playing games at 800p low/medium, then you suddenly don't feel so bad playing at maybe 1080p medium on your desktop.
Pay for memory once, and avoid all the copying around between CPU/GPU/NPU for mixed algorithms, and have the workload define the memory distribution.