You're right that they likely won't make 2 different dies. Desktop AM5 chips will just get a package with some of the memory controller pins unconnected. The big question is whether they'll also package the full width laptop chip with on-package memory in a package for desktop that's incompatible with AM5.
If they don't, somebody is going to solder that monster laptop chip into an ITX motherboard. People will grumble about a motherboard that can't upgrade either the CPU or the memory, but if the performance is there they'll still buy it.
I don't think putting the memory in the package helps much with practically achievable clock speed or bus widths compared to just soldering the memory nearby. (Consoles and discrete GPUs aren't doing on-package memory despite running GDDR at significantly higher frequencies than LPDDR.) And given that, there's even less reason to expect a messy hybrid configuration with half the memory controller connected to on-package LPDDR5x and half routed to DIMM slots, whether or not it uses the AM5 socket.
What I would suggest (and my suggestions are worth about as much as I'm getting paid to write this):
AMD should create an I/O Die (IOD) with 512 bit DRAM memory width. Then create an AM5 variant with 128 of those connected to the socket and the other 384 tied off. Then create variants with 256 bit and 512 bits connected to on-package memory and no off-package memory pins. Sell those to laptop manufacturers.
A big motivation on laptops is that slow & wide memory uses a lot less power than fast & narrow; on-package also uses a lot less power than driving the signal between chips.
Then in early 2026 introduce AM6 with 4 channel 256 bit memory support. Sell AM5 & AM6 in parallel for a while.
Why not? They already do the equivalent for SRAM. The big cost of a wider memory bus is routing it through the socket and the system board, which you're not doing since only the narrower bus goes through there. The wider bus is solely within the APU package.
You could also take advantage of the additional channels -- have e.g. 8 memory channels within the package and two more routed through the socket, for a total of 10. Now if you have 8GB on the package and 16GB off of it, you have 10GB striped across all 10 channels and another 14GB striped across two.
Continuing to have memory slots also allows you to sell chips with and without on-package DRAM and use the same socket. High bandwidth memory only makes much sense if you have a strong iGPU, since CPUs are rarely memory bandwidth bound. But low end systems and high end systems wouldn't have that, only mid-tier ones would. The low end system has a small iGPU where HBM is both unnecessary and too expensive. The high end system has the same small iGPU, or none at all, because it uses a big discrete GPU.
HBM is also more expensive than ordinary memory, but Windows will sit there eating several GB of RAM while doing nothing, so having some HBM and some DDR should lower costs compared with having the combined total in HBM.
You use HBM or GDDR if you want very high bandwidth, like top-end discrete GPUs would get. But then the memory itself is more expensive and you want the external channels to reduce cost, so the OS bloat can go in the cheap memory and preserve the limited amount of high cost memory for what needs it.
This is notably not what Apple does -- they're just using ordinary LPDDR5 with a wide bus, equivalent to having a lot of memory channels. It gets them several hundred GB/s worth of bandwidth, similar to a midrange discrete GPU. If you were going to do that, you could put most of the channels within the package and still have two of them outside of it.
That sort of configuration would allow some flexibility. The on-package memory might have lower latency (if they're both just ordinary DDR this isn't going to be much difference if any), but if you configured the system to only interleave between the on-package memory channels then the "close" memory could achieve that lower latency. Interleaving the external channels into the same pool would have a small latency hit but increase bandwidth by e.g. 25%. Which could be configured in UEFI based on your expected workload.