The only way to do this is to have one chip (It's always the GPU. The GPU needs more memory bandwidth) connected directly to the dram, and the second chip (CPU) has to send memory requests to the second chip.
Though, this console dates to a time when CPUs didn't typically have dram controllers onboard. PCs usually relied on a northbridge chip to have the dram controllers, along with the routing to all peripherals (PCI/AGP) and present a nice tidy Front-side-bus that the CPU understands. In the case of the xbox, the GPU is acting as a combined Northbridge/GPU (a design that was common at the time in low-cost desktops and laptops)
Unified memory has a large number of advantages for consoles. It lowers cost. It gets rid of copying delays between GPU and CPU memory and it allows the game developer to dynamically allocate memory to the GPU or CPU depending on their needs.
As I recall, the earliest Intel integrated graphics added almost nothing to the production cost of the chipset. The die needed a certain perimeter to support all the IO connections, and that left quite a bit of unused silicon in the middle. Putting GPU logic in that space was almost free (minus R&D), and only slightly increased the total pin count. Intel got to capture slightly more revenue per PC and deny a lot of revenue to competing chip companies making discrete GPUs.
The situation today is very different, with GPU and CPU on the same die and the GPU blocks taking up far more space than the CPU cores. The integrated GPU is an important part of the chip cost, and that means desktop processors often have worse (smaller) iGPUs than laptop chips.
It's a lot like what you see today with AMD's chiplets at 7nm connected to a central I/O die on 14nm. Just the classic systems were integrated at the board level instead of a special interposer.
EDIT: I wonder if the switch to serdes links instead of parallel buses (infinity fabric instead of hypertransport) is a large part of what made this idea useful again, reducing the number of off chip signals again for the CPU dies. I wonder if we'll therefore see a replacement for Intel's QPI if they switch to chiplets too.
Wikichip [1] says there are two versions of infinity fabric (which is a super-set of HyperTransport), one optimised for on-package communication that is 32 bits wide, and one optimised for inter-socket communication that is 16bits wide.
I'm not a hardware person, just a software person who dabbles in hardware, so I don't know if the term "SerDes" strongly implies serial and AMD have misused it here, or if SerDes is generic enough to apply to any SERialisation/DESerialisation.
But still, I think you are on the right track. The investment and development into low power, high-speed, and low-latency on-package links is probably what has enabled chiplets to become relevant now.
Because multi-chip modules have always existed.
The primary reason for this design is that at current frequencies it is essentially impossible to manufacture the physical parallel interface with equal-enough wire lengths. Interestingly, memory interfaces (like DDR4) use opposite approach: the interface is still mostly parallel, but memory controller measures the delays and mismatches of the physical wires and compensates for that in its timing.
Really? That's crazy! I thought DDR was serial connections. In fact, I thought parallel connections had mostly gone the way of the dodo. Serial is just so much less complicated.
I think Intel's trying to ensure they have the advanced packaging/interposer/bridge tech to handle wide parallel connections between chiplets. If it works out, they might even end up moving in the opposite direction—toward wider interconnects rather than narrower.
One fun thing was optimizing for the SPUs meant that your code was really cache coherent and usually saw significant gains on all platforms. Of course most people at that time wrote for PC+360 first and entered a world of pain when PS3 came along.
If there's one lesson to take away it's always build for your most constrained platforms first. There's still a few funky architectures out there(RPi I'm looking at you[1]) so it's always worth understanding where the hardware constraints in your system can come back to cause havok.
[1] https://www.raspberrypi.org/documentation/configuration/conf...
What do you mean by this?
Especially if the GPU already had custom silicon for it.
Whether it would be good engineering (cost, time to market, risks) is of course another issue.
It gets harder and harder to do such external muxing as the ram gets more and more complex. With multiple banks, row open delays, bursts and more complex signalling (fast and faster ual data rate at lower and lower voltages) it's near impossible to control modern DRAM without a proper controller.
And that controller has to live inside a single chip. It would be insanity to try and have two different dram controllers multiplexing the same DRAM chips.
I don't think thats true, in embedded architectures its not uncommon to have dual-port RAM.
I'm not aware of any designs which have large amounts of dual-port RAM as main system memory.
Ultimately designing a circuit board layout is an optimization problem. You usually have some constraints, like how close chips can be before they start interfering with each other magnetically, where the I/O will be, and where you need holes to mount the board. Then you either try to be a pathing optimizer yourself or you run a program that will layout your board for you.
I'm not sure about the XBox, but game consoles sometimes have faster Memory->GPU pipelines than normal PCs to speed up render times, which might be why the GPU is the most central component.
The only unusual thing here is that the GPU and Northbridge are the same chip.
Over time CPUs have integrated all those features on-die, resulting in today's SoC-like processors where the "chipset" is merely an I/O expander connected over a PCIe-like link.
Back in the years we're talking about, that would be AGP (https://en.m.wikipedia.org/wiki/Accelerated_Graphics_Port).