As an extreme example, when and where it was/is still viable, cryptocurrency mining operations use external PCIe cases with lots of slots, attached to only a few or even just one lane per case. For their purpose this was still viable, since mining uses very little I/O from/to the GPU, essentially only a small block of setup data into the GPU ("try hashing this block, vary those bytes within that range, report when result matches this pattern", a few kB) and a few bytes of results out ("found a match, varied bytes were as follows").
You are sending the exact same data to both GPUs. If the the PCIe switch and GPU drivers support multicast, then you would actually get the full 16x bandwidth to both.
Which is actually a huge advantage over the most common multi-gpu setup, which is forced to statically split the 16 PCIe lanes into two independent 8x slots. Even if you were playing a game that only used one GPU, simply having the second gpu plugged in meant your total bandwidth to each gpu was limited to 8x speed.
I really wish you could find x86 motherboards with PCIe switch chips... but switch chips are really expensive so nobody makes them.
I would think this would lead to bottlenecks for instance if you were trying to write to two different 16x PCIe SSD RAID cards simultaneously.
Finally Woz’s design for the Apple II color graphics was based on unified memory.
Dramatic shifts in technology sometimes obviates all our past learning and experience.
PCIe is a tree-shaped network (root is the root hub in the CPU, inner nodes are switches, leafs are devices) with point-to-point links (called lanes) that transmit packets. Bandwidth is determined by the number of lanes that a device's path up to the root hub has at the narrowest node.
Switches are far more complex. The protocol specifies addressing, bandwidth reservations, flow control, priorities and message ordering requirements, which the switches have to handle.
Not a switch thing per se, but also relevant for "single transistors won't do": On the physical layer, each port/link is also relatively complicated to initialize and operate. The transmission encoding, due to the relatively high speeds involved, requires a learning phase where both ends measure the link characteristics and adjust their (relatively complex) signal processing accordingly. The physical layer also handles different link speeds for down-/upward compatibility.
This may explains how they achieve the graphics blitting in the Vision Pro.
Not sure how this affects the use case of multiple video capture cards writing to fast internal storage that I see mentioned here and there. Anyone close to that industry around to comment?
Edit: i think I’ve seen mentioned that the PSU in this thing is “only” 1400W. That only covers one modern nvidia gpu right? :)
The graphics and processing speed of the Vision Pro makes me wonder if the GPUs are truly needed.
Most of the cards for the M2 Mac Pro are used to access the analog world. It seems that there is power inside the die and if you need an analog signal use a PCI-e port.
For GPU compute, I'm sure it can't beat 15 kW of nvidia dedicated cards.
But then you don't want 15 kW of cooling on your desk.
I'm mostly curious as to what this work load consists of?
but when I read the number 8k encode and decode streams 15 kW represents
How much of the power is just getting the data in and out of the card to the processor. Apple Silicon doing the same work on die is going to use less power. I don't doubt that there are faster processing units.