Also, given the insane cost premium apple charges per extra GB of RAM (at least when I was last shopping for a device), do you come out ahead?
Intel Core i9-13900F memory bandwidth: 89.6 GB/s, memory size up to 192 GB
Apple M3 Pro memory bandwidth: 150GB/s, memory size up to 36GB
Apple M3 Max memory bandwidth: 300GB/s, memory size up to 128GB
GeForce RTX 4090 memory bandwidth: 1008 GB/s, memory size 24GB fixed, no more than two cards per PC.
A GPU connected to a PCIe 3.0 x16 electrical uplink would be constrained to ~16GB/s, or ~32GB/s if it were a PCIe 4.0 uplink instead. Although those numbers imply slower bandwidth than CPU inference, that bottleneck would only be when paging in or out (or directly accessing?) layers overflowed to the shared system ram, so they don't really represent much on their own.
> no more than two cards per PC
I've seen quad 4090 builds, e.g. here[0]. What do you mean no more than two cards? Yes, power is definitely an issue with multiple 4090s, though you can limit the max power using `nvidia-smi`, which IME doesn't hurt (mem-bottlenecked) inference.
[0] https://old.reddit.com/r/watercooling/comments/16ed8fu/quad_...