GPU-Oriented PCIe Expansion Cluster
amfeltec.com
amfeltec.com
This is more or less useless, or if one wishes to be kind, an amazing emulation of building distributed GPU code over the craptastic bandwidths brought to you by AWS, Microsoft, and Google datacenters for now.
Also PCIE Gen 2? WTF? This might not even post with current GPUs.
Baidu just released a library implementing this approach between nodes. NVIDIA's NCCL is a library for passing such data between GPUs connected like this in a single node, and the deep learning framework DSSTNE has tensor collectives for implementing model parallel training and inference.
Do PCIe switches know the address windows of their children? Meaning that when one child does an access to one of its peers (another child), the packet is sent directly to the other child rather than all the way up to the root? I'm not too familiar with how PCIe switches get enumerated, but perhaps it gets configured at that time?
Or does it just behave as a regular bus where every device "listens" for accesses to addresses within their BAR regions (implying that every packet gets broadcast to everyone)? However, I suppose that would be more like a hub rather than a switch...I'll do some more reading!
The device could query the root complex (the OS) to find the bus address of the device with a particular host memory address but I don't know if that is a common idiom.
(How this works with Multi-Root IO Virtualisation I'm not sure and it gives me a headache.)
Ripping out the SandyBridge CPU and replacing it with an IvyBridge CPU magically fixes the problem. I think the difference here is that SandyBridge only supports PCIE Gen 2 while IvyBridge is PCIE Gen 3. One will never know for sure though.
Not a *coiner but I can guessL
With most motherboards even 3 GPUs will require you to have the cards right up next to each other. This is not great for cooling, which is important when you have the GPUs cranking away at full bore.
Also you can only get 4 GPUs at max in to a normal size PC case. With these you can have a single host PC running 16+ GPUs.
NVIDIA will do their worst to shut this down because it's a direct threat to DGX-1 running neural networks that aren't entirely communication-limited (long story), but if they could throw this together, I think they could make a great quick buck before the axe falls.
The 4x splitters go for about $200. But if you want guaranteed compatibility (i.e. a full build), the price (and margin for them) will skyrocket.
Or who are you going to believe? Your #PCIEAlternativeFacts or the above lyin' code libaries? But I'll cut you some slack, back in the day (2013-2014), it took a great deal of data to convince GPU server vendors like Quantum, SuperMicro, and Cirrascale that this was the case. And to this day, they still sell multi-GPU servers with chipsets where one set of GPUs on a PCIE switch cannot directly access the other(s).
There is 1 such bidirectional ring in a tree of PCIE hubs. For bonus points, figure out the 4 bidirectional rings inside the NVLINK topology within a DGX-1.
That said, I like the idea of this expansion cluster. If they were to quickly research and offer an 8796 switch-based variant, they could make top dollar in the 99% segment of the emerging Deep Learning market. Potentially also from hedge funds and big pharma, none of whom wish to pay $5000 for $1000 GPUs with premium trim.
And they would get to do this until NVIDIA gets its knickers in a bunch (once again) about people using GeForce GPUs instead of Tesla GPUs for the sort of workloads they have arbitrarily classified to require said Tesla GPUs and then either cuts off their supply of parts or extorts them in some way to shut them down.
Even with a dedicated 16x PCIe 3 connection there is a latency overhead compared to inter-CPU buses, like HyperTransport or QPI, which is why nVidia and IBM have scaled up the NVlink inter-GPU bus to become a memory speed interconnect.
https://www.ibm.com/blogs/systems/ibm-power8-cpu-and-nvidia-...
Though amfeltec is sadly a pretty unknown company. Probably best known for their "squid" PCIe "split" cards for multiple M.2 drives in one x16 slot. I've used one of them.
If the former, yes you will suffer a massive performance blow even with just one GPU - but if the latter, it's an easy way to upgrade your system.
https://www.microsemi.com/products/drivers-interfaces-and-pc...
The Dolphin systems are tuned better for computation: http://www.dolphinics.com/products/IXS600.html
I expect small PCIe NVMe external storage systems to become quite common in the near future, because enterprise systems need multipath storage for reliability; 8GB/s FC is too slow for SSDs, let alone NVMe, same with SAS bus expanders.
https://events.linuxfoundation.org/sites/events/files/slides...
Would require some modding to use more up-to-date GPUs (there are some mounting studs on the chassis floor that need trimming) but the forced airflow means that passive cards like the original Xeon Phis are also a possibility.
There is a bit of fiddling with drivers but it is possible to have one manufacturer working in compute only mode and other as graphics and compute.
You could build your own super long PCI Extender Cable, using superconductors...