So NVLink/NVSwitch pools multiple GPU resources on a single (very expensive) system. A cheaper alternative to that is "offloading", which is a technique that splits the inference process into smaller steps, so it can run on systems with much less resources available... and Petals is a 10x faster alternative to that.
Did I get that right?
This AI stuff is moving very fast and it's hard to keep up, but it's all fascinating.