Optical switching is about training at scale and replaces classical physical networking systems (people will say things like "electrical", "copper", etc.). Running models at edge devices is almost by definition without networking.
If I am getting this (still reading the paper... on a phone :/)... Stupid fast latency due to no processing overhead but slow switching speeds due to how it physically switches, this is basically good for any type of reconfigurable cluster of systems that need low latency without a fixed topology, not strictly ML specific or have I misunderstood that?