On the multi-process..yes, that's easy. I've removed it for the current main-line branch given the lack of demand. IfDef'ing the code made it much easier to proceed with getting it ready for alpha. The FIFO mechanisms are well tested using SHM and the forking code will be added back in soon. Another thing I commented out to get it working on multiple platforms is the NUMA placement code, now that I think hwloc will work on all platforms I'll get it added back in. Helps out on cross-socket communication quite a bit, as well as placing buffers closest to PCIe root for data transfer to accelerators.
In reality the data movement is no worse than any OpenMP or other parallel program. In as many places as we can, the data is left in place vs. pushed. Between nodes, it gets more fun...however it's still a rather well understood problem. Thanks again for the interest! I'll see if I can do a ShowHN before CPPNow 2017 for the beta release.