Looks very interesting guys, thank you a lot to release that.
I was working on something similar currently.
Any data pipeline complex enough finish soon or later multi-lingual (Python, R, C++ + MPI) and multi-runtime (JVM, native, python) and it becomes quickly impossible to execute everything in one single process space without problems.
You are right in your design, shared memory node-to-node data distribution is the answer to that to avoid the classical / inefficient data dump-load-dump pattern that we find usually in most heterogeneous pipeline.