Let's say someone want to use python first to create a few tensors with numpy, and then use another python package networkx to create a graph from the tensors just created and next to do some graph analysis with networkx.
It is easy to do it on a single machine in a single python runtime because data structures like tensors can be easily shared within a single python runtime.
It will become a bit harder to share the data across processes/Python runtime efficiently without expensive IO or serialization (doable with the help of things like plasma in Apache Arrow).
However, if we want to deal with "big data" that cannot be handled on a single machine, it will become very hard to share data without IO/serialization.
Vineyard can be used for such cases. It makes sharing big distributed data structures easy for different runtime.
There will be some added benefits with vineyard:
1. vineyard can handle the the IO, sharding/partitioning, failover and data migrations for the applications.
2. vineyard supports many out-of-the-box highly efficient data formats (such as Apache Arrow) and high-level abstractions which ease the difficulties of developing big data applications.
3. With stream data support, there is a possibility of enabling cross-process optimizations.
4. Fits K8s very well, where common data structures can live in a separated pod/container with its own resource limit.