Ah I was really hoping this was just a library that could be used in spark that implements probabilistic data stores. This still looks interesting, but it looks like this would take a lot more work to spin up.
It takes about as long as it takes to start up a Spark cluster, and you can interact with it entirely through Spark APIs if that's what you're comfortable with. It can also be used in "Split Cluster Mode" so you can use your existing Spark build instead of what's embedded within SnappyData if you prefer.
Right now it is easy for me to spin up a spark cluster on YARN, but hard for me to acquire a number of machines to run arbitrary code.