51 karma · joined March 5, 2019
One way to make Spark a bit easier to work with is through kubernetes or a tool like databricks that provides it as a service. Kubernetes, in particular, provides you a really nice amount of flexibility and composability when designing systems. One thing that we created to try to fill the gap of having to integrate System X with Spark was HTTP on Spark. This makes it easy to integrate Spark with other tools in a microservice architecture. When you couple this with containers, you can do a lot very quickly.
For datatypes I would look into different Spark connectors, these days there one for almost every database/streaming service/ cloud store under the sun.
This being said, Spark is a large piece of software that uses many different programming concepts which can be daunting. Our goal is to try to listen to feedback like this so we can try to make the Spark ecosystem a bit easier to use for everyone.
For the likelihood functions comment, I would totally agree. Autograd libraries are easier to build custom likelihood models in, which is why we created CNTK on Spark, and databricks created Tensorflow on Spark. These give you the flexibility of modern deep learning stacks with the elasticity of spark
But in the end Spark is a single tool in a collection of tools and might not be right for your project, but it's been good for a lot of our work here at MSFT :)!
https://github.com/Azure/mmlspark
Thank you!