HNHacker News
TopNewBestAskShowJobs

sudsmenon

10 karma · joined November 19, 2015

submissionscomments
sudsmenon··on Joining a billion rows 20x faster than Apache Spark
Being able to do things like churn prediction and net promoter score in real time was one of the motivations for creating SnappyData. You get the ability to mutate data (think KPI maintenance in memory without having to jump across products) , and do joins etc. on streams, which makes things a lot simpler
sudsmenon··on Joining a billion rows 20x faster than Apache Spark
SnappyData has a Zeppellin interpreter and the code is open source. So if adding a interpreter for jupyter is something that can be easily added, I am sure someone from the community would find it a interesting project to undertake. Agree that it would be useful
sudsmenon··on SnappyData: OLTP and OLAP Database Built on Apache Spark
The beauty of open source is the we do not have to guess. Take a look at the source. GemFire offers some amazing capabilities and it would be foolish to not use them, but the integration of Spark into the platform is fairly deep and we intend to contribute things back to Spark over time. But marketing is not complaining :)
sudsmenon··on SnappyData: OLTP and OLAP Database Built on Apache Spark
Fundamentally, any enterprise today has to deal with OLTP data, OLAP data (transactional plus other sources), Streaming and finally machine learning. Our premise is that you can choose to use 4 different platforms for each one or move to a unified platform. Spark offers streaming, it offers Spark SQL and there are a bunch of machine learning libraries available on Spark. Also, everyone who has anything to do with data has a connector to Spark so it becomes a good data integration platform. The API is uniform across these. So it forms a good substrate for what we are trying to do. As for bugs, it is a platform that is growing rapidly , going through some growing pains, and over a period of time, it will mature. We believe that the core capabilities will become powerful over time (SparkSQL, Streaming etc.) It is a somewhat opinionated choice but one that we think will pan out over time.
sudsmenon··on SnappyData: OLTP and OLAP Database Built on Apache Spark
1> Most of it is straight up Apache 2.0 license. The only thing that is not Apache licensed is the approximate query processing piece which is closed source right now. We intend to build out a community and that predicates that much of what we do has to be open source

2> Since we have spliced a database into Spark such that Spark and Snappy share the same memory space. Everything that Spark supports will be supported. The data can be accessed as data frames and accessed using the Spark API (and we will be a 100% compatible)

3> Absolutely. In upcoming releases. Both from a data standpoint and from a management and monitoring, and security.

4> I think becoming a full fledged tool is not on the cards. Think of this as a enterprise class real time operational analytics platform. But we intend to integrate and work closely with the leading tool vendors (especially the AQP piece is great for aggregate class visualizations)

sudsmenon··on Apache Geode: Distributed, in-memory database
FWIW, GemFire is all written in Java and has native clients for c++ and .net ( as well as REST clients). The product was built from the ground up in Java and runs highly scaled up , low latency systems across the globe. The place is called Beaverton and while there are many beautiful forested riverbanks, the Pivotal team does not sit next to one.