Very excited to see all these open-source projects take off in the internal tooling space. I regret how much time I spent building custom DIY tooling at previous jobs!
32 karma · joined March 1, 2019
Very excited to see all these open-source projects take off in the internal tooling space. I regret how much time I spent building custom DIY tooling at previous jobs!
We're showing new metrics (CPU & Memory) that you couldn't get in the Spark UI (you had to use something like Ganglia, and then jump back and forth between Ganglia and the Spark UI by comparing timestamps).
We alos hope to display this new information in a way that will make it more easy for most users to get insights. Make problems more obvious, more quickly.
Let us know your feedback!
I’m JY, co-founder of Data Mechanics (YC S19, https://www.datamechanics.co). But this post is NOT about startup core product (a managed Spark platform, deployed on a k8s cluster in our customers cloud account, see our HN Launch https://news.ycombinator.com/item?id=23142831).
This post is about a free and partly open-source monitoring tool for Apache Spark that we just released.
It works by installing an open-source Spark agent (https://github.com/datamechanics/delight) on your Spark infrastructure — whatever it is: commercial or open-source, on Kubernetes or on YARN, in the cloud or on-premise.
This agent streams event metrics from Spark (metadata about your Spark applications) to our backend, which then serves a dashboard listing the Spark applications, and giving you access to the Spark UI (Spark History Server) for each of them. The blog post has a lot more details about the architecture and security of it.
This release is just a first milestone for us, in our next release in January we will add new screens to gradually replace the Spark UI with a new monitoring view, you can see a glimpse of it at the GIF at the bottom of this page (https://www.datamechanics.co/delight)
We’d love your feedback about it — is it easy to install and use? What would you like to see in following releases? Thanks so much! JY
Sources: - end of presentation https://www.slideshare.net/databricks/reliable-performance-a... - https://issues.apache.org/jira/browse/SPARK-25299
Sorry to hear about the layoffs. I'd like to follow-up with you to get your feedback on specific roadmap items we have in mind. Would you email us at founders@datamechanics.co to schedule a call, or at least keep in touch for when we have an interesting feature/mockup to show you? Thanks and good luck as well!
Good luck with your venture :)
That's why we're working on new monitoring solution (think Spark UI + Node metrics) to give Spark developers the much needed high-level feedback on the stability and performance of their apps. We'd like to make this work on top of other data platforms (at least the monitoring part, the automated tuning would be much harder).
Case studies: Thanks, we're working on them. Check our Spark Summit 2019 talk (How to automate performance tuning for Apache Spark) for the analysis of the impact at one of our customers.
In the meantime you can book a time with one of our data engineers through the website to get a live demo: https://www.datamechanics.co
Optimization/Monitoring: This topic is very important to us, thanks for bringing it up. Indeed we automatically tune configurations, but developers still need to understand the performance of their app to write better code. We're working on a Spark UI + Ganglia improvement (well, replacement really), which we could potentially open source.
Would you mind emailing me (jy@datamechanics.co) or even scheduling a call with me (https://calendly.com/b/datamechanics/avk7bhxq) so I show you what we have in mind and get your feedback? Anyone else interested is welcome to do the same.
This being said, Databricks is a great end-to-end data science platform, with notable features we lack like collaborative hosted notebooks. A lot of people don’t want/need the full proprietary feature set of Databricks though. They choose to build on EMR, Dataproc, and other platforms instead. We hope they’ll try Data Mechanics now :)
Dynamic allocation is only enabled on our Spark 3.0 image (from the 3.0-preview branch, since the official 3.0 isn't released yet). It works by tracking which executors are storing active shuffle files. These executors will not be removed when downscaling. More info here: https://issues.apache.org/jira/browse/SPARK-27963
It's not perfect, but there are more improvements for dynamic allocation being worked on (remote shuffle service for Kubernetes).
Regarding costs. By autoscaling the cluster size and minimising our service footprint, the fixed cost for using our platform is around $100/month, which is negligible compared to the cost of most big data projects. We have some ideas on how to drive this fixed cost to zero, and offer a free hosted version of our platform too. It's in the roadmap!
If you're curious about our ML approach, we gave a tech talk about it at last year's Spark Summit: https://databricks.com/session_eu19/how-to-automate-performa...