Anyscale, from the creators of the Ray distributed computing project, launches
techcrunch.com
techcrunch.com
https://news.ycombinator.com/item?id=15481169
It would be compelling if Ray were to provide Horovod support versus the authors re-applying some of same research in their own thing. Ray programming API + code distribution + Horovod performance primitives is what the community probably wants.
It seems Spark is rolling on inertia more than being a modular system for low-latency cluster computing.
It's good someone is building a company around it. I could see them building services on top of it and build a SAAS like databricks did with spark.
I'll be curious to see how ray matures.
I’m worried about Ray as a SAAS Co because so far it looks to me like they’re riding reinforcement learning hype. They’d need to really penetrate the users of Horovod and Tensorflow Distributed to get beyond a beach head. And what if TPUs and Cerebras become more common? Because then the maker for multi-machine workloads becomes smaller (definitely not zero though).
One interesting thing that could happen is the hardware gets better, and then these distributed schedulers might not be able to keep up with all the different options on the market.
There is also the tension of the hardware vendors wanting to give away things that only run on their chips vs the software makers who want things to run on every chip. It seems like there will be a lot of competition among the various infra players in the next few years now that nvidia is starting to have real competition now (even if it's not big yet)
Ray shows expertise in multi-machine that's lacking in stuff like Jax, Tensorflow, and PyTorch. Horovod nailed down a lot of the performance issues for SGD in particular, but is missing the sort of rapid deployment / distribution stuff in Ray. If only they could all work together ...
And does "code distribution" just mean to release as open source?
I wish the Ray peeps would consider just trying to merge some with the Spark RDD API. Reynold (at Databricks) is kinda hard to deal with, but so far to me it looks like the aim of having Ray team build things from the ground up has simply re-validated a lot of the systems work (but not all) that’s already in Spark.
berkeley rise lab is such a powerhouse - spark, mesos, etc. just in the past few years. it really puts some of these big tech companies to shame.