Cloudera and Hortonworks merge
cnbc.com
cnbc.com
There are tons of them. Spark, Hue, Sentry, Kafka, the list goes on.
The fact that Cloudera, Hortonworks, MapR were all founded and raised $100m+ around the same time was a bit superfluous for the whole market.
Hard to say where this leaves MapR now. They seem to be the odd man out in terms of growth and adoption.
So, what does that mean for us? I am shellshocked. I expect prices to increase (good for us). What else should we expect?
They merged partially because they were both being cannibalized by services revenue they couldn't get rid of. Now they are struggling to move to the cloud (see Atlas now)
I'm not sure yet another hadoop distro with a bunch of 1 off tooling that is supposedly faster is the answer. Why is all that stuff even needed?
Dl4j itself has a decent sized user base. Ranging likely from your phone maker to your bank and retail store.
We have our own software distro too which is why I'm commenting on this. We don't try to boil the ocean with a bunch of tech though.
There's a whole new crop of companies focusing on solving bits of the ML problem well rather than trying to do storage and god knows what else.
My point here about you guys is you're trying to compete in what is largely a commodity market. People don't need all this stuff. Simplicity won here. It's not about better tech.
You guys have the same pitch MapR does and largely the same problem: Better tech is only part of the problem with adoption. You need customers, users, and a clear business model when going to market.
Cloudera and Horton ran one playbook that at least somewhat worked (it got them public) and now they can focus on competing with the cloud vendors, which made the right decision and just made commonly used software easy to use.
Your pitch is still about differentiated tech, not a large install base, a differentiated business model
and something related to people like a good partner ecosystem.
Your pitch here requires tons of services.
People don't know how to use all of this stuff especially on prem.
It takes more than just code to build a business.
I say this as someone who's been doing this since 2013. It's not easy.
If you want to train DNNs on a hundred GPUs today on-premise on TensorFlow, come to us, we can do it. They can't.
the same point. Tech doesn't matter. Simplicity does.
Even in our own product line, we only do a small
subset of this. We don't even require a cluster
to run. We also work with tech that people use.
You are currently competing with horovod
and kubeflow. eg: "competing with free"
You need more than that to survive.
Generally, that comes down to services.
Do you support Kubernetes too?
Reference: https://www.logicalclocks.com/millions-and-millions-of-files...
We have also redesigned the stack around our distributed metadata layer.
We are primarily targeting on-prem right now, but HopsFS would be the fastest DFS in the cloud if you ran it there today.
Define surplus. Why having 3 or 4 companies in 1B plus market is surplus?
Differentiation is of course key, then.
Both these companies are being challenged with a new crop of Databricks and the confluents of the world.
Working as a lead in a small team that deals with a colossal amount of data (human genomics), it was always easier for us to hand roll deployments with terraform/ansible in either baremetal or OpenStack environments.
In the public clouds like AWS we are using the managed services like EMR.
The whole sales pitch I've always got from either hortonworks or Cloudera always seemed more aimed at nontechnical stakeholders than the technical ones. Am I wrong? Have I missed out on some cool stuff?
(those were quickly googled, apologies if i got any dates wrong)
timing is important. cloudera had good, early timing and looked promising because of it. you are right that EMR definitely hurt all the other hadoop vendors, though i think people overestimate how comfortable big enterprises are with moving to public cloud. way more comfortable today; back then everything was a lot less certain. cloudera's name is unfortunate given they never got anything cloud-based successful, maybe they'd be in a better place now if they had.
but some of the key technologies you suggest as better options came 4-6 years later. that's 4-6 years of providing value, gaining traction, and building a committed customer base. 4-6 years is a long time, and even with how slow many enterprise projects run, more than enough to get entrenched, build tooling that makes a bunch of stuff easier, build mindshare, etc.
> Have I missed out on some cool stuff?
stuff that makes companies money doesn't always look cool.
“So what you get is really old versions of everything, plus some shaded jars that will break if your classloader ever changes load order. As far as cloud goes, we’ll give you this half-baked automation (Director) and we’ll also constrain you from using any features of your cloud provider (availability zones, load balancers, Amazon Linux 2). Finally, we require you to use Oracle’s JDK but we won’t distribute it so that you’re indemnified.”
Ansible
Pros:
- More flexible
- Easier to update separate components
- Bring-your-own monitoring & alerting
Cons
- Higher initial time to setup playbooks
- Rolling update configuration is a pain
- Upgrading components require more compatibility testing
- Lack built-in monitoring & alerting
Ambari
Pros:
- Easier to setup from scratch (step-by-step wizard for everything)
- Changing configuration is easy, you change one components, and other components are update accordingly. It also notify you which software in which nodes need to be restarted and allow you to do rolling restart seamlessly.
- Components integrate well with each others, so less time is needed for compatibility testing.
- Back-ported of important patches .
- Built-in monitoring and alerting
Cons:
- Older software version, albeit with back-ported patches
- Cannot upgrade separated components
- Less easy to integrate with your existing monitoring infrastructure.
I also find it easier to train new team members to use an existing Ambari installation than to maintain Ansible playbooks. We are now using Ambari to maintain the more stable parts of our big data stacks (HDFS, YARN, etc), and use Ansible for the part which are still improving rapidly (Kafka, Flink, Presto, etc)
I’m really intrigued to see whether Ambari or Cloudera Manager wins out in the long run. This is unexpected but interesting.
(I worked at Hortonworks 2014-2016, no current affiliation)