Cloudera taken private for $5.3B, acquires Datacoral and Cazena
blog.cloudera.com
blog.cloudera.com
Cloudera actually started in the cloud, but quickly moved to selling mostly on-premise software. Back then, cloud was a lot less cost competitive (it still is not really cost competitive) https://a16z.com/2021/05/27/cost-of-cloud-paradox-market-cap... The on-premise business had much higher margins, which kind of "took the oxygen out of the room" for the cloud side of the business. In retrospect, neglecting the cloud market was a big mistake, of course.
Cloudera management knew even back then that Hadoop was eventually going to be superseded by a newer technology. We tried to build that technology, in the form of Impala, an optimized SQL query engine. Unfortunately, it never really took off in the marketplace, for several reasons. For one, its dialect of SQL was not compatible with Hive, the traditional SQL frontend for Hadoop. For another, early versions of Impala were somewhat unstable, and didn't have the fault-tolerance of traditional Hadoop tools.
From my perspective, the main issue that cloud solves is that managing the server side stuff is really hard, and requires experts. Cloud also aligns expectations: it's really clear to both customer and vendor that a cloud contract is not a stepping stone towards self-managing. Unfortunately, Cloudera just didn't see the shift towards cloud coming, and they're now paying the price.
They both offered a Hadoop distribution but had different strengths e.g. Hortonworks had fine grained access control, Cloudera had a better SQL product with Impala.
Then AWS came along and built their own which was significantly cheaper and more flexible as you could easily scale your cluster up/down. And so companies moved to it when they over time began to move to the cloud.
The Hortonworks/Cloudera response to this threat was to put away their differences and merge together.
Over time Big Data has evolved from being Hadoop centric to being much more ML/AI focused i.e. not just manipulating and querying the data but doing something interesting with it. And AWS, Azure, GCP have really jumped in with a whole suite of products that are tightly integrated with the rest of their cloud offerings. And it's a large part of what differentiates their offerings so they compete very hard.
So Cloudera has no choice but to do things that cloud providers won't or can't do: (1) focus on non-cloud or multi-cloud and (2) offer a much more integrated and cohesive solution.
But having spent 10+ years in this space and deployed many Hadoop clusters I can tell you that Cloudera is going to struggle. Companies that I never thought would move to the cloud e.g. banks are figuring out the security and regulatory challenges and eagerly moving across. And so it's going to be a Cloudera versus Amazon/Google/Microsoft which is an impossible fight.
A good product is more valuable than a partnership.
IIRC MS/Azure is an early Databricks investor and their sales folks were heavily incentivised to sell it. They also pushed Snowflake in the early days until they had a competing product and their relationship status was upgraded to 'it's complicated'.
I’d like to learn more. What doesn’t work? What can vendors do to make it easier? My understanding is that lift and shift doesn’t mean one and done, no grueling manual testing required.
Once you set the lift and shift as long as the source schema doesn’t change, you could run it as often as you’ll like as you deploy, test, fix?
As more people moved to the cloud, Hadoop-style storage was extremely expensive (naively moving your Hadoop cluster to 3x replication on EBS volumes would result in a nasty case of sticker shock) so the data would move to S3 / ADLS / GCP. And now you've lost your data gravity.
Post-merger Cloudera focused less on on-premises clusters and tried to offer those same diverse workloads as a multi-cloud SaaS, with more focus on elasticity. This is hard because (a) there's a massive amount of surface area if you want enterprise customers to bring their own accounts, run all these managed open-source services in those accounts, and be multi-cloud, and (b) you're just competing more directly with the cloud vendors, on their turf as both a customer, partner and competitor.
You had to worry about the size of files since the NameNode would be overloaded. Being a Java app running on the older JVMs it would do a full GC under heavy load and cause failovers. And it was impossible to get data in/out from outside the cluster using third party tools.
I remember many companies seeing S3 and just being in shock that it was so cheap, limitless and that someone else was going to manage it all.
I think there are still a couple use-cases where HDFS dominates S3 (I think some HBase workloads?). But yeah, I scaled up and maintained a 2000+ Hadoop cluster for years, and I would never choose it over object storage if given any plausible alternative.
This is the wider issue with small files. On HDFS each file uses up some namenode memory, but if there are jobs that need to touch 100k+ files (which I have seen plenty of), that puts a real strain on the Namenode too.
I have no experience with S3 to know how it would behave in terms of metadata queries for lots of small objects.
Kerberos (quite popular on big enterprise clusters) is really what makes it hard to get data in / out IMO. I see generic Hadoop connectors in A LOT of third party tools.
[1] https://engineering.linkedin.com/blog/2021/the-exabyte-club-...
The meta data is not the only problem with small files. Massive parallel jobs that need to read tiny files will always be slower than if the files were larger. The overhead of getting the metadata for the file, setting up a connection to do the read is quite large to read only a few 100kb or a few MB.
The other issue with the HDFS namenode, is that it has a single read/write lock to protect all the in memory data. Breaking that lock into a more fine grained set of locks would be a big win, but quite tricky at this stage.
For e.g. even this site: https://blog.ycombinator.com/
Are you going to spend months upfront carefully modelling the data in order to ingest it making sure to handle schema and DQ issues etc. All to support one use case who only needs a handful of fields.
No. Which is why data lakes exist. Because it's cost effective. You simply dump the data and ask the Engineer or Data Scientist building the use case to do the heavy lifting rather than a centralised data team.
It has always revolved around the concept of a data lake, with data stored as objects, a series of data engineering pipelines moving data around and a query engine on top. And in almost every enterprise company this is the high level architecture you see today.
And this model only continues to grow in popularity as the use of siloed SaaS products drives data sprawl and the need for tools like Spark, Fivetran etc to move it all back to a centralised data lake for analysis.
A data lake is one way to deal with it. It's convenient but not always the best.
* Tuning for HDFS and HBase is hard. Too many files and GC pause will kill you. Hive Metastore is also a bottleneck.
* It's a nightmare to figure out how to configure some third-party software that's outside the bundle.
* It's even harder, if you want to use a more up-to-date version of a software included in the bundle. We still haven't figured out how to use Spark 3 with our old HDP clusters.
* Sometime the included software is a buggy mess. For example, HDP 3 provides HBase 2.0, which is absolutely unusable. Luckily, we were able to migrate to HBase 2.2 before putting too much data into it.
All in all, it still takes a lot of ops work to own a Cloudera's Hadoop clusters. I can't wait until the time when we can move to a better solution.
It is packaged open source software, from multiple projects (some of which are financed by Cloudera) all of which release at different paces.
And distributed systems are really hard to get right. There are lots of knobs to tune, all of which are needed in some circumstances, but.. there are lots of things to go wrong.
> It appears to have costed them the market, yes?
Cloudera is probably the market leader for on-prem Hadoop clusters.
This is a big market (even now - lots of defence clients who have issues with cloud).
But generally this market is getting eaten by Cloud SAAS products.
There are as probably as many services running inside a Hadoop clusters as a Micro Service meshes, except:
* The services can be huge: HDFS Namenode can takes hundreds GB of RAM, hours to start up in a multi PB cluster. Updating configuration requires restarting, so it is a huge pain if you care about zero downtime.
* The services are often bound to host and can not be easily migrated to other hosts.
* The communication interfaces between services are not well-defined: Hive Metastore, for example, didn't have a formal protocol documentation for a long time, yet everything depends on it to store metadata (I'm not even sure if they have the protocol now). As a result, services are often tightly coupled: you need everything to be compiled with the same libraries version, or stuffs might not work correctly. Furthermore, due to dynamic code loading, issues might not surface until days later - in the middle of the night, making life miserable for everyone.
Going private could give them the opportunity to make significant RnD investments away from the quarterly demands of a public company. On the other hand, the private equity backers could enact a bunch of cost-cutting, gut the company, cause its top employees to leave and load the company up with debt. It will be interesting to see how this shakes out.
PE usually is about discipline. I.e. cash based discipline. So PE change the debt/equity ratio from 30/70 to 70/30, thus enforcing the company to be much more cash efficient. I.e. take LESS risks.
To sum up, I do not envision more R & D.
PE has some scars, but they don’t typically take over healthy, well run companies. They take over mismanaged companies, and refocus on the pursuit of profit that you seem to take issue with, ultimately saving the company from bankruptcy.
There are plenty of PE deals of companies no where near insolvency, where the companies just had some more juice that could be squeezed (short/medium/long term depending, but usually short/medium term). If it’s a problem if THAT ends in bankruptcy or not depends on if they are able to get sufficient positive return/exit before they get hit by any losses - seen that many times.
If you can save a company from bankruptcy, that’s wealth that would otherwise have been destroyed.
It is purchased to make return, not to save them from bankruptcy. If it is a side effect, that’s fine, but that is not why it’s done or what the goal is no?
On Datacoral, it seems like a very good product, simplifying the overhead of data pipeline development but still staying modern. Datacoral has many cool features, such as treating ETLs as code with an option to develop via GUI. The industry is moving that way, has been for some time.
Good for Datacoral to get acquired and get exposure to so many new customers. But Cloudera has their work cut out for them to stay relevant in a market that is extremely competitive.
I'm too lazy to find the best reference, but this is one re: Skype - https://www.businessinsider.com/skype-scandal-silver-lake-20...
they raised a bunch, built a market... then Amazon and Google came for them. Impala was great - I was a big advocate, but then I worked on BQ for a bit, and now Omni.... Cloudera cannot compete.
The only stumble I can identify is that they didn't support Spark. They backed the Hadoop side too hard and left Spark open to Databricks. They should have signed those guys before any other investors got in (told Matie and co "go get offers, we will add 30%).
The "only" stumble I can identify is that they're selling a last-generation solution and most companies see Hadoop as tech debt nowadays. Which is to say, it's a systemic issue with their entire product, not a tiny mistake. This is like Mesos vs Kubernetes. One got squashed.
Over time it became increasingly popular to use cloud storage instead of running HDFS. This really destroyed Cloudera's moat, because there was no operational overhead to putting your data in S3 or GCS. You just needed to run some stateless compute, and if you fucked up it didn't matter. Nowadays your "data lake" is a bunch of files in commodity storage someone else runs.
There is probably a business-school case study there for branding yourself with the problem area rather than a single solution, esp. in fast-moving domains such as tech.
Also interesting is that these are market areas that AWS is starting to expand into.
Looking ahead, on the one side, we're seeing a privately hosted (on-prem + cloud) revival due to multicloud / lockin distrust + extreme cloud markup for AI/ML (GPUs etc) on the other, and for cloud, insane 30-50% YoY growth rates even at $B scales. On the other, I'm not sure how that lines up w/ Cloudera's current products and the investment team's R&D/M&A appetite
It also makes me wonder what is the value of being in n-th place. We're reviewing which new connectors to prioritize for our own viz tool dev, and cloud-native offerings like snowflake/databricks/redshift are generally ahead of on-prem ones, and we're not the only ones. So what's the value of current growth, and of paying premiums to leap frog (ML, GPUs, etc.), esp. when you have the kind of footprint cloudera earned?
Really?
(Maybe that's just a stupid rather than a serious question)
I believe their offering significantly improved during period.
(I was a Professional Services employee for a few years)
This is all definitely anecdata but, IMO, being backed by a PE firm forced us to focus on revenue (really EBITDA) alongside product growth. A mechanism that forced us to focus on the impacts of each product decision we made. I think this ultimately helped us keep a steady pulse on the market w/o chasing every shiny new trend that popped up.
The big bet we’re making at Splitgraph [0] is that the next wave of data engineering will take a more decentralized, “data mesh” type approach to enterprise architecture. “Data gravity” really does exist - it’s expensive to move, in terms of both cost and operational complexity. And with increasing specialization of analytical databases, a single source of truth will become unrealistic. So instead of bringing the data to the query, why not bring the query to the data? All we need for that is a set of read only credentials. And yes, it should also be easy to warehouse your data, but it doesn’t need to be the default.
Cloudera mentions they bought DataCoral to help with data integration and connectors. They’ve correctly identified the problem - data sprawl and fragmentation will inevitably grow - but I’m not sure they have the right solution.
Data integration is important, but it’s a moving target, which is why it calls for a collaborative open source solution. This is why so many new startups, like AirByte most recently, are coalescing around the Singer taps that Stitch left behind after its acquisition by Talend.
We also support using Singer taps to ingest data into versioned Splitgraph images [1], so we’re excited to see more collaboration on maintenance of taps. For us it’s a useful feature, but it should be just that — a feature. Is there really a need to replicate all of your data before you can even query it? Or would you rather experiment by directly querying its source?
[0] https://www.splitgraph.com
[1] unreleased and undocumented atm, but it does work. We’re hiring, especially on the frontend if you want to help build the web UI. See profile.
I'm a committer for Apache HBase, Apache Hive myself and I've been in this space for 13 years now. Yes, the hype is over but there are tons of companies using this stuff in production and tons of companies choosing it for new projects.
We're trying to tackle the biggest pain points our customers had: Lack of flexibility (i.e. locked into specific versions for ages), CDH/HDP not built on Infrastructure as Code principles, Security is hard to do, ...
The three-sentence (buzzword heavy) technical summary: Our distro uses the Kubernetes control plane but we've developed a custom Kubelet that runs software using systemd as its backend as well as a bunch of operators (all in Rust...). This allows us to leverage the best of both worlds and also allows hybrid scenarios (part in containers, part on "bare metal"). We've replaced Ranger/Sentry with OpenPolicyAgent.
If anyone's interested feel free to reach out to me (E-Mail in profile / https://www.stackable.de/en ).
Ultimate software is a recent example of this: https://www.forbes.com/sites/antoinegara/2019/03/01/an-insid...
How does this square with the valuations of companies like Amazon and Tesla whose share price performance is certainly not based on the next quarter or two worth of earnings.
Or pick any other top companyies' stocks which are collectively driving the market, presumably because they are expected to deliver years of more growth and cash flow.
This changed my impression to going-private moves, although in Dell's case the buyer is more like a founder with a help from the investor.
Being private allows for less accountability to people who might not be acquainted with the business, want short-term profits, or do not share the owners long-term vision for the company. All of these allow for more freedom in movements and limit the damages to those who understand the risks. On the other side, the company has more limited pool of liquidity sources and each of the investors might have stronger voice in shareholder meetings.
As a result, the public investor might not share the profits of the business, but it might not incur losses because of it as well. Therefore, private equity is just another business organization model which has its pros and cons.
But the latest boom in PE is a function of: higher than usual levels of credulity around private vs public performance (but PE funds have lower volatility???), and the wave of free money coming out of the Fed. In 2010, lots of PE funds got bailed out after making unbelievably bad bets with no economic rationale (the worst being Blackstone Real Estate, they should have gone bust, they now manage $250bn in RE).
Tech companies are hot right now because of Vista Equity's numbers. Tech companies appear nominally attractive in cash flow terms because so many tech companies pay staff non-cash. So you can acquire at an unreasonable multiple, load up with debt, and then leave employees holding your bags if it goes wrong...self-evidently though, no actual value is being created here beyond playing the capital cycle.
Most of the growth in PE is correlated to the growth in investors (both in the funds, and financing) who don't understand what they are doing. If you look at Europe, private equity activity has exploded higher with money-printing from the ECB and the growth in direct lending/leveraged loan markets (these have gone from 20bn to 150bn in five years or so...where it was in 2007, and lots of very unsophisticated private debt funds with poor incentives). Shadow banking all over again, no-one knows where the money is coming from, no-one knows where it is going.
When I see the name of KKR mentioned...
Hadoop on Google Trends peeked at 2015: https://trends.google.com/trends/explore?date=all&geo=US&q=h...
"Hadoop is dead" seems to be a popular topic in past few years. https://www.google.com/search?q=hadoop+is+dead
For example "computers" peaked pre-2004: https://trends.google.com/trends/explore?date=all&geo=US&q=c...
Javascript: https://trends.google.com/trends/explore?date=all&geo=US&q=j...
Machine learning: https://trends.google.com/trends/explore?date=all&geo=US&q=m...
YARN, Hive, HDFS, MapReduce have been replaced by Kubernetes, Snowflake, S3, Spark.
Kubernetes is overused right now, it has its place but it's not nearly universally the right tool for the job.
Snowflake will eventually fall to something else due it's poor economics.
S3 and Spark though I anticipate to be around for a good few years and if they lose out it will be to imitators or evolutionary equivalents.
Can you elaborate on this point? What’s wrong with their model?
Entry level DO compute instance (so boring, I know) is $5 per month.
There is a large gulf of pricing ranges that can undercut them in the coming years. It doesn't matter now because a lot of analysis projects are disconnected from market forces due to their projects mostly being darling child green field projects or new revenue streams. The moment the next AI winter comes along, a lot of projects by then will start to look like legacy code and the original thought process turns into worrying about cost centers.
And my understanding is they jacked up the prices to boost the earnings to boost the stock price leading up to IPO. They can be disrupted much faster than they will decide to let off that pedal.
Almost every Big Data tools only works with Hive Metastore (and Amazon Glue Catalog, but the compatibilities is not 100%)