MapR may shut down as investor pulls out after ‘extremely poor results’
siliconangle.com
siliconangle.com
These companies are failing because they can't establish a use case of running their product on cloud. It is ironic because when AWS started EMR, MapR was one of the distributions they offered via EMR[1]. Over period of time, they weened off MapR and stopped offering them as deployment option. So it is another case of AWS cannibalizing its own third party ecosystem.
Once this was gone, MapR got reduced just another marketplace partner [2] and their additional licensing cost didn't make sense when native EMR was sufficient for most use case.
[1] https://aws.amazon.com/emr/mapr/pricing/ [2] https://mapr.com/partners/partner/amazon-elastic-mapreduce-a...
Whereas AWS offering is ideally suited for transient clusters. It relies on other solutions like Athena, Redshift Spectrum to cater for ad-hoc use cases such as querying and reporting. In this regard, EMR has much better support for programmability and elastic resource provisioning which is really important for transient clusters.
Over the last 5 years the real money making enterprises solved their "it should look like a big file system" problem so "the rest of our stuff works" issue stopped being an issue ( in a process Isilon got bought by the EMC, EMC got bought by Dell ) either by buying those specialized solutions, moving to object store or building home grown systems that worked with their specific applications and big data people went with newer, more shiny big data solutions, leaving MapR with no market.
I also think cloud based data lakes and tooling around them seriously decreased the appeal of Hadoop.
[0] If I can fit your big data into my memory, it is not "big data".
And that adage of big data can't fit in memory is nonsense these days. We run clusters with hundreds of terabytes of RAM which is very much big data. It's pretty easy and affordable with the cloud.
When MapR was selling them the M5 and M7 were deployed in a data center on servers.
But the MapR sales people were impossible to deal with. They wanted to talk about Big Data. And all the wonderful things that we could do with it.
Me: I have lots of files. Think billions. Lets talk about what and how your stuff works to solve this problem because I'm hearing some people successfully used it on a million files scale.
Them: That's great. Let me tell you about big data that you can do using it.
Me: That's ok. I really want to solve my files problem.
Them: Big data is the future! You can change your application to do big data.
Needless to say it went nowhere fast.
Fundamentally, file system clusters are a very difficult business to get into and make enough money to justify being in that business. On a low end there's open source stuff that kind of sort of works. Selling consulting on a top of it is at best a ramen-profitable business. No one is buying a Ceph-consulting for a million dollars if they run a business that needs that size of storage -- the data is too valuable to do a semi-custom solution and be a test case for Ceph, Gluster, etc.
On the top end there are EMCs of the world with $350K/node pricing + yearly 10% support contract. Finally, there's a "rewrite the app" alternative ( call it 1 year, million $ price tag to move to object store ) that a customer company is always considering.
So the sweet spot is enterprise solution with a license cost of about $12k-$24k per year per node which has EMC/Isilon/NetApp-like functionality that unbundles software from hardware. The tricky part is that it needs to be a proven solution that works on the comparable EMC/Isilon/NetApp sized deployments. To do that one needs to have a lot of excellent engineers that cost a lot of money, a lot of sales engineers that intrinsically understand the commonality between the customers' requirements and can explain it to those engineers and a lot of very expensive test beds.
We really wanted to spend the money, but as far as I remember, we didn’t end up doing anything more than a test install.
I’d still love a middle option that could handle petabyte scale storage without the Isilon price tag.
Our lab has modest compute requirements, but need a ton of storage, so we were very interested in a mid-range storage option.
My first response to this news was "good riddance", my second thought was "I hope they open source the storage stuff."
The question is going to be: will anyone provide an intelligent way to maintain compatibility with applications that expect to interface with HDFS rather than POSIX? It's a bit of a gap right now from what we see.
fast forward 10 years and hadoop has basically been killed off by hosted storage services that are more expensive but 10000x easier to manage.
Then I'm guessing better cloud hosted options came out offering similar capabilities.
If that's the case was it really a big surprise that "cloud" hosting would eat any self-hosted platform's lunch?
But of course it didn't play out that way.
Most enterprises these days will have a massive distributed file system (S3 or equivalent) that has most of their data. And most will be running a decent percentage of data transformation jobs using Spark or some SQL layer e.g. Presto and then running BI tools like Tableau using Athena/Redshift Spectrum (or equivalent) as their SQL layer.
It's just that this is all playing out on the cloud instead of on-premise with some vendor. But you have definitely been seeing the decline of the core enterprise data warehouse.
Now what the cloud offered was so much more compelling. You had per hour pricing on the order of $20 for a minimal cluster. You had unlimited autoscaling so you didn't have to do capacity planning and go through procurement processes to pre-order hardware/licenses. And of course you had unlimited, ultra-cheap storage courtesy of S3. And it also allowed each team to have their own mini-cluster instead of everyone relying on some giant one.
I don't think the recent explosion in Data Science would've happened without the cloud.
Because Hadoop has not been killed off. In fact it is far bigger than it has ever been and growing each year. It's just that it is all transitioning to the cloud with parts like HDFS being replaced with more scalable solutions like S3/EMRFS.
But you can't say it's been killed off when on AWS you have managed Hadoop (EMR), managed Hadoop pipelines (DataPipeline) managed Spark (Glue ETL), managed Hive Metastore (Glue Catalog) etc. And similar on Azure or GCP.
> HDFS being replaced with more scalable solutions like S3/EMRFS.
Thank you for re-affirming my point.
I'm not going to go into the typical HN "I'm going to nitpick your point down to the bone and argue you over the pointless details", but 2012 Hadoop is not the same thing as the tools you're describing.
Simply put, Hadoop is not HDFS, it is a much larger Apache ecosystem that includes Spark, Hive, MR, NiFi, etc. All of which are doing fine.
Just because they're compatible doesn't mean they're synonymous. You can use spark with mesos or kubernetes instead of YARN.
Hadoop vendors are simply getting killed by the cloud. AWS EMR, Azure ML, Google Dataproc. And many people are simply swapping HDFS for S3 or similar object stores. Never seen anyone switching to very expensive, more HPC style NVME setups. That seems strange given the data sizes we are working with.
And 95% of what Data Scientists do is ETL, Data Curation, Feature Engineering etc and which Spark is still the overwhelming dominant player. And Spark uses Hadoop filesystem API underneath which is trivial to replace hdfs:// with s3://
Everyone's data volume is different as well...working a proposal for several petabytes of NVMe storage right now. It's not everyone, but use cases are out there for large volumes of high performance storage.
Fair point on the ability to swap the hdfs interface trivially, I hope it's that simple everywhere. Another issue we run into is that these companies have invested heavily in these Hadoop clusters, and would prefer to continue to get some use out of the mountains of HDDs that are captive in these environments. So tiering/HSM functionality is another facet of the issue that these environments will anchor for quite some time.
And then you just plug a ton of them into what kind of "motherboard" to interconnect them? I guess Infiniband based bus? or I am totally off-base here?
Any more details are appreciated, if only to satisfy my curiosity :)
It's at least 4.
https://www.logicalclocks.com/millions-and-millions-of-files...
Our business model is to build a data science platform, Hopsworks, around our distributed metadata layer. And yes we use YARN (training models) and also Kubernetes (serving models). The choice of resource mgr is really just an implementation detail, as the platform is backed by a REST API.
Hadoop was the floppy disk of big data. Ubiquitous, but always beat by other solutions.
All that is happening is that the vendors are being killed by the cloud.
DataPipeline is their managed Hadoop for ETL.
Glue ETL is their managed Spark.
And it all absolutely counts towards the market.
Plus everyone is moving away from Hadoop and HDFS, so it's Spark, S3, GCS, Databricks, etc
References:
Spark Streaming is just one part of Spark and the far less used part as well.
Spark hasn't been killed by anything and is still heavily used for ETL and ML.
Why spend $10k per node in an upfront licensing deal when you can buy it from AWS for a few dollars an hour.