Cloudera and Hortonworks merger means Hadoop’s influence is declining
venturebeat.com
venturebeat.com
I also don't agree with the author's assertion that Spark is "Scala centric". Yes, Spark is written in Scala, but PySpark is definitely a first class citizen. Databricks maintains a MLFlow project to make it easy to use Python with Spark: https://databricks.com/blog/2018/06/05/introducing-mlflow-an...
To be fair, prior to Spark Dataframes (i.e. the days of pure RDDs), the only way to get performance out of Spark was to write Scala code. The serialization overhead of PySpark precluded it from large-scale data engineering workloads. Most companies rewrote their PySpark code in Scala for production.
Now that we have Spark dataframes, PySpark performance is mostly on-par with ScalaSpark for many SQL-amenable operations. And with Apache Arrow in-memory support on Spark >2.3, the Python serialization overhead problem goes away.
But Spark is still to some extent Scala-centric. The documentation is trilingual, but there is still a distinct Scala-first culture.
If I'm not mistaken, Hadoop was created at Google to handle weblogs (side note: in 2014, Google announced that Hadoop was no longer being used internally)
Enterprise data however is primarily structured and relational, which is more suited to handling in a database-like system. Hadoop was never designed for this use case. Scalable cloud databases like Redshift, Aurora etc. always seemed to be a better fit. Cloudera created technology like Impala and Kudu to address this but not sure about the uptake there.
Hadoop was based on Google's published papers on the Google File System (GFS) and MapReduce. The later project, HBase, was a direct carbon copy of Google's BigTable paper.
By the time Hadoop reached maturity, Google had mostly moved onto newer technologies such as Colossus and Megastore, and mostly doesn't use MapReduce anymore.
To my knowledge, Google has never used Hadoop internally, although you can lease a hosted version of Hadoop on Google Cloud Platform.
Once IBM mainframe, the king of CAp, is put in a major AWS/MS/GCP data center expect them to gobble Cloudera. Or Principal corporation goes nuts and starts taking on Guidewire.
MapR is still out there.