How to Learn Hadoop for Free
johnwittenauer.net
johnwittenauer.net
You've got paid AND you got to know some Hadoop! (Worked for me; YMMV)
I agree Hadoop is no longer MapReduce. It's HDFS+YARN. That's it. Distributions package up Spark/Flink/Kafka/PrestoDB with the HDFS/YARN core.
At Hops, we've scaled the core HDFS by >16-37X ( https://blog.acolyer.org/2017/03/06/hopfs-scaling-hierarchic...) and we have a distribution called Hopsworks with support for Spark/Flink/Tensorflow. Nobody uses MapReduce on our platform.
The thing that has killed Hadoop, imo, is Kerberos. In Hops, we have switched to using TLS/SSL certificates instead of Kerberos, and that enables us to implement dynamic roles. Dynamic roles allows us to build a software-as-a-service platform, where projects are securely isolated from one another.
My own perspective is that there are lots of businesses that haven't yet needed the capabilities provided by a platform like Hadoop, but they likely will in the future. So the market may be saturated based on current needs but that market will continue to expand. Whether it's Hadoop (YARN/HDFS/etc.) that wins that market share or some other stack like Spark/Mesos remains to be seen.
You reference the MapR distribution for their training material, and its interesting that their version of HDFS is a reimplentation in C++ (MapR-FS). Its part of the reason I settled on MapR to use tools like Apache Drill, because the filesystem becomes usable to non-Hadoop tools via NFS (i.e. Awk).
Given a shift in some categories away from map-reduce to other approaches, could Hadoop eventually just become a collection of distributed filesystems and job schedulers?