(although it might look good on your CV)
"Not using a hadoop is like cutting down a forest with a single chainsaw. Using it is like cutting a forest down with unlimited hatchets. It will usually be faster and cheaper with the chainsaw."
I read the link and agree with most of it. All except for when he gets to the 5TB part, but I get that it was written in 2013 and probably not to a fortune 500 audience. If Teradata Cloud could get a reasonably priced entry level tier going, I think they'd be talked about a bit more.
And OP, if you see this, using raw M/R to do things like graph processing is going to be a HUGE pain in comparison to all the libraries built around it.
If you know Cassandra or other NoSQL, you can try your hand at Hbase. To do anything with it beyond adding or removing data from a key, you'll need to write an application of some sort. Cataloging tweets is a decently simple exercise.
In my work, the only time I accessed the HDFS directly was doing a put/delete of a flat file CSV that I was going to load into Hive. I'm not saying there's not use cases for using HDFS, just that in the set ups I've used, I've never seen it.
If you really have to learn how Hadoop works, you'll probably have to look at the source. Write the Hello World program and get it to run on AWS EMR (as others have suggested). Then clone the source to Hadoop and open it in your favorite editor. If you have to understand the nuts and bolts, the source code is the way to go. If you only have to know how to use it, then a book + online tutorial + simple project will get you up to speed within a week or two.
Do you want to learn how to setup Hadoop clusters and Zookeeper etc. on bare metal and understand all the maintenance of such a system. Or do you want to learn how to use the tool for enabling data science projects?
If it's the former, companies like Databricks are becoming popular because they abstract a lot of that complexity away: https://databricks.com/try-databricks
Because you're coming at it from scratch I'd strongly advise you look at the future trend and start straight with Spark. It is also made by Apache and is the next generation solution. https://spark.apache.org/
To give an idea of the difference, Hadoop uses Map (transform) and Reduce (action) higher order functions to solve distributed data queries and aggregations. Spark can do the same but it has access to so many additional higher order functions as well. This makes the problem solving much more expressive. See the list of http://spark.apache.org/docs/latest/programming-guide.html#t... and http://spark.apache.org/docs/latest/programming-guide.html#a...
The Spark documentation and interpreter are good places to start.
https://github.com/apache/spark/tree/master/examples/src/mai... (Scala) and https://github.com/apache/spark/tree/master/examples/src/mai... (Python)
In general, after you get some Hadoop fundamentals, I would recommend focusing on Apache Spark instead.
Disclaimer: I'm part of the Big Data University team.
If resources are what you were expecting, Coursera used to have a course named "Mining massive datasets" that covers some of the topics I saw you mentioning in your comments (MapReduce, HDFS, PageRank etc). It was ministered by Jure Leskovec, Anand Rajaraman and Jeffrey D. Ullman.
Although the course itself is not available anymore, the video lectures are still on Youtube, and a quick search returned me the following playlist: https://www.youtube.com/playlist?list=PLLssT5z_DsK9JDLcT8T62...
If you want some reading material by the same people, you should also check this book: http://infolab.stanford.edu/~ullman/mmds/book.pdf
Past the scope of your question - but I'd also recommend learning Spark as well, it's probably more relevant and marketable at this point than learning pure Hadoop.
We have not done Hadoop yet - is there anyone that would like to help? We are considering crowdfunding to pay for all trainers so that the videos and book would be free - or should we charge per course?
[1] http://NoSQL.Org
You have two/three main components in hadoop:
- Data nodes that constitute HDFS. HDFS is Hadoop's distributed file system, which is basically a replicated fs that stores a bunch of bytes. You can have really dumb data (lets say a bunch of bytes), compressed data (which saves space but depending on the codec you may need to uncompress the whole file just to read a segment), arrange data in columns, etc. HDFS is agnostic of this. This is where you hear names like gzip, snappy, lza, parquet, ORC, etc.
- Compute nodes which run tasks, jobs, etc depending on the framework. Normally you submit a job which is composed of tasks that run on compute nodes that get data from hdfs nodes. A compute node can also be an hdfs node. There are alot of frameworks on top of hadoop, what is important is that you know the stack (ex: https://zekeriyabesiroglu.files.wordpress.com/2015/04/ekran-...). So you have HDFS, and on top of that you (now) have YARN which handles resource negotiation within a cluster
- Scheduler/Job runner. This is kinda what YARN does (please someone correct me). Actually its a little more complicated https://hadoop.apache.org/docs/r2.7.2/hadoop-yarn/hadoop-yar...
Since hadoop jobs are normally a JAR, there are several ways of creating a jar ready to be submitted to an hadoop cluster:
- Coding it in java (nobody does it anymore) - Writing in a quirky language called Pig - Writing in an SQL-like language called HiveQL (you first need to create "tables" that map to files on HDFS) - Writing generic Java framework called Cascading - Writing jobs in scala in a framework on top of cascading called Scalding - Writing in clojure that either maps to pig or to cascading (Netflix PigPen) - ...
As you can imagine, since HDFs is just an fs, there are other frameworks that appeard that do distributed processing and that can connect to hdfs in someway: - Apache Spark - Facebook's Presto - ...
And since there are so moving parts, there's alot of components to put and get data on hdfs, or nicer job schedulers, etc. This is part of the hadoop ecosystem: http://3.bp.blogspot.com/-3A_goHpmt1E/VGdwuFh0XwI/AAAAAAAAE2...
Back to your question, I suggest you spin your own cluster (this one was the best I found: https://blog.insightdatascience.com/spinning-up-a-free-hadoo...) and run some examples. There's alot of details about hadoop such as how to store the data, how to schedule and run jobs, etc but most of the time you are just connecting new components and fine-tuning jobs to run as fast as possible.
Make sure you don't get scared by lots of project names!