I'm not even sure the mapreduce code has been deleted from google3 yet.
To be fair, MR was definitely dated by the time I joined- 2007- and I'm surprised it lasted as long as it did. But it was really well-tuned and reliable.
Also the MR paper was never intended to stake Google's position as a data processing provider (that came far, far later). The MR, Bigtable, and GFS papers were written to attract software engineers to work on infra at google, to share some useful ideas with the world (specifically, the distributed shuffle in mapreduce, the bloom filter in bigtable, and the single-master-index-in-ram of GFS), and finally, to "show off".
Even though Google codebase broke perforce scaling long ago and doesn't use it anymore, the replacement still borrows a lot of perforce names and sort of API.
Certainly even in 2013 MR was definitely being used; I launched a product at that time that ran an MR because we couldn't get similar performance out of Flume yet.
While the initial implementation at Google quickly got replaced with better things, the MapReduce pattern is everywhere in the data space, and almost taken for granted now. Hadoop is basically the same: a shitty (I think HDFS is still pretty good, just not the compute part) initial implementation of the pattern that was quickly iterated and improved upon.
Also, a big reason people stopped having to think about eg rack-local operations is that most people operating on huge amounts of data now aren’t doing it on traditional generic servers, they’re using something like s3 on VMs in Public Cloud datacenters if they’re doing something relatively “low level” or more likely just using something like Snowflake, Spark/Databricks (pretty close to OG mapreduce…), etc.
We had an impl of the Pregel paper on top of the Yarn manager.
The API was painful and easy to err on but it did provide quite a bit of functionality.
Now, of course, that stuff is all out of date. Where I am now we have custom job engine and it's way better. I imagine others have something like this too.
Things have just changed. Interconnect is now cheap and fast: 40 Gbps is commodity.
Outside of Google, most organizations with large distributed data processing problems moved on to Hadoop2 (YARN/MapReduce2) and later in present day to Apache Spark. When organizations say they are using "Databricks" they are using Apache Spark provided as a service, from a company started by the creators of Apache Spark, which happens to be Databricks.
Apache Beam is also used outside of Google on top of other data processing "engines" or runners for these jobs, such as Google's Cloud Dataflow service, Apache Flink, Apache Spark, etc.
Some info on flume: https://research.google/pubs/pub35650/
To quote from there: "MapReduce and similar systems significantly ease the task of writing data-parallel code. However, many real-world computations require a pipeline of MapReduces, and programming and managing such pipelines can be difficult."
So map reduce is in the DNA of many data computation flows instead of a thing in off itself.
In terms of usability the other two main innovations were to make it easier to program a workflow that chained MapReduce operations (without an intermediate, expensive, blocks-until-all-nodes-done disk write step, nor a jankass orchestration engine) and subsequently to declaratively specify the desired output (eg SQL) without requiring the user to specify the implementation.
They’ve since added more stuff like streaming, ML, whatever, but the biggest change from 1st to 2nd gen is really in the data topology.
Rama seems like if you are a fullstack or backend dev then it can provide you an easy way to have a(low latency) view of your data to build upon. If you are a Data Scientist you can use the thing to pull necessary data for analysis and slice and dice it.
The best place to end up is something like PySpark/Snowpark as a better API for SQL is really useful when doing complicated things.
You still need to have a standard SQL layer though, as otherwise you'll cripple adoption.
For streaming there is flume and beam, or just load important data into Spanner.
Parquet is a format, and not execution engine or paradigm?.. You can totally mr over parquet.
Urs Hölzle, the former head of Google TI (Technical Infrastructure), discussed in public some of the challenges and reasons for creating Google Cloud Platform as a platform, and backing projects like Kubernetes.
Over time, Google has become a proprietary tech "island" in severals ways and arguably more fragmented than other large tech companies, such as Microsoft and Amazon, which happen to both have commercial cloud offerings, and Meta/Facebook. While all of these companies certainly have challenges with not-invented-here ("NIH") syndrome, and lots of internal, proprietary tools, as a software engineer at one of these three, odds are you will use and touch more commercial and open-source technologies than you would at Google. Google itself still struggles with having teams and projects use GCP for internal work versus Borg/etc; and there are plenty of valid reasons why Google teams don't use GCP.
The proprietary tech "island" issue is a non-trival concern when you need to hire new software engineers from industry/outside and ramp-up time with some of these systems may be 6-months or even greater; today Alphabet/Google is at around 200k+ FTE, and you aren't going to be able to find many engineers outside that have experience with Borg/Flume/Spanner/Monarch/etc. Likewise when you are an experienced Google software engineer looking to work elsewhere, you need a translation map to figure out what tools outside are similar to the ones from inside.
Google's proprietary tech island has its legitimate reasons for existing, and when people say 'xyz' commercial/open-source thing is "better," they often mean it is better for their problem at hand.
At Google, a decade-plus ago many of the problems it had to solve were problems that few other organizations had, such as large-scale data processing (to be made cost-efficient on commodity hardware), and it needed to create a number of tools/platforms as solutions such as Map Reduce/GFS.
Many of these tools and platforms were discussed via papers, and inspired open-source work. In the Map Reduce case, it changed how Apache Hadoop itself took shape, and the lessons from all of these later led to things like Apache Spark.
The idea of losing a battle can only be applied with the benefit of hindsight, and many of the Google examples given were created at a time where there were no peers, nor at that time was Google interested in selling these things as commercial products at the time (i.e. GCP vs AWS vs Azure); it built these things according to its unique internal needs that few other organizations could relate to. I acknowledge that I am intentionally leaving out organizational politics, and culture (e.g. PERF) as non-trivial contributors for this result).
Went to a GCP event once, expected it to be like the aws one… it was 100% marketing and the Wi-Fi didn’t work. So yea, they drop the ball a LOT when trying to interface with the developer community.
This is a valuable comment and I didn't mean to nerdsnipe you.
Interesting opinion but not supported at all by evidence. Most non-Google datasets are small and stored on off-the-shelf heterogenous hardware, so HDFS / Mapreduce for streaming OLAP is a great fit. Cassandra (BigTable) and Parquet (Dremel) plus Cloudera’s Impala had much quicker time-to-market when large-scale BI became more relevant.
“Obsolete” for Google problems sure, but Google problems largely only happen at Google. Stuff like ad targeting and ML look a lot different for products outside the Chocolate Factory.
1. created the software for their own needs 2. maintain a developer team to improve it and address requirements/pain points/integration 3. have an internal pool of experts in the form of the developer team and “customers”/early adopters 4. most likely have other proprietary systems like Borg or Colossus which integrate with the software very well, which OSS like Hadoop may not (another example: OSS Bazel vs Blaze+Forge+Piper+Monorepo structure).
Something like HDFS was hugely painful for many teams because they had no idea how it worked or how to debug it, had no idea how to fix it or extend it, and didn’t have any good tooling to understand why something was slow. All they could do was try to configure it, integrate with it, and find answers for their problems online. That’s because HDFS was “free” but a team capable of properly maintaining, supporting/operations, and developing HDFS was extremely expensive.
Wish more organizations understood this part before adopting Bazel.
Despite this misuse of Hadoop, another guy really loves it, and decides to start a project rewriting everything to use MapReduce. A new guy started, got assigned to the "MapReduce project"... he worked on this for over a year. It never made it to production.
In the HDFS case with Hadoop's ecosystem, consider Hive, BigTable, Drill, and even Spark when running on YARN.
In the peak days of Hadoop, many organizations were primarily on-prem, and S3 or S3 compatible object stores were mostly reserved for people using AWS.
what is the better fit if you need to store 1PB of data for cheap?..