MapReduce, TensorFlow, Vertex: Google's bet to avoid repeating history in AI
supervised.news
supervised.news
I agree with another commenter, where Google missed the boat was cloud. Ironically, AppEngine came out pretty early in 2008, but its original incarnation was much too "Google-specific". Google just didn't have the corporate DNA to understand that, sure, Datastore can allow infinite horizontal scalability, but most people don't want to deal with eventual consistency, and they just want a plain-old SQL DB. It took Google a long time to come around to understand how other businesses use their infrastructure.
That was the first project which trained me not to use Google products. I remember reporting by a ton of stuff and then never hearing back and things never improving until much (multiple years?) later when I’d stopped using it. I was profoundly unsurprised when AWS ate their lunch.
Amen, amen, amen. This is so frustrating to me because in general I really like GCP as a platform, and I think they're doing great things with stuff like Cloud Run.
Case in point, I like Firebase a lot, and I think Firebase Auth (aka Google Identity Platform) is a great product. But it took them about 2.5 years between releasing their initial support for a second factor, which was originally SMS only, and allowing a TOTP (e.g. Google Authenticator/Auth) 2nd factor. And they still don't support "remember this device" functionality, so every time someone logs in they always need to enter the 2nd factor. I literally can't think of any major site (especially Google's own) that doesn't support "remember this device", and yet it just must be a detail that's not "important enough" to get someone at Google a promotion.
They did close my textbook example of that but it took half a decade to implement HTTP to HTTPS redirects, which basically every large customer wants to make compliance easier:
Here is my similar (though not as long) example - point-in-time-recovery on postgres: https://issuetracker.google.com/issues/78448400. Interestingly, it had the same problem with 0 communication from a Google PM until early 2020, so maybe they hired someone at that time who convinced them you can't treat business customers the same way you do consumers.
Sure it wasn't enough alone to outlet outplay AWS, but Google would have been worse off without it
GKE (with Autopilot) is firmly set in my "day 1 toolbox" if I were to start a new company, assuming K8S was the right tech choice.
As long as you don't rely on implementation details like a specific Ingress Controller that needs special annotations, this is exactly what Kubernetes does, and the reason why you have Helm Charts that don't need to take into consideration which the underlying cloud is.
Going back to what we had before, you had to be very aware of exactly which cloud provider you'd use. And since AWS had about 80+% of the market then, that meant that instructions for how to run things on Google's cloud were not very good, if they existed at all.
That was not the point of Kubernetes.
And if it was, it did not realize that point for many years.
The problem is that GCP for some reason just doesn't align with many corporate customer requirements (ie: support, stability, never kill anything, available resources, etc.) and an engineer friendly product (GKS) does nothing to alleviate that. It makes GCP better in an area it was already better in and doesn't at all fix the tangential reason companies actually avoid it (or rather from what I hear either migrate off of GCP eventually or use it as just a secondary cloud provider).
Google famously uses lots of custom versions of infrastructure software while the rest of the world uses something else. Obviously that has served then extremely well, but the rest of the world wants stuff they are familiar with. It really took Google years to fully grasp that.
Most of Google's problems in B2B reflect that the organization has little concept of customer service, in the sense of serving customers. It's not the technology.
The current problem with AI for Google is that the good stuff costs a bit too much to give away with search. That problem may be solved, or evaded. Most searches or questions don't need a large language model. I expect to see systems where you ask a question, and if it's hard, you get something like "That's a hard question. Give me a minute to work on that. Meanwhile, here's a word from our sponsor."
This, exactly. Google is fantastic at research, great at development. It's terrible at product. Or, specifically, all the bits of product that aren't R&D. MapReduce, Tensorflow and Kubernetes are all examples of technical successes that Google has spawned. None is a product. You can't buy them, buy support, or, you know, actually talk to someone if something goes wrong.
That's why Google cloud has failed so far. It's not the tech; big G can go toe to toe with Amazon and MS on that. It's that enterprises want contracts and support and relationship managers and all that stuff. That just doesn't seem to be in Google's DNA. Unless they can adress that, they're not going to win with big org customers.
They might win the startups - where the customers are themselves techies who are OK with self supporting, looking up Stackoverflow, asking ChatGPT, etc. But unless they can address the people angle, I don't see Vertex gaining meaningful traction against Amazon or OpenAI/MS.
Perhaps, but they seem to be in a better position than many to get that cost to a reasonable level.
They've got the scale, the research side, lots of experience with models of different sizes, they own TensorFlow, they produce and deploy their own Tensor chips (so I would guess are not dependent on Nvidia's tech at 10x markup), and have lots of experience in datacenter power efficiency.
This is all just based on what I've read in tech news. I happen to work at Google but don't have any insider info here. I'm biased, but they seem to be in a good place.
Outside of Google, most organizations with large distributed data processing problems moved on to Hadoop2 (YARN/MapReduce2) and later in present day to Apache Spark. When organizations say they are using "Databricks" they are using Apache Spark provided as a service, from a company started by the creators of Apache Spark, which happens to be Databricks.
Apache Beam is also used outside of Google on top of other data processing "engines" or runners for these jobs, such as Google's Cloud Dataflow service, Apache Flink, Apache Spark, etc.
Some info on flume: https://research.google/pubs/pub35650/
To quote from there: "MapReduce and similar systems significantly ease the task of writing data-parallel code. However, many real-world computations require a pipeline of MapReduces, and programming and managing such pipelines can be difficult."
So map reduce is in the DNA of many data computation flows instead of a thing in off itself.
In terms of usability the other two main innovations were to make it easier to program a workflow that chained MapReduce operations (without an intermediate, expensive, blocks-until-all-nodes-done disk write step, nor a jankass orchestration engine) and subsequently to declaratively specify the desired output (eg SQL) without requiring the user to specify the implementation.
They’ve since added more stuff like streaming, ML, whatever, but the biggest change from 1st to 2nd gen is really in the data topology.
Rama seems like if you are a fullstack or backend dev then it can provide you an easy way to have a(low latency) view of your data to build upon. If you are a Data Scientist you can use the thing to pull necessary data for analysis and slice and dice it.
The best place to end up is something like PySpark/Snowpark as a better API for SQL is really useful when doing complicated things.
You still need to have a standard SQL layer though, as otherwise you'll cripple adoption.
For streaming there is flume and beam, or just load important data into Spanner.
Parquet is a format, and not execution engine or paradigm?.. You can totally mr over parquet.
I'm not even sure the mapreduce code has been deleted from google3 yet.
To be fair, MR was definitely dated by the time I joined- 2007- and I'm surprised it lasted as long as it did. But it was really well-tuned and reliable.
Also the MR paper was never intended to stake Google's position as a data processing provider (that came far, far later). The MR, Bigtable, and GFS papers were written to attract software engineers to work on infra at google, to share some useful ideas with the world (specifically, the distributed shuffle in mapreduce, the bloom filter in bigtable, and the single-master-index-in-ram of GFS), and finally, to "show off".
Even though Google codebase broke perforce scaling long ago and doesn't use it anymore, the replacement still borrows a lot of perforce names and sort of API.
Certainly even in 2013 MR was definitely being used; I launched a product at that time that ran an MR because we couldn't get similar performance out of Flume yet.
Interesting opinion but not supported at all by evidence. Most non-Google datasets are small and stored on off-the-shelf heterogenous hardware, so HDFS / Mapreduce for streaming OLAP is a great fit. Cassandra (BigTable) and Parquet (Dremel) plus Cloudera’s Impala had much quicker time-to-market when large-scale BI became more relevant.
“Obsolete” for Google problems sure, but Google problems largely only happen at Google. Stuff like ad targeting and ML look a lot different for products outside the Chocolate Factory.
In the HDFS case with Hadoop's ecosystem, consider Hive, BigTable, Drill, and even Spark when running on YARN.
In the peak days of Hadoop, many organizations were primarily on-prem, and S3 or S3 compatible object stores were mostly reserved for people using AWS.
1. created the software for their own needs 2. maintain a developer team to improve it and address requirements/pain points/integration 3. have an internal pool of experts in the form of the developer team and “customers”/early adopters 4. most likely have other proprietary systems like Borg or Colossus which integrate with the software very well, which OSS like Hadoop may not (another example: OSS Bazel vs Blaze+Forge+Piper+Monorepo structure).
Something like HDFS was hugely painful for many teams because they had no idea how it worked or how to debug it, had no idea how to fix it or extend it, and didn’t have any good tooling to understand why something was slow. All they could do was try to configure it, integrate with it, and find answers for their problems online. That’s because HDFS was “free” but a team capable of properly maintaining, supporting/operations, and developing HDFS was extremely expensive.
Wish more organizations understood this part before adopting Bazel.
Despite this misuse of Hadoop, another guy really loves it, and decides to start a project rewriting everything to use MapReduce. A new guy started, got assigned to the "MapReduce project"... he worked on this for over a year. It never made it to production.
what is the better fit if you need to store 1PB of data for cheap?..
Urs Hölzle, the former head of Google TI (Technical Infrastructure), discussed in public some of the challenges and reasons for creating Google Cloud Platform as a platform, and backing projects like Kubernetes.
Over time, Google has become a proprietary tech "island" in severals ways and arguably more fragmented than other large tech companies, such as Microsoft and Amazon, which happen to both have commercial cloud offerings, and Meta/Facebook. While all of these companies certainly have challenges with not-invented-here ("NIH") syndrome, and lots of internal, proprietary tools, as a software engineer at one of these three, odds are you will use and touch more commercial and open-source technologies than you would at Google. Google itself still struggles with having teams and projects use GCP for internal work versus Borg/etc; and there are plenty of valid reasons why Google teams don't use GCP.
The proprietary tech "island" issue is a non-trival concern when you need to hire new software engineers from industry/outside and ramp-up time with some of these systems may be 6-months or even greater; today Alphabet/Google is at around 200k+ FTE, and you aren't going to be able to find many engineers outside that have experience with Borg/Flume/Spanner/Monarch/etc. Likewise when you are an experienced Google software engineer looking to work elsewhere, you need a translation map to figure out what tools outside are similar to the ones from inside.
Google's proprietary tech island has its legitimate reasons for existing, and when people say 'xyz' commercial/open-source thing is "better," they often mean it is better for their problem at hand.
At Google, a decade-plus ago many of the problems it had to solve were problems that few other organizations had, such as large-scale data processing (to be made cost-efficient on commodity hardware), and it needed to create a number of tools/platforms as solutions such as Map Reduce/GFS.
Many of these tools and platforms were discussed via papers, and inspired open-source work. In the Map Reduce case, it changed how Apache Hadoop itself took shape, and the lessons from all of these later led to things like Apache Spark.
The idea of losing a battle can only be applied with the benefit of hindsight, and many of the Google examples given were created at a time where there were no peers, nor at that time was Google interested in selling these things as commercial products at the time (i.e. GCP vs AWS vs Azure); it built these things according to its unique internal needs that few other organizations could relate to. I acknowledge that I am intentionally leaving out organizational politics, and culture (e.g. PERF) as non-trivial contributors for this result).
This is a valuable comment and I didn't mean to nerdsnipe you.
Went to a GCP event once, expected it to be like the aws one… it was 100% marketing and the Wi-Fi didn’t work. So yea, they drop the ball a LOT when trying to interface with the developer community.
While the initial implementation at Google quickly got replaced with better things, the MapReduce pattern is everywhere in the data space, and almost taken for granted now. Hadoop is basically the same: a shitty (I think HDFS is still pretty good, just not the compute part) initial implementation of the pattern that was quickly iterated and improved upon.
Also, a big reason people stopped having to think about eg rack-local operations is that most people operating on huge amounts of data now aren’t doing it on traditional generic servers, they’re using something like s3 on VMs in Public Cloud datacenters if they’re doing something relatively “low level” or more likely just using something like Snowflake, Spark/Databricks (pretty close to OG mapreduce…), etc.
We had an impl of the Pregel paper on top of the Yarn manager.
The API was painful and easy to err on but it did provide quite a bit of functionality.
Now, of course, that stuff is all out of date. Where I am now we have custom job engine and it's way better. I imagine others have something like this too.
Things have just changed. Interconnect is now cheap and fast: 40 Gbps is commodity.
I remember vaguely one task where i tried to join the index of the internet to some other dataset, which required joining across datacenters. The job ran with some warning. Then I got a friendly message from an SRE saying something like "it can cost thousands of dollars to run a join like that, so, no worries, just only do it if its something important." Of course, I wasn't doing anything important. I was just enjoying running MapReduce.
My only thought is, I wonder how well the nvidia/google partnership will do against Azure/Intel (I believe Azure invested heavily in FPGA's for their ML use cases).
And Microsoft is seemingly shooting for their own AI hardware: https://www.nextplatform.com/2023/07/12/microsofts-chiplet-c...
After all if you compare Nvidia’s success with H100 sales versus GCloud TPU sales, it would be easy for Sundar to say “if you can’t beat em join em” and just maintain TPU team for inference which is more closely tied to wall street margins.
XLA (backing Jax) has been supporting the NVIDIA GPU for a long time. (Otherwise TensorFlow were TPU-only, which cannot be true.)
I think the announcement is more ceremonial than technical, maybe expecting this kind of superficial reaction.
LLMs are shaping up to be the second such failure.
Amazon, on the other hand, has no qualms about converting their internal expertise into products ("turn every major cost into a source of revenue" [0]) and giving customers what they ask for.
[0] https://twitter.com/BrianFeroldi/status/1284795114187919362
What Google missed on, was taking cloud outside of itself as a visible customer product. AWS leapt into the breach, but the irony is that we want to use AWS to run a technology platform which Google has significant DNA in.
More of a proof of how much it was google's to lose than anything against AWS.
So? Does it translate to them escaping the reputation of mostly making money from ad?
Go? Inside. They hired Pike and Thompson amongst others. Kubernetes, inside. Pike and Thompson had been working on plan9, which in many respects foreshadows Kubernetes. The language couldn't really be proprietary (google maintain very strong control over it) and Kubernetes went out as an open source quite quickly. QUIC also underwent this transformation from in-house to shared.
wait a few years. then google will be blamed for it, not praised.
I never expected google to end up in the innovator's dilemma but here they are.
Seriously though, the AI race has just started. On a horizon that matters for large scale and ongoing adoption (and thus persistent corporate profits and valuations as opposed to manic hypes) nothing has been decided.
The hardware and software mix that will deliver this large scale adoption is not clear yet. Past experience (and reason) suggests it should be relatively cheap (commoditized) and easy to use.
Achieving X but at 0.1% of the cost will be the game that people will thrive at, betting on gargantuan volumes instead of gargantuan prices.
Given "AI" is more or less linear algebra the world is crying for commoditized vectorized compute. Its a solved problem. The world will get what it wants.
We can be happy that at Google, at least, these things can seep out and be of some benefit to the rest of us.