Additionally, the “two pizza teams” style of Amazon’s development was more amenable to the AWS model of development where each piece of Amazon could run effectively in AWS [1]. Google is structured differently, and it’s only recently that both AWS and GCP could even approach running the largest systems from a company like Google.
[1] Yes I know that AWS wasn’t built specifically for Amazon engineers, but the fact that it could accommodate them over time has been a boon for their cloud, and it gives enterprises confidence that of AWS’s strategic importance and desire to make it world class. There isn’t a large hidden layer of services that Amazon has access to that the general public doesn’t.
Steve Yegge talks about it on his podcast: https://youtu.be/9v4z46Ea35Q
Amazon actually works backwards like they say - they ask customers what they want, and build it.
Google thinks they're smarter than everyone else so they can't comprehend that developers out there in the world want different things than the developers inside Google.
You can see this permeate even in this documentary, which is ostensibly about a success - when Google actually did launch something that developers really liked - K8S.
They didn't start from developer needs or wants. They started by noticing AWS was eating their lunch, and wanting to compete with them, but realizing they couldn't win by doing it directly.
Also, while Amazon does work backwards from customers, part of them working backwards is how to translate customer requirements into a successful business. Amazon is doing this as a for profit company, not spending money on altruistic initiatives.
E.g. purely for the benefit of customers, having AWS services be multi-cloud and open source or reducing egress cost would be best.
To the point that k8s is base operating system for... Chick-Fil-A restaurants, deployed locally, controlling the industrial kitchen hardware.
Amazon didn't try - they just knew that they were good at building high scale distributed systems, and thought they could build services for others. That's what S3 and SQS and EC2 were - built from scratch, not trying to figure out how to externalize the internals.
They had years of headstart on most infrastructure technology that’s ubiquitous now, but didn’t bother to offer it to others. Perhaps they thought only Google should benefit from it, while Amazon saw a business opportunity with AWS.
The empirical evidence seems to point that way, a lot of people were noticing that many of the Google stuff became open source or got open source equivalents from Google (Borg/Kubernetes, Blaze/Bazel, etc.) only after a critical mass of former Google employees left Google and since they liked Google tooling started replicating it as open source software outside of Google. In some cases that open source software got traction and Google seemed to be falling behind in some areas.
Which basically tells me that in the software world, you can't really have trade secrets. I could be wrong, though.
For Prometheus they could have just chosen to aim for smaller businesses, since that's 99% of businesses out here.
So it could just be a tradeoff.
I just find a kind of perverse humor in the fact that outside google Prometheus is viewed as alien technology - it even has that condescending name - but inside Google borgmon was considered the punch line of memes and/or a hazing ritual.
While the text based metrics aren't best, they are a reasonable compromise, and I found that it provides a good enough common interface where if needed I can change what collects from it later on.
Are there better options? SURE! I can, however, usually get my 80% of the problem solved with Prometheus integration and worry about scaling later.
If we just talked about Borg, but didn't ship code, someone else might have set the agenda, rather than Kube.
Brian writes very succinct and clear descriptions of problems encountered inside google, breaks them down to “what worked, what didn’t work”, and then would summarize a design seed in a github comment (of which several such subsystems subsequently were implemented that have been fairly successful).
I would keep mental track of some of these comments and assign a score based on how much R&D investment at google was compressed into a 200-300 word comment. I guesstimated some at being valued at upwards of $100k a word, and jokingly asked brian if he was authorized to share $10M worth of google’s secret sauce. He laughed, but reminded me of some of the same things mentioned in the documentary (which we all knew in the project at the time) - paraphrased:
“if you can seed a community project that includes some of these patterns built on that research, and that enables $10 billion of application team value, and google can capture $1B of that in google cloud, that’s a pretty good return on investment.”
And definitely among the key contributors I worked with from google was the very human and very relatable (for HN) desire to build that shows those internal ideas off and “do a successor to borg that really takes what we’ve learned and doesn’t just copy it”.
Even if it was a business strategy to help catch google cloud up… the beauty of open source is that you can benefit from it regardless, and it moves the industry forward in a meaningful way.
Most places I have seen, architects seem to spend most of their time writing out overly verbose documents, detailed specifications, making presentations etc etc. There is pressure to show you are in charge of the project and accountable for it so you end up with all these ceremonies where you create heaps of content and dump it onto the teams.
Whereas the sort of situation you described is really where the role of an architect shines, in making sense of the chaos and showing light to the teams. You often don't need more than a short message or a brief meeting to convey these ideas, but such things don't lend upward visibility to what you're doing. I've always wondered if the life of an architect in Google involves these things or is quite different from most organizations?
They protect this aggressively, and the tech is impressive.
I remember the first time I accidentally saw a custom google node when I went there to pick a friend up from there, and I had gone to his desk and I saw a really nice looking PCB on his desk he forgot about, and I exclaimed "WHATs that' -- I was pretty familiar with a wide range of server boards, and I knew that was custom immediately...
He panicked covered it and we didnt really talk about it much, but it was known really early that they were building a lot of custom kit.
~2005 maybe?
Is this true? There was the story about the resources they need for Christmas shopping that went unused for the rest of the year.
But recently somebody mentioned here (I don't remember where) that AWS was a separate business and that Amazon itself initially didn't use any resources from AWS.
so while there’s the ideal state of running amazon on a large multi tenant cloud, there was absolutely no safe and reasonable path to making that happen by extending the existing infra. it would’ve also hampered aws by forcing them to focus way too large way too early.
But they're the least diversified, after Facebook, agreed.
So yeah they weren't profitable but they weren't just scraping by either. I think that's one of those things that people miss when they compare a not-so-profitable startup to Amazon. The size, scale, and huge number of product markets that Amazon plays in is really pretty astounding for a company that is only 25 years old.
Not actually the case. AWS was a completely independent stack (from hardware on up) - we couldn't use it internally for quite some time (years); and even then, only in a limited fashion (e.g. only S3 for some time)
A lot of people used to think that Amazon.com was subsidising AWS. Financials from other VPS providers suggested this to be the case.
Google app engine was definitely NOT competing with AWS. They were not selling to enterprises.
i think the interesting thing is that while you’d think that is the ideal everyone should strive for, the reality is that’s inapplicable or detrimental to run most businesses that way.
I'm confused, did you mean "without" instead of "with"?
Also:
> i think the interesting thing is that while you’d think that is the ideal everyone should strive for, the reality is that’s inapplicable or detrimental to run most businesses that way.
Why would you say that is?
> Why would you say that is?
many/most businesses don’t have the margins/scale/time to focus on it. in a lot of businesses (amazons included) the goal for the business is to try as many things as fast as possible and scale that which works and abandon what fails. the goal of infrastructure there is to get out of the way and enable the layers above to build things as naively and simply as possible with little to no churn.
A traditional FOSS project is not limited in this way. The developers often just start a new major version and overhaul everything. But they don't have to answer to "the business" or huge institutional users. Commercially driven open source will always be a product beholden to stakeholders, rather than technology designed to operate as well as it can.
And regular users hate this.
Jamie Zawinski describes this funnily, eloquently and also offensively:
https://www.jwz.org/doc/cadt.html
> This is, I think, the most common way for my bug reports to open source software projects to ever become closed. I report bugs; they go unread for a year, sometimes two; and then (surprise!) that module is rewritten from scratch -- and the new maintainer can't be bothered to check whether his new version has actually solved any of the known problems that existed in the previous version.
But he's right, and that's why:
* regular application developers who don't care about ideology
* and mass market users
hate desktop Linux.
Users just want stable software. Developers want software that works better and solves more problems. The compromise is for users to stick with old major versions. At some point that's no longer tenable, but at least this way the software is free. Certainly no worse than when the same thing happens but you're paying for it.
AWS being the dominant player means that if they make any silent changes there will be a SO/GitHub issue/HN thread about it. But with GCP without their pricy professional tier support you can feel a bit unsure and lost all the time about your understanding of it.
I don't know about you, but in 48 hours I've usually already built a workaround and moved on to other projects.
Not saying that AWS support is worse or even equal to GCP support, but just commenting that cloud support in general feels pretty useless and an afterthought at best.
And let's not even get started on the horrible, awful UI in AWS console.
I've had this happen multiple times for multiple services and the person's i was able to work with had enough knowledge and access to either see detailed logs and walk me through some undocumented feature or file a bug report by days end.
I'm spoiled in that regard but I've never had non-enterpise support with anything other than the 4 hour sla.
But at the same time I can refer to enterprise-tier support at AWS, which has gone past ticket stage into "regular calls with AWS support engineers to help architect our solution" and certain critical topics ended up ghosted (my first reaction was that support had no idea about the functionality we asked about, later experienced taught me that at best, the AWS documentation lied by omission. At worst it was divorced from reality.)