I don't regret hand-installing Kubernetes on AWS, if only because I know a lot more about how it works now. But I rather just spin it up with GKE.
We do invest a lot in making Google Container Engine a great experience, including integrating it with other parts of GCP (e.g., IAM [1]), but at the same time, the core is plain vanilla Kubernetes.
[1] https://cloud.google.com/container-engine/docs/iam-integrati...
Dropbox built out their own environment (and did it migrating 500PB out of S3) [1] [1a]. As did Twitter [2]. And Facebook [3]. And GitLab [4] (too soon?) As well as Mixpanel [5]. Even Twilio is multi-cloud (last time I checked it was split between AWS and Rackspace; this was several years ago during an interview, so maybe its changed). Sure, start in Google, or AWS, but at some point you will either need to use multiple compute/storage providers (redundancy) or go to your own gear (redundancy and cost).
Example: "In 2014, Moz CEO Sarah Bird said that it was spending “$6.2 million at Amazon Web Services, and a mere $2.8 million on [its] own data centers.” Simply put, the cloud killed its margins." [6]
EDIT:
simonebrunozzi: Forgive me, but when you're talking about hundreds of millions of dollars in spend, "easy" is relative. It is much easier when you're not relying on underlying primitives that are difficult to reproduce on your own at another provider (witness how terrible Open Stack is; no one wants to do that if they don't have to).
Am I minimizing the effort involved for this discussion? For sure. But the money involved...it solves most problems you would have migrating between providers.
> It seems to me that you have no serious experience in the real world.
You are entitled to your opinion. I have seen the pain, and it is relative. Its easier when someone says, "Here is the budget, just fix the problem", and your vendor's (AWS/Google) margins are 20-40% (these are real margins pulled from earnings reports); that's a lot of money you can put back in your own (or your shareholders') pockets.
If you we're spending $2 billion dollars, and I told you I could save you $400 million by spending $100 million, wouldn't you take that deal? Even at $200 million, its a bargain!
For less than what Snap is spending on Google cloud infrastructure, SpaceX built a rocket that can take a payload to orbit and return the first stage successfully (SpaceX has taken on ~$1.2 billion in funding over the last 14 years). Moving out of a cloud provider is comparatively hard?
EDIT: Maybe this is a roundabout way to kick back to Google in order to get preferential treatment on the Ad network. It sure isn't a logical decision.
EDIT 2: @ashayh: I'm not saying go back to good ol' bare metal. For $2 billion, you could build your own cloud provider out as an internal operation. The amount that's being spent on Google Cloud is egregious, and worse yet, common shares have no voting rights to push back against poor decisions like this.
EDIT 3: @hueving: HN throttles my posting; editing this comment is my only way to respond. Sorry about that!
[1] https://www.wired.com/2016/03/epic-story-dropboxs-exodus-ama...
[1a] http://www.informationweek.com/cloud/cloud-storage/how-dropb...
[2] https://blog.twitter.com/2016/overview-of-the-twitter-cloud-...
[3] https://code.facebook.com/hardware/ | http://www.zdnet.com/pictures/facebooks-data-centers-worldwi...
[4] https://about.gitlab.com/2016/11/10/why-choose-bare-metal/
[5] https://code.mixpanel.com/2011/10/27/why-we-moved-off-the-cl...
[6] http://www.thewhir.com/blog/moving-away-from-aws-cloud-dropb...
https://cloud.google.com/container-engine/
Also them starting the project along with the knowledge they have internally scaling containers helps.
I'm going to go through all of the service offerings this weekend - from Docker Inc's Docker for Azure and Docker for AWS to the native container services on each.
FWIW, Most of this are advertising gimmicks though and Google has a pretty different internal infrastructure for orchestrating containers that has hardly much to do with K8s.
And it's not just the rather-large core team directly on GKE and k8s, nor the related products like Container Registry [1], Container Builder [2], and Container-Optimized OS [3]. GKE and k8s benefit in other ways too: Google's internal kernel team helps debug customer issues when we trace them to the kernel, and people like Kees Cook are helping with the upstream Kernel Self-Protection Project [4] that make container technology more secure. In addition to that kernel work, Google also has rather-decent security teams and they work with us to improve security in other ways too.
Finally, re: toomuchtodo's question, "Why opt for Google if you're going to use containers in Kubernetes?" Because we hope you find that Container Engine is the best place to run Kubernetes -- and benefit from the other parts of Google Cloud Platform. If you ever find GKE is not that place, and you don't derive value from the rest of GCP, then exactly as toomuchtodo puts it: "You can even move to your own datacenter at some point (relatively) easily."
[1] https://cloud.google.com/container-registry/
[2] https://cloud.google.com/container-builder/docs/
[3] https://cloud.google.com/container-optimized-os/
[4] https://www.linux.com/news/google-developer-kees-cook-detail...
Am I totally off-base, if you're able to speak to this? (Maybe it exists and I've missed it, but I'd love to see a blow-by-blow of the differences and their rationale, too, because that'd be valuable insight on how Google learns.)
When I say k8s is like borg I mean: it has the same concepts of tasks, jobs, and allocs. The scheduling of those is handled by a k8s scheduler which resembles the borgmaster scheduler (a lot of hand waving here), and the containers themselves execute in an environment much like the borglet provides for containers.
Many of the valuable features provided by borgmaster and borglet are provided in k8s and you configure them through similar mechanisms.
Beyond that, how they are implemented specifically, there are a ton of differences but for an end user who is just using k8s, not setting up and managing k8s infrastructure, it's conceptually isomorphic.
Due to the cache and read dependend nature of the database queries the latency impact is worth it.
Moving from one cloud to another, even with containers, is never easy at large scale.
(source: I have worked at AWS for 6 years, at VMware for 2, and I've seen hundreds of clients go through this exercise)
If moving operations around is insurmountably difficult, you built operations incorrectly. Put another, even broader, way with a few more implications: if you are totally reliant on one vendor for continuity of any part of your operations, you built operations incorrectly and are introducing unnecessary risk. If us-east-1 goes down and you cease generating revenue as a result, you have built operations incorrectly. That's really all there is to it. And yes, I realize this means 80%, maybe more, of the operations in in the world is built incorrectly. We just learned Snap's is[0]. Maybe even yours! And that's fine as long as you're working on it. Good news: said exercise is a good chance to fix it!
Now that half the crowd is inhaling to bombastically retort that undoubtedly controversial, yet completely true, paragraph, allow me to quickly redirect:
What's "the real world," anyway? Most of HN forgets a Windows/.NET ecosystem exists, not to mention extra-valley gigs in, say, Nebraska. Would you say the lone sysadmin holding together a hospital in Des Moines is gaining "real world" experience and able to meet you in discussion? Seriously, I hate "the real world" and the people who fire it as a volley during an argument. Even your career is not indicative of "the real world." (Nor is mine.)
[0]: Flagrantly so. I've spent the better part of an hour trying to concoct a scenario where that deal is even remotely in the win column for Snap. Still trying. You enshrined a business disincentive (nay, prohibition!) toward optimizing your opex into a five year contract and $400mm a year operating Snap was some kind of win? ... How on Earth? Even with a quarter billion DAUs...
If you're ever in Tampa or Chicago, let's grab a beer or dinner my treat. Would love to share war stories.
An additional point is that even the second tier PaaS/IaaS providers (like GE/Predix, for example) are trying to get out of the DC ownership business. There's no compelling reason for them to keep their own DCs when a) it's expensive to run & keep fresh, b) it's CapEx, and c) it costs them expensive heads to organize and manage everything.
That shows about as much of a lack of real world experience as you are bemoaning. It doesn't work that way in real life.
"Real world" and "real life" dismissals of said ideal are stupid, lame excuses for people who want to make a stupid, lame argument to cover up stupid, lame technical debt in operations by either (a) assailing the credentials of the speaker, as this entire thread has spent much time doing, and/or (b) providing a "well, everybody else isn't investing in this, neither should we" cop-out. And yes, you are doing it wrong. Rather than getting defensive, pulling the knives out, downvoting on sight, and trying to wipe away or justify doing it wrong, why don't you instead realize that it's motivation to do it better? We should all do things better, and every time a single AWS region goes down and takes out half the Internet I sigh because it's this exact, stupid, lame justification fest that results in that situation.
Work toward correct operations. You will never reach it. The cognitive dissonance of these two statements is totally acceptable. If you build a new service today and don't account for DR and security up front it just goes to show me that you haven't learned from the very public failures of those who have come before you, and that to me shows a lack of real world experience.
And yes, I'm aware of several shops that can literally flip a switch, and we're talking cage and ASN, not hobby-scale iOS backends. It's not like finding the holy grail and requires developers who "get" operations. It does exist. And it exists in real life! \u1F632!
When it comes to in house DCs, very few companies seem to have that top talent.
For example, most companies are simply buying regular Dell/HP servers, Cisco switches/routers, slapping them in cabinets and calling it a day. They simply do not know how to take advantage of high density platforms like open compute, or of SDN. They also throw millions at software vendors like Vmware/CA, instead of building their own provisioning or CPU/RAM/Disk aggregating solutions with open source or custom tools.
If they don't know how to do it, then the theoretical savings for certain companies, of bringing things in-house simply won't materialize. And then its better to throw money at AWS if your in-house "engineers" are also breaking things 10x more often.
In a company with a badly run DC, developers very quickly latch on to cloud benefits like not having to wait for 3 days for DNS request, and 2 weeks for a VM.
At this point, with datacenters having been in wide production for the last 20 years, there should be plenty of 7+ year veterans.
(though yes, I have known in-house service providers that couldn't set up a new VM to save their life)
I've seen innards of very large DCs for more than a decade. At one of my first jobs at a fortune 500, I was responsible for everything from rack and stack to the command prompt. The expected turnaround time for a single physical server was 4-6 weeks until the application could be installed on it! One of the reasons was that they did not have automated DHCP/PXE provisioning. I started the process to enable it.. going through all the political, security mess, it was 9 months until it was enabled. I was gone by then.
An extreme example for sure, but if AWS revenues are growing like they are, then surely such issues are everywhere to some extent.
I don't understand this statement. I have built systems that work with openstack and then scale up to public cloud. The primitive is the same as EC2, which is a virtual machine. What did you have difficulty with?
Brilliant.
Dropbox built out their own environment (and did it migrating 500PB out of S3)
Like most discussions in this thread, this statement is way too general. Dropbox moved their storage from S3. What about EC2 or other AWS services they were using? Did they abandon all of those too?