Why We Chose Kubernetes
code.haleby.se
code.haleby.se
The first issue is that I've probably set up 10 or 15 app development and deployment "systems" if you will. I've found that it's very beneficial to automate the simple stuff but it quickly reaches a point of diminishing returns. A super-custom system always works great for a while until some big change or library upgrade or refactor or whatever comes down the pipe. Then we spend a ton of time resetting up everything. We have to keep the build system components up to date so it doesn't turn into an ancient mystery box. Sometimes an upgrade breaks the whole thing and then we're on stack exchange all day debugging a parser library or some other thing that we don't care about. Basically spending hours and days and weeks on the build system so we can have that sweet one-click (or fully automated) deploy.
The other thing is that we release frequently but we tend to double check everything before it goes to production. Our staging server is auto-deployed except DB changes which we do manually. Right now it's about 2-3 clicks for us to deploy to production and it works fine. We still do DB changes manually though. It takes a minute or two to deploy. I feel like the process encourages that final check that everything is cool.
I guess I'm nervous to set up something that deploys to production simply by adding a tag to a slack message or the git commit message. Should I get over myself? If I change my thinking is it possible that deployment to prod could be a non-event?
Sometimes adding some friction to the deploy process can be good IMHO. Continual deployment isn't good for every product or every team.
If a typical 2-3 click deploy generously takes an hour, and they do 40 per year... then it would take 1 year to break even presuming that a fully-automated system could be built and deployed in 1 man-week. If the deploy takes 10 mins, can it be built in less than 1 day?
Ignoring the development time for a fully automated system, I think the real question is, "how does a rollback and unscheduled downtime impact the ROI due to unforeseen problems?" because it will happen, eventually.
Automation is not just for saving time. It is also preventing the system from human errors.
People make errors due to fatigue, inattention etc. performing even simple tasks.
For instance, in a networked file system, the on/off switch for “move from version A to version B” ought to be about as simple as swapping a directory symbolic link. That way, everyone can see exactly what it is pointing to now, what it used to point to (as old version targets are probably in the same parent directory), and anyone can figure out how to roll it back instantly.
Given trivial on/off switches, the details of the rest of the system can start to gain complexity.
Also, it helps a lot to have something like a “beta flow” that is essentially a parallel replica of your production environment.
Google Cloud, on the other hand, was truly easier to use. Redeployments didn't take forever, and it wouldn't try to fail over and over for minutes before returning an error like AWS.
Disclaimer: I work on Compute Engine, but didn't work on GCR.
And for public images, it would be nice if Google would foot the bandwidth tab :-).
Having used both AWS and GCE, I find GCE is just a better experience.
I love the cloud console and the cloud shell. The CLI tools work well. VMs start quickly.
For development, preemptible VMs are an incredible bargain.
This is slightly unfair as the setup and configuration of kubernetes on your own kit is fairly difficult, at least it was the last time I looked, especially the networking side of things.
(chromium, linux)
Currently you can run it on AWS, OpenStack or vSphere; Azure and GCE support are being worked in concert with Microsoft and Google respectively.
Disclaimer: I work for Pivotal, the company which donates the largest chunk of engineering effort to Cloud Foundry.
I wish more people would spend the time to look into it instead of just defaulting to Docker + tooling, just because.
I'll definitely have a look at it.
FD: I work on Cloud Foundry
But IBM also has such an interest, so too HP and Fujitsu and NTT and Anynines and Intel and other members of the Foundation I've rudely neglected.
In a given team you can often find engineers from multiple companies. Or you'll find that different companies will provide a team. For example, the `cf` CLI was previously a Pivotal team, then a mixed Pivotal-IBM team, then an IBM team, now I believe Fujitsu are assigning engineers.
Voting rights in the Foundation are proportional to contributions, so it's in the interest of each participant to be generous in providing people, resources and IP. A defined high-level process, modelled after Pivotal's, is used to ensure a common approach to working on the platform -- pair programming, TDD, product management from a Tracker backlog and so forth. Every engineer goes through the same "dojo" onboarding.
I don't think anyone has ever tried anything quite like this before. At the corporate level it's press releases at twenty paces, but in the trenches it's a lot more seamless.
Did it?
For simplicity, http://deis.io should be preferred for modern "deploy-by-buildpack" architectures.
Strictly, BOSH solved it. But BOSH was originally written for Cloud Foundry, so through the fuzzy lens of distant history it looks the same.
In Pivotal we upgrade our public-facing Cloud Foundry installation, Pivotal Web Services, within a day or two of a new CF release being blessed. Typically nobody ever notices. Hundreds to thousands of VMs (I don't know how many we are running now) are upgraded to the most recent version of Cloud Foundry in a few hours and nobody notices.
Yeah. It's solved.
Recently installed a three vm cluster in Vagrant (actually super simple, some things actually have improved a LOT in 2016) but I still need to understand a lot it seems to get rolling upgrades etc.
https://github.com/kelseyhightower/coreos-kubernetes-talk
That might not be the right repo, though. Somewhere he has complete walkthroughs, including command by command, but I'm not sure if that's the right one.
Regarding the bug in Mesos found by Aphyr, that has been fixed: https://issues.apache.org/jira/browse/MESOS-3280
The internet should thank people like him for finding these issues.
Say I already use ECS or Kubernetes for orchestration of some of my services. If I'm working on a new Rails/Node/Python app that doesn't talk to the rest of my services, would it make sense to stick the app into my existing cluster? If not, what would be an easy way to launch, deploy, and manage (non-PaaS) these kind of stand-alone services?
Disclaimer: I work on Compute Engine (and I've always just used a single cluster)
Not sure if it's better to use larger clusters for unrelated stuff or this though.
I highly recommend you check it out before going down the long winding road that is kubernetes.
You'll actually be able to use Rancher to manage Kubernetes clusters[0] in the next few weeks, if you want the best of both worlds.
[0] http://rancher.com/introducing-kubernetes-environments-in-ra...
So they didn't actually "choose" Kubernetes. They just chose to use something that is run by someone else.