Scalable and resilient Django with Kubernetes
harishnarayanan.org
harishnarayanan.org
I've run django and other web deployments with simple shell scripts and occasionally some python to glue it together.
Most recently I'm running a django web server and some other custom stuff (a real time stateful server, database, nginx, etc).
A somewhat complex setup (or at least as complex as it needs to be), and I have no need of docker, kubernetes. If I want a new server, I just change a few parameters and run the script to deploy. It's straight forward and without magic.
The deployment flavor of the month before were DSLs like ansible, puppet, etc. Did these solve real problems? I found they added complexity without adding much. How is the new generation different?
Caveat: I'm not google, and I don't deploy huge server farms. There is a use case but the vast majority of us aren't that.
I echo exactly what you say in a giant caveat way up top in the piece. :)
Most of the first half of the article beyond that point basically tries to motivate why you'd want to try this beyond just using a classical VM approach. But the basic idea is that it raises your level of abstraction from working with machines to working with your application components (on abstracted hardware).
The problem I have is that most people feel like theyre doing 'serious work' so they need a serious solution when the reality is that it really is just a small percentage of companies - not usually even successful startups, for instance - that require this sort of automation.
Mostly, it adds complexity rather than reduces it.
I know that sounds like we're getting into the weeds, but if you consider that most companies are taking on additional complexity because of this, it becomes a problem.
Containers solve this problem nicely.
And once you're going down the container route, you need something like a Kubernetes to easily schedule and network them.
My personal site -- the one you're reading the article on -- is literally just a static nginx server on an old Linux box (with the files copied in by rsync). Even the curmudgeon in me now thinks it makes sense to use solutions of different complexity for different situations.
Don't think of Kubernetes as a replacement for Puppet. Think of it as distributed init/supervisord with service discovery, which you are clumsily building with Puppet. Your automation ends up writing init, units, monit, supervisor, and so on, right? (You shouldn't be deploying apps with Puppet anyway because CD is not CM, but nothing like Kubernetes or Mesos really existed in the open source world until now so a whole generation learned to do it with Puppet.)
Kubernetes exists for when your scripted machine fails or you exceed the capacity of your scripted machine and need to horizontally scale. It is not a deployment tool per se; that's only part of what it does as a scheduler and resource manager. Every project, from personal to Fortune 500, needs to plan for those scenarios. I'm not saying Kubernetes is always the answer, I'm saying it solves a problem that every project universally has, despite your claims. The bonus of working that way is now you have an API to your machines and can treat them as a single unit of resources. Kubernetes is about half of building your own mini Heroku.
You've heard snowflakes versus cattle, right? Your way is for snowflakes. Kubernetes is for cattle. The threshold where one becomes more productive than the other is the constant debate, and as an SRE, my opinion bucks must people who talk about this on where that threshold lies. If you self-identify as a sysadmin you'll likely have one opinion, devops another, and then you have my group of crazies that generally want to crush all snowflakes. I can speak with experience that our crazy SRE ways scale pretty well to the hundreds of thousands of nodes case but also work quite well for toy systems.
Subscribe to everything CoreOS is doing. There's a spectrum of quality there, but they're running with the programmable infrastructure ideal and have largely bet the company on the Kubernetes ecosystem.
[0]: Kubernetes makes some interesting choices versus Borg and I think they're going to have trouble scaling it to that, and Borg really shines at Google because of the global filesystem layer that exists on every machine (they don't realize this and are doing gymnastics around storage in the open source side), but if they can mature Kubernetes it'll be a solid building block for platform development and eliminate a lot of the clumsy scripts and automation that you are defending and push us toward programmable infrastructure.
I already have AMIs, autoscaling groups, health checks, ELBs, etc. Why am I adding another unnecessary layer of abstraction?
Disclaimer: Devops/sysadmin who would like to get off the hype train and get real work done.
People don't realize how deep Amazon's hooks are at scale. We were dropping half a million USD a month at that successful startup I mentioned and could do the same work with four or five racks of hardware, storage included. I bet you could cut your opex in half on something else, and it's a safe bet because (a) I've seen it be resoundingly true even coming from three year reservations and (b) half is actually a conservative estimate. I am floored by how much money the industry throws at the fundamental inefficiency of multitenant virtualization atop AWS.
What a strange argument, considering Amazon a primitive. That's flirting with vendor koolaid. Physical and DigitalOcean/cheapo deployments are an immediate counterexample, too. Hell, if network latency is low enough, you can straddle both and schedule Kubernetes pods on whichever provider is cheaper on a given day. Spot instances, too. Lots of possibilities.
Note I said "AWS primitives", not AWS as a primitive. I think they're overly expensive unless you're using them for prototyping, as you should eventually move off to your own gear once you know what your load/compute/storage profiles look like.
I appreciate you pointing out that this is most valuable when you're _not_ running in AWS (I work at a startup that wants to use this tech in AWS, hence my not understanding why we'd waste additional abstraction for little additional benefit).
Disclaimer: I _want_ to move off of AWS to our own physical gear at some point, but our dev team is married to it because they're afraid of physical infrastructure (but its okay when critical services like EC2, autoscaling, and IAM aren't working for several hours apparently).
> but what about when you want to move to GCE to arbitrage price differences for certain services?
Already doing so without additional (unnecessary?) tooling, like Docker or Kubernetes (build images at respective cloud providers, use existing orchestration tools to spin up or terminate spot (AWS)/preemptible (GCE) instances).
It's a step up from puppet for sure but I don't really see the benefits of using it over ansible.
But also, you can use the same cluster to run your preproduction enviroment, or use it for CI/CD. (Check deis) As an example, we have an small cluster with two machines and on it, there are:
- the main app, - the preprod environment, - two more feature branches (that had to be reviewed) - And also commercials can deploy playground environments to make demos and trials.
All are independent apps, sharing resources, on the cluster.
As a summary, on daily, we maintain 7 to 10 independent apps instances, and we do regular updates (as new revisions arrives). As an example, the trials or feature branches, are single pods (all included, db, redis, app and worker).
Kubernetes is not for deploying pet projects, but as soon as you start working on a real project is a must. You can manage complex deploys with a lot of services.. keeping the costs as lows as two n1-standard-1 machines. And as a plus on gke, you got monitoring, and logging aggregation.
If the same (deployment and upgrades for production, staging and testing environments, for any number of hosts) could be achieved with a small shell script - doesn't this mean that cluster thing isn't really useful?
I think a lot of "real" projects (smaller ones) work perfectly well without any complicated cluster management tools, and aren't really hindered by lack of what those offer.
My guess is that 2-3 years from now it will be easier to set up a kubernetes cluster than to set up even the single VM at the top of your piece, and a corresponding dev env will run seamlessly on Windows, Linux, OSX. The average user won't even know or care what OS they're running in production; it'll just work (think Heroku++)
Of course once it becomes ubiquitous, then containers become the new (language agnostic!) way to deploy "libraries", accessed by HTTP, and DLL hell is just moved up an abstraction level. So what have we really achieved....
1) it usually doesn't have sshd
2) you change would be thrown away at nearest container restart or migration.
The only way for it to persist is to do the right thing, i.e. apply the change where you should and have the container rebuilt.
Disclosure: I work at Google on Kubernetes.
They're a bit more verbose than `scp ./files/nginx.conf $HOST:/etc/nginx/` or `ssh $HOST sed -ie 's/.../.../ file.conf' for most basic stuff, but way less verbose for anything even a bit more complicated.
I don't use Docker or Kubernetes, so no idea about those. But I think configuration management tools can save time. But their usefulness is debatable if, say, all you need to deploy is to scp and untar a simple archive then poke init system to spawn or reload a daemon.
- How do you make sure the version of Django on your machine matches the one in production? - What if you do brew update on your Mac and the package for your repo in production is out of date? - How do you roll out a new version of your app (even allowing for downtime) and make sure it did it "completely"? - How do you make sure your app keeps running even if your VM goes down (let's say you need to do a kernel upgrade)? - How do you use one password for your local development database and one for your production database (and not hard wire it into code)?
These are super common problems - you're going to have to solve them no matter what. Lots of folks do this with scripts, but one small change can cause all kinds of annoyances.
This is why containers have taken off, and all of these are addressed via the combination of containers & Kubernetes. It's a different way of thinking about development, but I've been able to accelerate my velocity substantially once I adopted them. I'd argue that the complexity that you accept as 'normal' for a small site is an unnecessary tax, but I'm biased :)
Disclosure: I work at Google on Kubernetes.
To check for versions of Django, and other libraries, you write simple scripts to test for them. Run them in unit tests, and as a pre-check stage in deployment scripts.
Our apps do keep running if any single VM goes down. We don't need Kubernetes or anything to handle that. Just a load balancer and an ElasticSearch cluster.
Dev hosts obviously have different passwords read from config files only present in development systems.
However, in practice, if you're using AWS or GCloud this is usually a bad idea - just use the managed database solutions provided. They have things like backups, snapshots, restores, upgrades, HA, monitoring, and alerting baked in. These are non-trivial to do yourself.
https://cloud.google.com/python/django/appengine
However there are a lot of tricky issues with Django on App Engine so I would also look into Djangae:
https://github.com/potatolondon/djangae
As far as Kubernetes, the OP and I had some independent discovery and started working on similar material at the same time. I also have some Django/Kubernetes content, both CloudSQL (MySQL) and Postgres in Kubernetes:
https://cloud.google.com/python/django/container-engine (CloudSQL)
https://medium.com/google-cloud/deploying-django-postgres-re... (Postgres)
My biggest reason for running the database in Kubernetes is a) CloudSQL doesn't currently support Postgres and b) it's interesting. Seems like OP had similar motivations.
Obvious disclosure: I work for GCP.
I cited your work as additional reading because I really enjoyed watching your talk on this recently! While I'd figured out a bulk of this stuff out, there were a few things I learnt from it too.
Definitely happy to have you contribute to the ecosystem. It seemed like a content gap that I wanted to fill in, but at Google we would of course prefer to have a healthy external ecosystem of people writing content like yourself!
I am personally experimenting with this within the context of containers because I'm trying to construct something like Vitess[1] from first principles as an intellectual exercise.
Running our own Postgres install costs roughly half what it would on AWS, takes very little engineering time to maintain, and can be tuned as needed.
When the costs go down and persistence-as-a-service is provided by more cloud hosts, it will make more sense but right now it often doesn't.
I wonder about opportunities to provide some of these services as a white label to hosting providers like digital ocean.
Is anyone doing that? What stops them?
If the service is easy enough to provide (this means not just ease of deployment but also up-time, maintenance, etc) while fitting into their margins without issue, it doesn't provide stickiness.
If the service is sufficiently complex to provide stickiness, they either need to charge for it, or it will likely eat heavily into their margins.
Neither of us know the specifics of Digital Ocean's business plan, but I strongly suspect that were it as simple as you put it, where it's hardly any cost to them, and a huge gain due to increased stickiness, they'd already be doing it, or be moving toward it. And they might be. However, if they aren't, then it's hardly a mystery as to why they're not providing such services.
[1] https://flynn.io
I often find that posts like this are a bit too meta and in trying to write things in a general way they leave out some critical step which makes replicating their ideas difficult.
I wanted to start from something I assumed people knew, and tried to motivate why one might want to improve on that. Only then did I introduce the new solution.
A lot of tutorials I found jumped too quickly to the 'how' without spending enough time on the 'why' I should care.
I'm starting to play around with Google Container service as you suggested, having a good time.
One is containerizations, where developers are responsible for maintaining fat application stacks that can easily be redeployed and moved around.
The other is towards serverlessness with things like [Django-Zappa](https://github.com/Miserlou/django-zappa), where scalability is handled automatically by cloud providers.
My bias is quite clear - deploying apps should be easy.