Deploying a Django App with No Downtime
medium.com
medium.com
I arrived at this procedure: Spin up a dyno, worker, caches and warm everything up. Run smoketests to validate everything from the app-server onwards is lit and firing on all cylinders. Redirect nginx traffic to the new setup. If everything turns to shit in a second, reverse traffic change and re-evaluate. If everything works and nothing is broken after an hour, take the old dynos/workers/caches offline.
The trick was if there were migrations that broke the older version that you did it in 2 or more steps to isolate the changed model interactions.
And nobody should be using gunicorn, it's way less performant than uWSGI (citing:http://blog.kgriffs.com/2012/12/18/uwsgi-vs-gunicorn-vs-node...).
Don't forget to warm up Solr or ES as well.
uWSGI is a beast to learn but once you got the hang of it, nothing else will compare.
What I loved most about the course is that the presenters were extremely knowledgeable and gave great tips on how to manage your schema with ES and configure the clusters so that you don't encounter performance issues later on.
You can do some amazing things with ES, that you wouldn't necessary even know you could.
One example is the percolator[0], which can be used to slam new documents against existing queries and essentially classify the document based on search queries.
[0] https://www.elastic.co/guide/en/elasticsearch/reference/curr...
The way you solve the migration problem (much harder than it sounds) is to do only non backwards breaking migrations between a blue green switch. So you end up doing your migrations in multiple deploys (e.g add a field, and start using the new field with no references to the old field in the code, then remove the old field in the next blue/green deploy). I've generally found that it might not be worth the effort though depending on the context (especially if your initial healtcheck is before the migration).
And nobody should be using gunicorn, it's way
less performant than uWSGI
uWSGI may be different now, but 4 years ago I built a browser toolbar C&C API server that peaked at about 2k requests per second. We used uWSGI to front our Django application with SoftLayer HTTP load balancers in front of uWSGI. uWSGI was speaking HTTP and translating to WSGI, but did so only in one thread/process. So while 8-16 Django threads were tootling along, the single uWSGI HTTP->WSGI thread was pegged. We switched to gunicorn and had dramatically better performance.I know that it takes time and effort to change out software pieces that have a long history within an organization (I have some experience in that department) even when a clear and technically superior alternative is available.
I have a drawing of my favorite infrastructure that solves a ton of problems before I code one line of anything.
nginx -> uWSGI -> Python(Django) -> Celery(via RabbitMQ) -> Postgres
It's a bit more complicated than this and redis is missing but it's a nice linear representation of how I like to structure things. The only thing there not lightning-fast is Django but it doesn't need to be. I'm spending most of my time in the other pieces of software.
Sounds more like a crappy setup than a problem with uWSGI itself.
No. As I detailed, it was a deficiency with uWSGI itself. Complicating the setup as you suggest is a crappy workaround for a uWSGI deficiency. We simplified the setup by switching to gunicorn. I would still use uWSGI, but, as with all software/tools, I would be aware that it has a deficiency.uWSGI docs on production usage of HTTP functionality don't mention this issue so I wanted to let others know. https://uwsgi-docs.readthedocs.org/en/latest/HTTP.html?highl...
I have a drawing of my favorite infrastructure
Wish I had the luxury. We're running 20-30 projects/applications (~3000 tasks) across our clusters. Some are Django fronted by gunicorn, some are directly fronted by flask, but 80% are non-HTTP services.Here's a must-read if you use it: http://uwsgi-docs.readthedocs.org/en/latest/ThingsToKnow.htm...
It's feels pretty hacky to me though, so I've been recently reading up on Fabric and deployment patterns for Django, great timing with this article being posted!
In all other cases you're procedure seems only reasonable. I don't understand why anyone would do it differently.
I submitted a patch to gunicorn just this week about this: https://github.com/benoitc/gunicorn/issues/598.
Ideally I'd love to see all server software that listens on a port have a SO_REUSEPORT option for hot-swapping. This feature makes operations so much simpler.
Looking at this article is just seems to make ports operate either with any new server initialisation just automatically takes over the port (or maybe they become effectively load balanced between all running processes, it's hard to tell conclusively)...
http://freeprogrammersblog.vhex.net/post/linux-39-introdued-...
Can you work around this by instructing the first server to stop calling accept() before launching the second server? (But keep the first server's listening socket open, in case the second server doesn't start, so you can tell it to start calling accept() again.)
What I'd suggest is to add a "stop accepting" and "start accepting" signal, maybe a URL route that checks that it came from localhost or something and sets/clears a flag. If the flag is set, the mainloop skips all calls to accept().
So the full set up is, first the old process stops accepting, but keeps the listening socket open, so new connections queue in the kernel. Then you start the new process and make sure it works. Then you shut down the old process. If the new process doesn't work (e.g., accepts connections and returns 500s or something), then you've lost some requests, but the old process is still around. So you can signal it to start accepting again.
Maybe this is overengineering, but the case I'm worried about is where something in the state of the old process is important (it has an open database connection, it has a module loaded into memory that got corrupted/deleted on disk, etc.). So if you shut down the old process before you test the new one, you have no guarantee you can get the old process back up if you need it. You may have lost the node entirely.
https://uwsgi-docs.readthedocs.org/en/latest/articles/TheArt...
That said, if you boot a VM and it's bad, you are going to lose requests then too.
Second: smoketests should show if the engine is up and serving requests not under load.
Third: you have the fallback of reversing the flow.
You will always lose requests. All you can do is minimize the damage.
EDIT: Want to clarify what I mean by lose requests since someone is downvoting. In the process of administrating a site with that many requests and a moving infrastructure you can plan as much as you want and try to employ hedging strategies all you want. You will STILL run into problems once in a while that will drop messages. You don't have to lose requests to handing load-balancing off to another handler and you can reload apps without losing requests if the app server is properly coded.
Django Deployments Done Right by Peter Baumgartner: https://www.youtube.com/watch?v=SUczHTa7WmQ&list=PLE7tQUdRKc...
What I miss usually are two parts of deploys that are often ignored, and are critical in production environments:
1. Revert/rollback to older versions. We're humans and despite having all the processes sometimes murphy's law applies on moderate to complex server setups.
2. There is git/svn for code tracking, but even more important is the database consistency. Versioned backups and restores should be also part of the whole setup.
I'm currently looking into having the full setup with no downtime with open ears.
If there's a service/tool that automates a lot of this or makes it safer, I'd be really happy to learn about it!
Of course a database rollback may loose information from any newly created fields.
1) Add a Migration that adds the new field, but allows null. Deploy.
2) Add the field to the model. Also make sure it sets it's default on write. Deploy.
3) Execute a background task to set the field to whatever it's default value should be.
4) Add migrations to enforce integrity and add indexes. Deploy.
5) Actually deploy code that needs the new field.
Yes, this is super annoying, no, most people don't do this. Separating out #1 and #2 means you can always roll back all the way to right after #1 without losing any data. An extra, nullable field on a model with no indices on it shouldn't hurt anything.The problem is with restore. If there is no change in database it is more or less straight forward to download and import the db with the n-1 code tag. But if there are any inserts in database then a migration back should be applied to the latest stand of the database. I can imagine it's going to be a hefty workflow to do.
When deploying new version - backup the transactional db and keep a pointer to the last event in events db. Keep the events db running.
If restoring - restore transactional db backend based on backup and "play back" all events from the events db.
The distinction can also be made using one db and backing up / restoring subset of the tables.
In our own Django web app, we basically just use the load balancer for deploys. Our service provider (Linode) has a nice API, so when deploying we just do one server at a time and direct traffic away from the ones doing the upgrade. It's not complicated and works just fine... At least when there are just two machines.
It sends traffic to a new instance only after the healthchecks are passing. Makes for a really nice rolling restart workflow.
>kill -HUP <pid>
EDIT: That probably applies to nginx too.
nginx -s reload
also works, you don't need service/systemctl at all for this.However, I'd say that this gets a whole lot easier as you add even a "local" load balancer: that is a load balancer that lives on the same box but does not require a restart during a deploy. In this case you can start up your gunicorn/apache2/uwsgi/whatever on a new internal port, then remap which port is "live" using your firewall or by updating your load balancer's config and reloading it. This is how dokku works with Docker containers and I love that I have zero downtime deploys with it out of the box.
Also make sure you're on pip 7+ to take advantage of the automatic wheel cache to speed up building new virtualenvs.
As another commenter noted, I discussed a similar approach at DjangoCon US this year https://www.youtube.com/watch?v=SUczHTa7WmQ&t=1083
rsync -avz ./ user@remote:/home/user/web-app
ssh user@remote 'cd /home/user/web-app; \
venv/bin/pip install -r requirements.txt; \
venv/bin/alembic upgrade head; \
supervisorctl pid web-app | xargs kill -HUP'
The full deployment script has a few extra options which I've omitted for clarity, but this is basically it.The real way to perform clean update is using a load balancer, offloading instances for upgrade one at a time.
Also please people stop deploying using github. Package management & versioning IS a thing.
> rsync -avz ./ user@remote:/home/user/web-app
This is not an atomic update, meaning that you could end up loading modules from the new version from code of the older version. Use a new directory to upload the new files.
> venv/bin/pip install -r requirements.txt
You do not delete the old requirements, this is incrementally bloating the virtualenv. Just create a new one, virtualenv are designed to be created/removed quickly.
> venv/bin/alembic upgrade head
Huu, that's nice if you only have 1 instance to upgrade at a time. Otherwise it will just blow up.
#!/bin/sh
dest=/var/www/myproj
GIT_WORK_TREE=$dest git checkout -f
$dest/manage.py collectstatic --noinput
$dest/manage.py migrate
touch $dest/myproj/wsgi.py
Running a git push to the prod server fires this hook and that's all. I could improve it by first checking out to a test environment to run django tests and, if everything passes, do a git checkout to production.
Plus, am I the only one using apache?
As a side note: I've run copious tests against Nginx -> uWSGI setups and Gunicorn, etc. setups, and Nginx -> uWSGI is in a whole different league (way faster, way more requests per second, way less memory, and you CAN (despite what I'm reading herein), very EASILY with uWSGI, deploy without losing a single request). However, I still use Gunicorn, because it's quick and easy, it works well, and I don't have to sys admin the thing.
It makes me wish for some kind of "advanced postgres" guide.
While benchmarking this, I learned that with the default and safe settings (synchronous commit), PostgreSQL can only do few hundred write transactions per second, however simple. Set synchronous_commit=off in postgresql.conf and TPS goes into thousands.
As every transaction has to be durable at commit, fsync (well, fdatasync() if supported) of a sequential log is required. Most single drives can do a ~100 (rotational disks, normal speeds) to a few thousands fsyncs/sec. If you have a battery backed raid controller you can do tens to hundreds of thousands, since the data will not actually be written to disk.
Concurrency matters because several commits can be made durable with a single flush to disk if they're happening concurrently (sometimes called "group commit").
EDIT: spelling, grammar, minor clarification.
At least such cases should IMHO be considered before doing such upgrades.
I would indeed sleep better if the service monitored my cron jobs. But I would sleep even better if I knew you were making some money from this, because this gives you an incentive to keep providing the service.
For instance, you could offer a paid plan for users with >N/day requests or something like that...
Not sure if you want to go there or not (and I am not currently in the market for such a solution as I have solved my need with a proprietary solution), but the fact that there is no paid plan lowers my expectations of the service.
EDIT: clarified my point.
Or is he using Python 2 to deploy a Python 3 application?
For a single project running on a single server without many updates, you don't need virtualenv.
For multiple projects with various age running on multiple different distros and you have dev/staging/prod you might need virtualenv badly.
If you manage globally the dependencies you would probably brake the running version of the site during deploy, because it may expects different (old) API from his dependencies. For example the new code may NOT require a previously used dependency, and thus the deploy would remove it and the running code would collapse.
Also, if the deploy fail midway you'd like the global environment to be unaffected.
^/(\w{8}-\w{4}-\w{4}-\w{4}-\w{12})/?$
https://github.com/mutalyzer/ansible-role-mutalyzer
healthchecks.io looks pretty nice, I might start using it!
Even if you can guarantee no processes will restart in this period (maybe you pause your daemons and you've configured your webserver to not spawn any new workers for a while) you've got the issue of migrations.
Personally I stop the webserver before this sequence and restart it afterwards as I can't reliably reason about all the things that can go wrong in between.
The automatic reload behavior is a thing peculiar to the dev server. If you tried using it in production, especially on a large repo on a loaded server, chances are that Django's auto-reloader would notice that files have been changed as soon as git updates the first file, and it would reload that new file but the remainder of your old files. So some of your requests would go through to a splinched version of your website.
Of course, the reloader will figure out soon after that the last file has been changed, but if you were willing to tolerate a brief period of incorrect code / requests returning errors, you can tolerate actual downtime.
That's why this method, and all correct zero-downtime deploy methods, involve running a new Django server in a new directory, waiting for it to start, and then signalling the web server to route requests there.
> Aside: With regard to technology choices, the guiding principle I’ve been following is to keep the stack as simple as is feasible for as long as possible. Adding things, like load balancers, database replication, key value store, message queue and so on, would each have certain benefits. Then on the other hand, there would also be more stuff to be managed, monitored, and kept backed up. Also, for someone new to the project, it would take more time to figure out the “ins and outs” of the system and set up everything from scratch. I see it as a nifty challenge to stay with the simple, no-frills setup, while also not compromising performance or features.
> Let’s set some constraints: no load balancer (for now anyway).
But yeah, that would work otherwise.
That's not how you verify it. Use a load benchmark like "ab -c 10 -n 3000 -q domain.tld/" and if that doesn't report non-200 responses while deploying the new code, then you can talk about no downtime.
My solution for this problem (tested and deployed on Django sites) is this: https://github.com/stefantalpalaru/uwsgi_reload
Erlang lets you update live code, while it's running.
Still, I'm a big fan of Python and Django, and it's nice to see people realizing the value of "live" upgrades, even in a round-about way.