These guys avoid the problem by never rolling back the database, and never making changes that might require that.
These guys avoid the problem by never rolling back the database, and never making changes that might require that.
As for reducing the risk of damage to the data itself from badly behaved application code, I think the best approach is to design your architecture such that you can't lose important data.
There are various other techniques than can help. I wrote about some of them here http://benjiweber.co.uk/blog/2015/03/21/minimising-the-risk-...
Releasing less frequently actually only makes the problem worse. Infrequent changes are often too big to have a chance of understanding their potential affect on the production system. You're also less likely to immediately know how to respond to a problem.
I would not roll out new code with migrations without ad-hoc db backup. Any new code as well.
1. Change your code to write to old and new schema; keep reading from old schema.
2. Migrate data to new schema in background.
3. Add a flag to control where you're reading from. Default it to read from old location. Keep writing to both.
4. Flip the flag on some subset of your jobs. Ensure everything is still running smoothly for as long as you like.
5. Change the default to read from new schema. Wait as long as you like to be comfortable that the change is working properly.
6. Delete the code that reads from the old schema.
7. Delete the code that writes to the old schema.
8. Drop the data in the old schema.
At any point prior to deleting the old data, if you encounter problems you can roll back to an old version of the code. If your schema changes are incompatible, you can make an entire new database with the new schema. This may temporarily waste some storage space, but it's very safe.