Django zero-downtime migrations for Postgres that respect database locks
github.com
github.com
Very useful, especially for long migrations, but I would like to see a bit more detail about how this library achieves this, what the caveats are, etc.
EDIT: Never mind, there's a comprehensive "how it works" section farther down, I just needed to scroll far enough. This is very useful.
- it doesn't use transactions, so if migration will down, then you will need to fix state manually (one point of improvement), however you cannot run `CREATE INDEX CONCURRENTLY` in transaction.
- it can use `CHECK (column IS NOT NULL)` constraint for safe `NOT NULL` replacement for huge tables, that isn't compatible with standard django behavior.
So all this cases highlighted in README.
There are definitely some wonderful ways to mess things up when starting out with postgres migrations. Nothing quite like the surprise you get the first time you rewrite every tuple on a table with 20M rows because you added a column with a default value (no longer an issue with the latest postgres).
As pointed out in the docs, your code must be prepared to support the schema both before and after the migration. Not that I use django anyway, but normally I'm more worried about the interlacing of the code and db changes to keep things running smoothy. Migrate a bit, release some code, migrate some more etc.
Another alternative would be to use postgres savepoints, which are like transactions inside transactions, as a wrapper around each migration. You can do the same thing - set lock_timeout and catch errors when those values are exceeded, and try the transaction again.
Provide an option to run the handful of operations that can't be run inside a transaction as an escape hatch, and then you can retain the ability to run migrations inside transactions, which is usually a good thing.
Here are the "implicit commit" statements:
https://dev.mysql.com/doc/refman/5.7/en/implicit-commit.html
It's predicated on the fact that you don't use foreign keys. Now why would someone use MySQL without FKs... is beyond me, but I'm sure they have their reasons.
MySQL has this https://dev.mysql.com/doc/refman/5.6/en/innodb-online-ddl.ht... for Online Schema migration and
MariaDB has ALTER ONLINE https://mariadb.com/kb/en/library/alter-table/
They both work with triggers + ghost tables, so they don't need transactions.
If you want to have referential integrity there's no other choice. Otherwise you create tables that have no relation to each other. No one stops you from doing this, but then you probably don't want an RDBMS in the first place.
It's madness, I know.
1. https://api.rubyonrails.org/classes/ActiveRecord/ConnectionA...
Thanks for that tidbit, I have no XP with Rails, but I would never imagine they do something so atrocious.
If you have constraints that can be enforced by the DB, you simply use the DB's constraints, because the DB guarantees they're gonna work 100% of the times and your data will be correct.
If they are more dynamic or require custom business logic... well you do it on the application layer. That's what everyone does.
I mean you can probably implement your own transactions, that doesn't mean that you should. And if you do, then just admit that there's no point in using an off-the-shelf RDBMS.
This has some benefits - it allows for a uniform definition of relationships regardless of database backend, allows for constraints that can’t be expressed by the database itself, and allows constraints to be used as a first-class concept for things like presenting error messages to users. But it also means that data integrity is not guaranteed - modifying records concurrently or outside of the application can result in a data model that’s valid according to the database schema but not according to the application model.
FWIW I exclusively use Postgres as my Rails database backend these days, and foreign key relationships are extremely easy to include in migrations. This still requires a companion definition in application code so that good error messages can be presented, but that seems acceptable to me. I’d hope that this eventually becomes the default for databases that support these keys.
I am just that crazy person that believes that default options should err on the safety/integrity/consistency side.
I also love abstractions, but abstractions can't change their underlying fundamental reality. So I would like to be the one who makes the compromises and I'd like those compromises to be explicit rather than implicit.
Some of that is compatibility reasons; not every DB Django supports would allow creating all the desired types of constraints in the DB, so Django has no choice but to ship application-layer enforcement.
You can work around that with manually-created migrations (and I've worked some places that did this, in order to get particular constraint types that were needed at the DB level), but I don't know of a good generic solution to that problem.
Django 2.2 is going with a somewhat-manual approach to allow specifying richer types of constraints, so maybe that'll lead to progress.
rails g model Post user:references
Then this line is generated `add_reference :things, :this, foreign_key: true`
and then we call the migration, and there is the foreign key.
By default the relationship constructors do not create foreign keys, if they did it would be redundant to include it in the generated output surely? At least that's what the docs seem to say.
Here are some of Github's reasons for not using them: https://github.com/github/gh-ost/issues/331
I think this is pretty normal for heavy MySQL users, honestly. When I was at Eventbrite we didn't use them for the same reasons.
Deciding to implement constraints on the application/business layer is something that github can probably do, in the same vein that Facebook can create their own PHP compiler.
I just think that you can't use these examples as effective arguments on whether someone should use FKs or not. I guess whatever brute-force solutions these companies would apply, on any given problem, would probably work one way or the other.
It's a bit tricky as you cannot use it in a transaction though
We use it by changing the models in code and then having Alembic automatically create a suitable migration for us. Then we change the migration as required (adding steps for data migrations etc). It's really simple:
alembic revision --autogenerate -m 'changing something...'
That generates a python file with an upgrade and downgrade function. You can then customise fully as you require.Then on the db server you can run:
alembic upgrade head # or downgrade etc