This was scary.
This was scary.
But I give them a pass because they are a young company. My company was similarly reckless early on, but as we scaled, we had to tighten things up and turning to an immutable deployment approach has saved our asses so many times.
Two days later the CTO sends everyone a stroppy email about "column bloat that should've been a table", ssh's into the personal instance that they've been keeping alive† since before you had funding and learned to launch servers as immutable black boxes, and whilst trying to prove a point by rolling it back manually, drops all tables by mistake when a cat treads on the keyboard
--
† excuse: "it's for reporting"
What happens if something goes really wrong after the production deploy? Is there a way to skip steps if you need to quickly push an emergency fix?
Being a young company doesn't mean you ignore all the mistakes other people have made and figure them out for yourself.
I really surprised someone has access to the prod DB, and that it's possible for them to connect to it in dev (Meaning they have a copy of the credentials???).
* Production should be immutable
* No one doing dev in a dev environment should have such trivial access to prod
* Are there still good reasons for a migration to drop all tables? I guess it's for the dev environment to etch-a-sketch to a known state?
Yikes.
It’s the new and hip ‘cloud’! Probably using planetscale or something like that, which (last I checked, maybe it changed but wasn’t on), doesn’t even have ip protections outside the mysql user settings (while bad, would’ve protected them).
> Are there still good reasons for a migration to drop all tables?
We haven’t found any.
ofc i'd think differently if i was also putting write-permission prod credentials into my machine, but luckily i haven't been in many places doing that
1) PlanetScale has IP ACLs, which locks down passwords to specific IP addresses. [1] Additionally, with TailScale or another VPN solution, locking down based on IP isn't necessary foolproof.
2) They also have Safe Migrations. When enabled, it prevents DDL from being run directly on a database. [2] Additionally, using deploy requests for zero-downtime schema migrations also allows you to use reverts, which will revert the migration. [3]
[1] https://planetscale.com/blog/introducing-ip-restrictions
[2] https://planetscale.com/docs/concepts/safe-migrations
[3] https://planetscale.com/blog/behind-the-scenes-how-schema-re...
We can restore a 10TB disk in about 12 minutes. its much faster to snapshot, do migration, then if necessary, drop disk and remake from snapshot. (and then replay the replay any other WAL changes up to the exact second you want with a tool like barman, wall-e, pg_backreset, etc.
Postgresql backups are critical for disaster recovery, but the restores are so very, very slow, they should be a last resort.
Today if a developer can bring down the operation accidentally, that’s a problem with the org more than the developer.
(On the other hand if a developer screws up the shared dev environment, it is his or her fault and they deserve the wrath of their coworkers.)