Incident report for February 21st, 2024
resend.com
resend.com
This is the second incident that can be characterized by a very hazy delineation between development and production environments. The first incident had to do with an attacker gaining access to private credentials due to devs leaving keys set in NEXT_PUBLIC environmental variables on their site.
>accidentally destroyed production database on first day of job
Same symptoms: while developing a feature locally, they accidentally pointed to the production DB, which was destroyed when running tests.
I remember as a child thinking adults had everything under control. That they know what they're doing. I guess I assumed a day would come when I too would know. That day never came. It's easy to think when you look at shiny websites and the first paragraph of this comment that other adults do know. But I'm often reminded of the truth: nobody knows. Everyone is always operating at least slightly outside their comfort zone or, in the case of the article, wildly.
This was scary.
But I give them a pass because they are a young company. My company was similarly reckless early on, but as we scaled, we had to tighten things up and turning to an immutable deployment approach has saved our asses so many times.
Being a young company doesn't mean you ignore all the mistakes other people have made and figure them out for yourself.
I really surprised someone has access to the prod DB, and that it's possible for them to connect to it in dev (Meaning they have a copy of the credentials???).
Two days later the CTO sends everyone a stroppy email about "column bloat that should've been a table", ssh's into the personal instance that they've been keeping alive† since before you had funding and learned to launch servers as immutable black boxes, and whilst trying to prove a point by rolling it back manually, drops all tables by mistake when a cat treads on the keyboard
--
† excuse: "it's for reporting"
What happens if something goes really wrong after the production deploy? Is there a way to skip steps if you need to quickly push an emergency fix?
* Production should be immutable
* No one doing dev in a dev environment should have such trivial access to prod
* Are there still good reasons for a migration to drop all tables? I guess it's for the dev environment to etch-a-sketch to a known state?
Yikes.
It’s the new and hip ‘cloud’! Probably using planetscale or something like that, which (last I checked, maybe it changed but wasn’t on), doesn’t even have ip protections outside the mysql user settings (while bad, would’ve protected them).
> Are there still good reasons for a migration to drop all tables?
We haven’t found any.
1) PlanetScale has IP ACLs, which locks down passwords to specific IP addresses. [1] Additionally, with TailScale or another VPN solution, locking down based on IP isn't necessary foolproof.
2) They also have Safe Migrations. When enabled, it prevents DDL from being run directly on a database. [2] Additionally, using deploy requests for zero-downtime schema migrations also allows you to use reverts, which will revert the migration. [3]
[1] https://planetscale.com/blog/introducing-ip-restrictions
[2] https://planetscale.com/docs/concepts/safe-migrations
[3] https://planetscale.com/blog/behind-the-scenes-how-schema-re...
ofc i'd think differently if i was also putting write-permission prod credentials into my machine, but luckily i haven't been in many places doing that
We can restore a 10TB disk in about 12 minutes. its much faster to snapshot, do migration, then if necessary, drop disk and remake from snapshot. (and then replay the replay any other WAL changes up to the exact second you want with a tool like barman, wall-e, pg_backreset, etc.
Postgresql backups are critical for disaster recovery, but the restores are so very, very slow, they should be a last resort.
Today if a developer can bring down the operation accidentally, that’s a problem with the org more than the developer.
(On the other hand if a developer screws up the shared dev environment, it is his or her fault and they deserve the wrath of their coworkers.)
Hours later they run some script that does DROP DATABASE from that same shell they used to troubleshoot, which takes a little longer than usual…
Anyway I can totally see it happening to me in my little one man shop, but now i think I might look at removing some privileges from my prod account :)
Does that make sense?
I mean, I know there are some bad practices out there - but connecting local dev environment to prod database server would be insane for any reason!
[0] https://resend.com/blog/incident-report-for-january-10-2024
Maybe this company doesn't need to exist, and shouldn't.
> I like that Prisma drops the database.
‘Like’ ; so if it didn’t but still worked perfectly, you would complain that it didn’t?
Yep I would, because it maintains the idempotency invariant, plus it flushes out any bugs that may exist if databases are not idempotent from seed data, that is the "why it does it."
The commands that suggest a reset are explicitly designed for use with development databases, not with production. The "deploy migrations to production database" command on the other hand only applies already existing migration files (that you have reviewed before) and does not suggest a reset, ever.
> While building a feature, we performed a database migration command locally, but it incorrectly pointed to the production environment instead, which dropped all tables in production.
Different color scheme on prod consoles.
Always use “BEGIN TRAN”.
But it's better not to have these shells open and available to accidentally choose them once my brain switches from ops-mode to dev-mode.
Luckily for me I'd been working on some other things related to it and had taken a backup not long prior, but it was pretty nerve wracking to hear the CTO/CEO nearly running through the halls to find out what had happened.
Pretty sure they never implemented more stringent access controls at that company, either.
But yeah, both the incident and the report are really tough to read. It would be great if they can do a follow-up with further actions they're taking.
There's a neo-bank called Revolut that allegedly at one point had just two teams: "go fast" and "don't screw it up". I feel like an infrastructure play needs some dedicated hires in camp 2.
What strikes me in the incident report they focus on failed migration, where the real issue is not planning or not testing for recovery if migration goes very wrong.
Even if backup recovery would take just 6 hours would it be acceptable ?
Aside from mitigating local dev accidentally pointing to the prod db, if you have the db accessible externally means it’s susceptible to network attacks and password attacks
Fortunately we had backups available.
Often this isn't a problem with the individual developer itself, but points to a problem with the organization. Frankly most developers shouldn't have access to a production database, let alone mutable access.
One major concern is loss of data, but another is privacy.
It's so frustrating to see this happening when there are tools that solve this like Snaplet (I'm a founder), and Replibyte that allow you to generate or obfuscate data for usage in dev-environments, and Neon that allows you to branch your database.