As if this was not enough, I didn't understand what was going on because tests that used to take 5 seconds were taking 3 minutes at that point, so I ran them at least 10 times to make sure this was not a pipeline glitch. This also helped me to make sure that the database was completely wiped.
I think Kubernetes has some nicer features that could make this easier but I haven't had a chance to try this out yet.
In fact, the CI environment and production environments are on different AWS accounts. This way if you accidentally execute tests with the wrong connection strings (and we have done this before), they just fail because they can't connect.
yep, that's one way to structurally enforce isolation (unless you go out of your way to configure cross account permissions)
there's probably a bit of a tradeoff between things being safe and isolated, and things being seamlessly automated.
I did think of a way to make this more secure but I haven't done it yet - you could write a script that waits for input from a website where you have to log in via SAML. The script then passes the AWS credentials for the role back to the deployment task. Effectively, it's a second step authentication process where you'd need to be authenticated using a security key for the deployment to proceed.
I haven't quite figured out all the details yet but that might help you get the assurance you need.
The next day I woke up and realized I forgot to enable password protection on the staging environment. I opened up my laptop, navigated to the dashboard, add the flag and reload the config. 5 seconds later, my phone starts beeping. Tons of alerts that the production website was down and responding with "403 access denied". Immediately tens of people jump on Slack "website down?!".
Turn out, with my sleepy head, I added the flag to the production website instead of the staging website. And thus enabled password protection on the live website. As soon as I noticed, I undid it and in the end, downtime was roughly a minute.
Suffice it to say, I got slapped for this fuck up.
Sounds worse than it is. Got to work in the afternoon, got pizza and took a couple of extra free days exchange.
It took 2 parameters, the first one is the IP address of the production database and the second one my IP address.
I guess it was inevitable for those to get mixed up, and I guess no one did bother to prevent write access to production.
Found my mistake months later, was able to fix it before disaster struck.
The system still booted, because the boot sector pointed to the sectors that the deleted file had occupied, and the contents of those sectors had not been over-written. I ran that way for two days until I got it to a point that I could restore the boot file. I spent those two days afraid that any step would over-write the wrong sectors and the system would be dead.
For a combination of reasons there was one order which was picked up every time the job ran.
This went on for a very long time and nobody in the company realised. The warehouse staff were all agency workers who dispatched B2B and B2C so sending a pallet wasn't unusual for them. It was only when the customer got in touch asking us to please stop sending the same thing that we were alerted.
Apparently she had ordered the items to her friend's office and they were getting very frustrated having to deal with them.
Even worse the customer was in China and the company had to order a shipping container to retrieve the stock.
Hyper-V live migration from a HP EVA4400 to a 3PAR Storage. 45 hosts migrated without a hitch. The one that couldn't fail, a SAP production server, failed hard. The EVA crashed hard, both controllers went offline in the middle of the migration. After a couple of seconds later, one of the PSU's shut off. The other one was waiting for replacement part. My face turned white. Huge downtime to recover everything from a Tape backup. A couple of days later, I had a major burn out.
It was a really bad day at work :(
Thankfully all but the last day's work was on Git, but I was using the laptop as a temporary photo storage (only device with an SD card reader at the time) and lost 9 months of important photos.
Yeah, I forgot the WHERE.
This was 15 years ago at my first job and I haven't done it since. I was lucky that we had backups.
An alternative would be for databases to add an option to prevent accidental deletion. When enabled, would make it impossible to truncate or delete entire tables. I would enable such a setting on my production databases.
At least for PostgreSQL, someone tried:
https://www.postgresql.org/message-id/12104.1150425319%40sss...
But it doesn't seem like it got much serious attention.
I think front-end makes sense, where SQL is a UI, not back-end, where it is an API.