Lost your luggage? That's nothing – we just lost your whole flight
theregister.com
theregister.com
Finally a supervisor figured it out. Somebody had cancelled the flight, moved everybody to other flights, and then killed the flight record. For some reason, my ticket was still booked on a now non-existent flight, and that confused the hell out of the system and made it give weird info (such as the flight was in the hanger, or that my ticket didn't exist). The supervisor had to re-create the flight, move my ticket off it, and then delete the flight again.
We had no development or test environment so development were done in production. The system lacked the necessary functions to correct errors so there were frequent updates/corrections made by writing SQL-queries directly towards the production database. The incident was when I should update a single record so the system would resend it to the customer, but ran the query before writing the WHERE-clause.
Luckily there were working backups (incremental backups that ran every 15 min, almost like someone knew they might need them!) and no one had made any changes since the last one. We did a quick restore and let out a sigh of relief.
We did eventually write a tool to do those updates (and deletes) we were doing multiple times a day and were able to hand that tool over to one of the users, saving us multiple five minute interruptions each day and boosting our productivity.
Worth noting that this was part of a business that we bought from another company and we did transfer all customers to our own product with dev environments and no need to manually updating databases.
Would have been good management...had they not decided to replace said senior with the new hire junior DBA.
In MEDITECH 6.x, your LIVE system instance was blue. TEST environment was green, and any management interface or screen on a server console that was RED was "this will fuck up LIVE" including giant
** WARNING ** LIVE ** WARNING ** LIVE ** MEDITECH
warnings in the title bar for the window..
The paradox is that every device I use interrupts me non-stop with modal dialogs ad nauseam.
But the production delete? You said,"drop", we said, "gone".
Either production database deletes need to adopt alert fatigue or the ops are already so fatigued that it doesn't matter.
Also frankly I think most software does it wrong, if it is possible it shouldn't ask before delete but just allow to undo the step easily. Then work can go fast and uninterrupted while still being resilient to mistakes. Obviously not possible everywhere but where it is it should be the default.
Obviously. When undoing is not easy or possible automatically, there is the "Are you really really sure?" dialog.
Pop-up warnings that are too common are a problem itself. People get so used to accepting them immediately that they stop serving their purpose and just become an annoyance. With one exception: they can prevent misclicks.
https://en.wikipedia.org/wiki/Alarm_fatigue
The misclick point is beyond truth, particularly given the storm of marketing that is mobile web.
Even fatigued, "DELETE 1182900" tends to stick out, if you're used to, and expecting to, see "DELETE 1".
("BEGIN" appears to be PostgreSQLism, "START TRANSACTION" is more standard.)
I consider myself an "experienced" developer, yet on the first day on my current gig I dropped a '--recursive' flag in an AWS command and deleted a bunch of stuff I shouldn't have. ¯\_(ツ)_/¯
A simple --max-really-delete added to a command line is subject to "replay attacks" on yourself when you use CTRL+R.
Even though I take all care to differenciate environments, always cross check, we had to live drop a table while on a meeting, and due to ongoing conversation the wrong window was selected, and multitasking ended up causing a major distraction.
The reason? Well, the admin interface was written in ASP. According to the guy that called me in panic said that ASP version was so very old it was extremely insecure. I didn't have access to the database so could not verify anything but the webserver logs had a lot of "Bobby tables" in them so I have a strong suspicion it wasn't all ASP's fault...
Anyway, we ended up with the database disappearing in the night or containing random data until I got that phonecall asking why we were linking directly to the admin interface :-)
They were very good at restoring the database though, just took an hour or so every morning until they found the reason.
Never had another incident..
"No worries, we have backups", I thought. Turns out those were more like replication, copying over every change immediately (including deletes).
In the end, fixed with a short script that cross-referenced tables and checked for missing relationships, but terrifying nonetheless.
Open office. Boss rolled over on his chair and started saying "so..." and I dismissed him with a "deleted all users leave me alone". He rolled away, kinda like Homer in that GIF where he disappears into the hedge. Luckily it was back up in minutes, but a new phobia was unlocked.
I think I'm going to start unplugging my desk phone when one of those moments happen to me. "Hey, there's something wrong with..." "YES I KNOW I WAS FIXING IT BEFORE YOU INTERRUPTED ME"
We ran the booking websites for lots of airlines (Which was really just a frontend that connected to Sabre or Amadeus). Airline data standards were bad. They're probably still bad. You were lucky if you could get an airline's upcoming schedule or pricing rules in some format other than "manually typed into an email by a human". We had people who's job was solely just to take these emails and put the info in the database.
We had a bunch of arbitrary safe-guards against data entry mistakes. Ticket prices above 10k or under 10 were flagged as suspicious and needed someone more senior to approve them.
So this airline, headquartered somewhere with a 7 hour time difference from us, wants to do a 9 euro flight special for some holiday. There's a big marketing campaign, apparently lots of excited customers, etc.
I don't know if you've connected the dots yet but 9 euros was less than our suspicious-data-threshold of 10. So about 95% of their available flights just disappeared. Searching for them would turn up a "no results" page. The only available flights were expensive ones excluded from the sale, usually codeshares with a larger airline for a more distant destination.
The airline never notified us. It wasn't until we waltzed into the office 7 hours later that any one of us saw the problem. One of the DBAs wrote a quick "approve all suspicious entries" query and that was the end of it. Thankfully it was a 3 day sale.
Afterwards, I remember the chosen method to prevent this from happening again was to change the low threshold for a suspicious price to 8.99 or below.
That doesn't sound like they deleted a flight, how often do you hit a flight that's currently interesting to airports around the world? Of course there are exceptions, e.g. Emirates would route you via Dubai, so a flight Dubai to Frankfurt could have connecting passengers from a few dozen different originating countries, but not "every airport around the world".
OK, come to think of it, maybe the inconsistency in the DB halted everything, e.g. the system saw a few hundred passengers booked on flight with ID 463523 but there's no record of a flight with that ID...
I wonder what the 2023 equivalent is.
I never bothered writing about the person putting tipex on the screen to delete something!! (true)
or the time another person swore blind that they've not touched the keyboard - dispite all the keys being in the wrong order (they took the keys off to clean because they spilt coffee on it and put them back in the wrong order)
or the time I nearly wiped out 4 man years of code.....
That reminds me when one user pissed one of our then-senior admins so they reordered his keyboard keys alphabetically.
Or the fun we had where in old system user needed to come to admins to type their first password (that then they are supposed to change after) on account creation and one of admins had blank keycaps keyboard. Reaction was anywhere between "didn't even notice" to "bewilderment and panic"
Been there, done that.
> and one of admins had blank keycaps keyboard
That's just plain evil, even at BOFH scale
"With Dilbert being so popular, I get a lot of email anecdotes. Once, maybe twice a week, I get a story about a customer support tech who found a user using a CD tray as a coffee holder.
You know what functionality I'd like my email software to have? The ability to gather all those people who claim that really, truly, it happened to them personally, lock them in a room together, and make them fight it out until there really is only one of them."
Why not make it a team effort?
I tasked them with monitoring and dealing with all the non prod environments. You have to realize in a big corp there are an infinite number of ways for things to fall apart. Much of it beyond your control. My goal was to minimize the blast radius.
One time I took vacation for the first time in years. When I came back I was blamed for an outage while I was on holiday. Their reasoning was that they didn’t push any new code into production so it must be the database’s fault. After a day of troubleshooting, I learned the application code was not using a framework like Microsoft Linq to read the SQL output but was reading it raw. I also learned that the Windows admin had been pushing Windows updates into the servers. The issue was that a Windows update (not a SQL Server update) changed how timestamps were recorded. It was beyond stupid. Instead of yyyymmdd hhmmss like the application expected. It was now yyyymmdd hhmmss timezoneoffset +- hh. They now had some weird hour offset and the actual values changed. To fix the issue the windows update was rolled back. And the Windows admins were told never to touch the database servers without my permission. And the application was rewritten.
rm -rf vv\ *