What we learned after I deleted the main production database by mistake
medium.com
medium.com
It's called Disaster Recovery and it has a best mate called Business Continuity. If you don't actually have a process for something then it will fail unconditionally, without you even knowing it is going to happen, until it does.
OK, it's easy for me to snipe but nowhere did I see terms like those. Backups are mentioned almost casually and there is this: "but no process was implemented for ElasticSearch databases".
So, not only were critical parts of the system not actually backed up but there seems to have been no attempt to even discuss how to put Humpty Dumpty (1) back together again if the silly sod falls off the wall.
Then we get the meat of the "Processes are to blame, not people" section. It discusses avoiding fucking up by making some small changes to working practices and so on but completely misses the real point. The blogger lacks a process for recovery. It's all very well worrying about avoiding a fuck up involving a DB deletion but how do you recover? That should be only one entry in your DR plan. BC also needs some work ...
I own a small company and I spend quite a lot of time worrying about an awful lot of things. This "contrite" article that implies that the writer has actually learned a lesson worth divulging is extremely concerning to me. I suppose it's a good start but I feel it falls rather short of identifying the real problem, learning from it and implementing a proper ... process.
(1) Humpty Dumpty: UK nursery rhyme nominally about an egg. Bit more to it than than but good enough for this discussion.
If you can't recover in peace time, what makes you think you can do it in war time?
Reminds me of Douglas Adams take on 1 in million odds
Until you have a provably working restore, the backup is nonexistent for all intents and purposes. The sort of calculation you’d need to perform to justify not doing so borders on alchemy. Unless your infra is extremely exotic this should be rather straightforward and inexpensive, and you’re one failed restore away from this this process changing immediately anyway.
How much should you value untested plans? You should value them as worthless.
But of course, testing creates more value.
there's an interesting piece on the web about how HD is not necessarily an egg, that that idea comes from an illustration added long after the nursery rhyme was well established.
https://literature.stackexchange.com/questions/1489/how-do-w...
Well maybe the people that enforced those processes, wouldn't you think?
Think safety first - this has to come from the top and it is a bit boring until you suddenly find yourself single handedly rescuing quite a few people's livelihood in the face of a disaster of some sort.
There are no real shortcuts but you can build yourself up to a decent position incrementally and erratically or you can do a formal analysis and create a plan and follow your plan - yeah right!
Start off with the basics: Do you have backups? Actually, do you have enough backups? You should have a complete copy of your data available on site (not a cluster replica) and another copy off site that might be a bit older, depending on your taste for data loss. Really work on evaluating how much data you can afford to lose. You should also have an offsite copy of your data that is immutable - ie can't be deleted or encrypted.
If you can get yourself into the safety first mood but don't know how to do it online then get a removable, USB connected disc and use that for your offline backups that you know can always be recovered from.
Now check your backups. Do some recoveries of files.
I don't know how important your company is to you but I suspect it is very important. Take some time out every now and then and do some due diligence "doo dill".
I also should do it more often ...
Instead it comes across as a veiled condescension. It doesn’t impress.
At some point a company needs to adopt practices like Netflix's chaos monkey or Google's DiRT (disaster readiness testing) to purposefully excerise continuation of business plans as well as recognize the effort that is required to keep things running. Otherwise other incentives will drown out any intrinsic motivations individuals may have to improve reliability.
... and tailoring. And "agile" when some people are aleeady working on a new version when the old one is not even tested.
It seems that with processes is like with deaths: a disaster needs to strike for people to implement a working process.
And the bureaucracy expands to meet the needs of the expanding bureaucracy (-:
But all those processes and certifications are the kind of things people bemoan in big corporations. And start-up don't have the resources nor time to do it, even if the sooner it's implemented the cheaper it is. That's IMO where incubators could add some value: have a team of people whose job is to setup and follow those processes for your smaller start-ups.
From TFA, next sentence after the one you quoted:
Also, that database was a read model and by definition, it wasn’t the source of truth for anything. In theory, read models shouldn’t have backups, they should be rebuilt fast enough that won’t cause any or minimal impact in case of a major incident. Since read models usually have information inferred from somewhere else, It is debatable if they compensate for the associated monetary cost of maintaining regular backups
The article discusses how they executed an actual recovery process from the other data sources, but it took longer than it should have (6 days), so...
We ended up with a mix of the two. We refactored the process to a point it went from 6 days to a few hours. However, due to the criticality of the component, a few hours of unavailability still had significant impacts, especially during specific timeframes (e.g. sales seasons). We had a few options to further reduce that time but it started to feel like overengineering and incurred a substantial additional infrastructure cost. So we decided to also include backups when there was a higher risk, like during sales seasons or other business-critical periods.
Don't do stuff in a rush like this. That's when I almost always make my worst mistakes. If there is a "business urgency" then cancel or get excused from the upcoming meeting so you can focus and work without that additional pressure. If the meeting is urgent, then do the other task afterwards.
I work at an energy company at the moment, the servers are overloaded because more old-fashioned IT and there's an energy crisis ongoing tripling people's monthly expenses and putting them into poverty. Everyone's panicking, but I hope they're not trying to force quick fixes or skipping due process. We're still doing code reviews and going through our regular deployment processes.
Hugo has some other interesting takes on his blog.
I don’t know him. But just because you haven’t done the exact same thing — or never been is a position of responsibility where your mistakes can affect many other people, doesn’t mean that everyone doesn’t make mistakes.
It just means that not everyone can learn from others mistakes.
I'm not claiming to have never made a similar mistake. I have messed up in production a lot. But I would definitely recognise from this episode that the lack of good tooling led to risky, rushed, manual processes that can easily go wrong.
There you go.
If I calculated this right the time they mention comes down to 30 items per second. Which is maybe not unreasonable for something that queries a whole bunch of services via HTTP, but is kinda ridiculous if you compare it to directly querying a single RDBMS.
You could probably fix this by scaling everything horizontally, if that is possible. But the real solution would be as you say to have bulk processing capabilities.
In one case, the application was running on a laptop over WiFi, which increased the network latency by 10x. Suddenly a 30 sec job turned into a 5 minute job.
Since one can easily implement a singular version using a batch size of 1, it's a drop-in replacement in most cases.
Also, since one can easily implement a batch-style API using a singular version, you can write the API batch-oriented but implement it using the singular version if that's easier. This allows you to easily swap out the implementation if needed at a later date.
threads exist, use them. If you're waiting on 500 full sequential RTTs that's your fault. network requests on local fabric can be faster than storage.
The thing I've found for object retrieval (as opposed to search) is that you might want to break GET semantics and have people POST in a list of IDs. Otherwise you might hit the query string size limit. Random tip.
I had already called out how sub-optimal this entire setup was before the incident occurred but it rang hollow from then on since it sounded like me just trying to cover for my mistake. The footguns were only half-fixed by the time I ended up leaving some time later.
That's what blameless post mortems are supposed to prevent. The only valid things to consider are questions like "is it a cost-effective prevention?", "what is the timeline of implementing this prevents / how should we prioritize?", and "does it meaningfully overlap with other changes that would prevent this?". Things like "well person A is new and that's why this happened" is not an acceptable position. Now of course, people can be not honest / not self-aware enough and are coming from that position without stating it anyway. That being said, that's why you try to focus on the objective evaluation of the recommendation and nothing else - that tends to make the bias irrelevant. Any EM/tech lead worth their salt would instantly say "is it actually important that staging and dev databases share creds"? Certainly prod should never. The other follow up to track down would be "if DNS flushing is so important when changing hosts, how have we automated it so that it happens automatically before every script that's run". I would also recommend to get rid of (ab)using the hosts file immediately and switch to explicit configuration files that you have to specify (perhaps automatically picking up "<username>.config" as a default so that scripts only ever run on your dev zone by default & locking down production.config so that only CI has the ability to touch that).
I’ve now taken 3 companies from horrid MVP (literally just got seed funding) to something a team enjoys working on. Each time I’m amazed that the app went without issues long enough to get funding (and long enough for me to refactor the whole thing.)
(Which I'm not OK doing, because that tells me that they want to track me, presumably for advertising purposes.)
They've been flailing around for a decade in search of a model that works. If UX hostility is what it takes to have a story that retroactively justifies the investment, UX hostility is what we'll get.
At least I assume that's what's happening, I haven't seen a medium paywall yet.
We want an internet with less ads, but good writers deserve to get paid. They can get paid via Medium (though how much, I don't know) through subscriptions. Is that the worse than ads or newspapers?
But if the choice before me was to pay for writers directly (like Medium), or let non-book writers as a profession disappear, I'd opt for the latter. You may criticize this attitude. I assume the responsibility for that and I'm being honest.
https://support.brave.com/hc/en-us/articles/360021123971-How...
I've seen other pay walls posted in HN that take more to get around.
Lots of people hate the ad system, won't pay for an individual blog, and get tired of hearing/seeing thr same 10 generic internet sponsors (nordvpm, express vpn, skillshare, etc.)
Medium is just painful to use.
Not a front-end dev obviously
There's a videogame streamer I like. I probably watch 4-8 hours of her content, as she live translates Japanese games while playing them. At $5/mo, that is the cheapest source of entertainment available aside from used books.
For blogs I read once a blue moon, I don't typically contribute unless it's worth supporting.
If from the comments here, or wherever is referring me to an article, I get the impression that it will be interesting/useful to me, I'll maybe go look, if not then I won't. Or sometimes the discussion the link starts is enough such that I get useful references of things to look at in other places instead (or just learn what I might want to from the discussion directly). Heck, sometimes just the subject is enough to start searching for other references.
I did something like this 23 years ago. I could bore you with specifics, but I'm sure you can guess, or worse, can fill in the details from your own similar experience. It's sorta like grabbing a hot iron skillet with your bare hand - you're not likely to ever do it twice, and like Mark Twain's cat, you won't even pull one out from the cabinet without getting an oven mitt first. Sometimes we have to learn lessons the hard way.
In this case, the datastore uses a REST API, so that should be fairly easy to implement. You could even do it in Nginx or Envoy.
Postman doesn’t know that sending a single DELETE request to that URL will delete 17 million records.
Arguably, REST interfaces shouldn’t allow deleting an entire collection with a single parameterless DELETE request.
How do you produce repeatable test results when you’re just passing around Postman configurations (they don’t want to commit the to GitHub in case there are embedded credentials)? How do you know your services are configured correctly?
`requests.get(url)` is a lot harder to mis-type as `requests.delete(url)`.
At $dayjob we would sometimes do this sort of one-off request using Django ORM queries in the production shell, which could in principle do catastrophic things like delete the whole dataset if you typed `qs.delete()`. But if you write a one-off 10-line script, and have someone review the code, then you're much less likely to make these sort of "mis-click" errors.
Obviously you need to find the right balance of safety rails vs. moving fast. It might not be a good return on investment to turn the slightly-risky daily 15-min ask into a safe 5-hour task. But I think with the right level of tooling you can make it into a 30 min task that you run/test in staging, and then execute in production by copy/pasting (rather than deploying a new release).
I would say that the author did well by having a copilot; that's the other practice we used to avoid errors. But a copilot looking at a complex UI like Postman is much less helpful than looking at a small bit of code.
And because the main branch was protected, getting the script to run meant being forced to do a PR review.
But it made things easier to not fuck up. The URL that's being modified is right there, in the code. The action being performed is a delete, it's in the code.
Makes things slightly more inconvenient but adds extra safety checks.
Its funny when the safety checks aren't enough, though. Back in the olden days, I had to drop thousands of records from a prod database because they were poisoning some charts.
Well, I was smart, you see. I first did a SELECT on the records I expected to be safe. I looked at the results, everything is okay, I'm in the right db and this is the right table. Then I did a SELECT of the records I wanted to delete. Only a couple thousand records, everything looks good.
Now all I have to do is press the up arrow key once, modify the SELECT to DELETE, and run the command.
So I pressed the up arrow key but nothing happened. I must've not pressed it hard enough to register so I pressed it again and it worked. I see the SELECT command, change it to delete, run it, aaaand it deleted hundreds of thousands of records.
What must've happened is there was a bit of lag and all my up arrows registered at once, taking me back to the select command where I looked at the good records.
Obviously I wasn't being safe enough, because I should've double checked the DELETE command I was about to run. But I thought I was being safe enough.
I had done a backup before all that, so everything was fine. But I'm still traumatized like 10 years later. I quadruple check commands I'm about to run that will affect things in a major way.
DataGrip also allows you to color-code the entire interface of a connection, so even my cowboy-access read-write connection is brightly red and hard to miss.
Pretty recently corporate changed something on my work laptop that resulted in a bunch of temporary files generated during the build getting redirected to OneDrive. I went in and nuked the temp files and shortly thereafter got a message from OD saying 'hey noticed you trashed a ton of files, did you mean to do that?'
The developer side of me thought 'of course I did, duh' but I can imagine that's useful information for most users that made an innocent yet potentially costly mistake.
* checklists/runbooks with sample commands so that you copy/paste instead of manually typing or relying on your shell's history
* explicit scripts, named in a verbose and clear way, to be executed for dangerous steps (e.g. to delete something or deploy to prod) instead of allowing typing random commands
* for really critical operations, a mandatory code review/a second pair of eyes in some way (the blog seems to mention something of this nature)
One place I worked we had detailed step by step checklists with ready to copy paste commands for patching server. The Windows process was forever drifting and getting out of date while the Linux process was being kept up to date.
The difference was that the Windows admins were all morning people who enjoyed showing up to work at 7am or earlier if they could justify it and just doing it from memory. While the Linux team were all night owls who preferred to roll out of bed 5min before the update schedule, remote in, execute the process with as much automation as possible in the hope of getting some more sleep before having to got to work.
Basically the Windows process was only ever actually followed when someone outside of the Windows admins had to cover for them and the Linux process was used every time by the Linux team.
Testing / separation of dev environments is almost black magic.
From the article, it sounds like it would have helped to formulate the correct command in a non-production environment before moving to production
Outside of that, I highly recommend using a managed service, either AWS ElasticSearch service, or Elastic.co to make recovery and management easy. AWS does a snapshot on every index every couple hours and it's relatively simple to restore a deleted index.
Also, I'm sure the author knows this now, but don't ever run any command on a live production data store without triple checking it.
I agree. And luckily that was the case for them.
You can use a tool like elasticsearch-curator or even cron to manage running backups or use the built in scheduling (snapshot lifecycle management)
It is interesting they have all that architecture and an endpoint that can just delete everything.
My biggest concern about restoring that Elasticsearch backup would be that the restored backup would be inconsistent with the real source of truth and it might be hard to reconcile to bring it up to date.
In other words, it only has to be good enough for a few days (ideally - hours).
Also, what happens if somebody nukes the event table? How is your precious event sourcing supposed to save you then, huh?
This made me appreciate how desperate is the industry for software engineers. :)