The Day I Took Down 100,000 Web Sites
sef.kloninger.com
sef.kloninger.com
Even if I'm doing a mass file delete, if I've got the time for it, first thing I'll do is move/rename the files to a backup / marked-for-deletion directory. Watch for strangeness. Then commit.
If it's site accounts or features, disable them, prevent them from being served, mark them inactive in the database / server. Leave them for a week or two. Then nuke.
Legally proscribed deletions (illegal content, etc.) are another matter, but most other content is eligible for a graduated approach.
(Assuming Linux) If you're concerned with open files, you can do an 'lsof' and screen for files of interest. 'fuser -v filename' will show the process using a file, 'fuser -vk filename' will kill it (xargs or loops for many files, obviously).
And not everything that you're cleaning up is necessarily files, so adapt methods accordingly.
I don't take credit for that -- it was the work of smart folks who came before me. But it sure saved my bacon this time.
I owned a pay per click management agency from 2005-2009. In 2008 I decided to buy some Enterprise software to help me better manage big clients in competitive spaces. I migrated these clients over to the software, set up a bunch of automated rules to change bids and called it good. Then Saturday hit and we suddenly exceeded a single account budget by $40,000.
That was the day that I lost $40,000. I took responsibility for every penny of that loss.
After a thorough internal investigation, we determined it was a bug in this Enterprise software. We took our records to the software company and they admitted or confirmed that it was indeed a bug on their side. Unfortunately, they were such a big company that we couldn't get anywhere when seeking a refund. They lawyered up real fast upon our request for our money back. I didn't have the means to pursue it at that point, so we took the hit as a business and continued on our way.
The lessons I learned: audit every automated task, find ways to lower your risk while implementing changes, and hire/retain a business lawyer for the duration of your business.
I have had the experience of shaming a company with a little essay that, over the years, garnered a few responses from other customers, a few employees, a particularly irate investor. And eventually watched the firm file for bankruptcy liquidation.
Fun, that.
Monday sucked.
And if at all possible, don't put major changes live on a Friday afternoon; you're just going to ruin yours and everyone else's weekends.
Things were probably moving so fast that revising it was always an option. Instead of changing how that main file worked, it was "best" and part of Ning's policy to just note that its "important" not to mess with.
That said, if you have the resources to create a team of people to revise early code and how it relates to the whole system, its important to just get it done instead of waiting for something to "blowup".
- real costs - database storage, blob storage, (minor) perf impact
- hidden costs - tougher to create test environments that look like production if production is so large, but largely invisible / unusable.
There's a reductio ad absurdum open in the room-cleaning premise. I could write a million lines of code that don't do any measurable harm -- calling a no-op -- and by your standard they ought not be removed. Not doing any harm isn't the same as actively contributing value, which is what most people want their codebase to do.