I Accidentally Deleted 7TB of Videos Before Going to Production
blog.thevinter.com
blog.thevinter.com
It always does.
> Well, it teaches me to do more diverse tests when doing destructive operations.
Or add some logging and do a dry run and check the results, literally simple prints statements:
print("-----")
print("Downloading videos ids from url: {url}")
print(list of ids)
...
...
...
# delete() dangerous action commented out until I'm sure it's right
print("I'm about to delete video {id}")
print("Deleted {count} videos") # maybe even assert
...
Then dump out to a file and spot check it five times before running for real.This is basically the approach that was taken: log before and after every action exactly what data or files is being acted on and how. Don't actually do it. Then have multiple people inspect the logs. Once ok'd, run again, with manual prompts after each log item asking to continue, for the first few files/bits of data. Only after that was ok'd too did it run the remainder.
In other things I've worked on, I've taken the terraform-style plan first, then apply the plan approach, with manual inspection of the plan in between.
Multiple reviewers here didn't catch the mistake
https://www.bloombergquint.com/markets/citi-s-900-million-mi...
It's used rather extensively in safety-critical public transportation in Japan [1] and to a lesser extent in New York (along with many other countries) [2]. This can easily extend to software without overcomplicating by just setting the expectation that engineers, Q&A, etc. do this even when alone.
[1] https://www.atlasobscura.com/articles/pointing-and-calling-j...
https://rachelbythebay.com/w/2020/10/26/num/
The take away by both is there is actually something to do which can wake people up when the stakes are high, and they might not be doing what they expect.
Also, I now realized that aviation checklists seem to tend to be done similarly with gestures - at least from what I saw on YouTube, not sure if that's representative or only used during education (?)
I'm honestly trying to think of the way how I could approach this for myself, just I don't see a clear solution yet that wouldn't require me to spell out everything I type in my terminal window.
If your application has space concerns, you can modify this approach to be like a recycle bin where you delete records which are no longer valid and have been invalid for over a month (or whatever time frame is appropriate for your application). However, I think this is unnecessary in most cases except for blob/file storage.
Sure, but we can only do so much. I find its good bang for buck and alternatives that might prevent that are not always available, so we do the best we can. You gotta make a call on whether its enough or not.
For database entries, flag for deletion, then delete.
In the files case, the move or rename also accomplishes the result of breaking any functionality which still relies on those file ... whilst you can still recover.
Way back in the day I was doing filesystem surgery on a Linux system, shuffling partitions around. I meant to issue the 'rf -rm .' in a specific directory, I happened to be in root.
However ...
- I'd booted a live-Linux version. (This was back when those still ran from floppy).
- I'd mounted all partitions other than the one I was performing surgery on '-ro' (read-only).
So what I bought was a reboot, and an opportunity to see what a Linux system with an active shell, but no executables, looks like.
Plan ahead. Make big changes in stages. Measure twice (or 3, or 10, or 20 times), cut once. Sit on your hands for a minute before running as root. Paste into an editor session (C-x C-e Readline command, as noted elsewhere in this thread).
Have backups.
And yes, copy, verify, delete. And make sure by the code structure that you either do the three on the same files, or their fail.
Also, do it slowly, with just a bit of data on each iteration. That will make the verification step more reliable.
Anyway, for a huge majority of cases, only having backups is enough already. Just make sure to test them.
Example:
cd datadir
mkdir delete
mv <list of files to be deleted> ./delete
# test to see if anything looks broken.
# This might take a few seconds, or months, though it's usually reasonably brief.
rm -rf ./delete
The reasons for mv:- It's atomic (on a single filesystem). There's no risk of ending up with a partial operation or an incomplete operation.
- It doesn't copy the data, it renames the file. (mv and rename are largely synonyms.)
- There's no duplication of space usage. Where you're dealing with large files, this is helpful.
The process is similar to the staged deletion most desktop OS users are familiar with, of "drag to trash, then empty trash". Used in the manner I'm deploying it, it's a bit more like a staged warehouse purge or ordering a dumpster bin --- more structured / controlled staged deletion than a household or small office might use.
> ... Then have multiple people inspect the logs. Once ok'd, run again, with manual prompts after each log item asking to continue...
This sort-of reminds me of some "critical" work I had to do a couple of decades ago. I was in a shop that used this horrifically tedious tool for designing masks for special kinds of photonic devices-- basically it was tracing out optical waveguides that would be placed on a crystal that was processed much like a silicon IC.The process was for TWO of us to sit in front of computer and review the curves in this crazy old EDA layout tool called "L-edit" before it got sent to have the actual masks made (which were very expensive). It took HOURS to check everything.
The first hour was tolerable but then boredom started to creep in and we got sloppy. The whole reason TWO people got tasked with this was because it was thought that we would keep each other focused-- 2 pairs of eyes are better than one, right?. Instead, it just underscored the tedium of it all. One day someone walked in and found us BOTH in DEEP SLEEP in front of the monitor. Having two people didn't decrease the waste caused by mistakes, it just bored the hell out of more people.
Was it worth it? No, I don't think so from an opportunity cost perspective-- even though we were the most junior folks there. A mind is a terrible thing to waste!
I think that this is the most important part of any check. Your parent refers to checking the log five times, but, at least in my experience, I won't catch any more errors on the fifth time than the first—if I once saw what I expected rather than what was there, I'll keep doing so. Of course everyone has their blind spots, but, as in the famous Swiss-cheese approach, we just hope that they don't line up!
See PDCA for more a more time critical decision loop. https://en.wikipedia.org/wiki/PDCA
I was running commands manually to interact with files and databases, but was quickly shown that even just writing all the commands out, one by one gives room personally review and get a peer review, and also helps with typos. I could ask a colleague "I'm about to run all these commands on the DB, do you see any problem with this?". It also reduces the blame if things go wrong if it managed to pass approval by two engineers.
While I'm thinking back, another little tip I was told was to always put a "#" in front of any command I paste into a terminal. This stops accidentally copying a carriage return and executing the command.
For a one-liner sure, but a multi line command can still be catastrophic.
Showing the contents of the clipboard in the terminal itself (eg via xclip) or opening an editor and saving the contents to a file are usually better approaches. The latter let’s you craft the entire command in the editor and then run it as a script.
[For Bash] Ctrl + x + Ctrl + e : launch editor defined by $EDITOR to input your command. Useful for multi-line commands.
I have tested this on windows with a MINGW64 bash, it works similarly to how `git commit` works; by creating a new temporary file and detecting* when you close the editor.
[0] https://github.com/onceupon/Bash-Oneliner
* Actually I have no idea how this works; does bash wait for the child process to stop? does it do some posix filesystem magic to detect when the file is "free"? I can't really see other ways
It will prevent any number of newlines from running the commands if they're pasted instead of typed.
You can enable it either in .inputrc or .bashrc (with `bind 'set enable-bracketed-paste on'`)
Anyway, I selected what I though was a "merge all duplicates" option without previewing results. What I had actually done was "merge all selected". So, the system proceeded to merge a very large % of the database... Into One. Single. Record.
Luckily the vendor kept very good backups, and so I kept my job. Because I also luckily had a very good boss and I had already demonstrated my value in other ways, he just asked me "Well, are you going to make that mistake again?". I wisely said no, and he just smiled and said "Then I think we're done here."
I have been particularly fortunate throughout my career to have very good managers. As much as managers get a lot of flack here on HN, done well they are empowering, not a hindrance, and I attribute a lot of success in my career to them.
I think that, if you've only learned something like that the easy way, then you haven't learned it yet. As long as everything's only ever gone right, it's easy to think, I'm in a rush this one time, and I've never really needed those safety procedures before, ….
It's simple to manually test corner cases, and then when everything is smooth I can just
script1 | xargs script2
It's also handy if the process gets interrupted in the middle, because running script1 again generates a shorter list the second time, without having to generate the file again.When I'm trying to get script1 right I can pipe it to a file, and cat the file to work out what the next sed or awk script needs to be.
Over-simplified example:
1. Copy stuff from A to B
2. Delete stuff from A
(Obviously you wouldn't do it like that, but just for illustration purposes.) It's all fine, but (2) assumes that (1) succeeded. If it didn't, maybe no space left, maybe missing permissions on B, whatnot, then (2) should not be executed. In this simple example you could tie them with `&&` or so (or just use an atomic move), but let's say these are many many commands and things are more complex.
You could always generate an intermediary set, inspect/test/etc, and then apply it with Python. I've done that too, works just as well. The important thing is to separate the planning step from the apply step.
* where "complicated" means more complicated than, for ex, `rm some_path.txt` or `DELETE FROM table WHERE id = 123`.
- make sure the sum of their lengths == number of total current items
- make sure items_to_be_kept.length != 0
- make sure no two items appear in both lists
- check some items chosen at random to see if they were sorted in the correct list
At this point the only possible mistake left is to confuse the lists and send the "to_be_kept" one to the delete script; a dry run of the delete list can be in order.
This is the most reliable way. Bash has a few niceties for error handling, but if you are using them, you would probably fare better in another language.
If you do insist on Bash, quote everything, and use the "${var}" syntax instead of "$var". Also, make sure you handle every single possible error.
It's like the difference between "do what I say" and "do what I mean"...
- -WhatIf to see what would happen if you run the script;
- -Confirm, which asks for confirmation before any potentially destructive action.
Moreover these arguments get passed down to any command you write in your script that support them. So you can write something like:
[CmdletBinding(SupportsShouldProcess)]
param ([Parameter()] [string] $FolderToBeDeleted)
# I'm using bash-like aliases but these are really powershell cmdlets!
echo "Deleting files in $FolderToBeDeleted"
$files = @(ls $FolderToBeDeleted -rec -file)
echo "Found $($files.Length) files"
rm $files
If I call this script with -WhatIf, it will only display the list of files to be deleted without doing anything. If I call it with -Confirm, it will ask for confirmation before each file, with an option to abort, debug the script, or process the rest without confirming again.I can also declare that my script is "High" impact with the "ConfirmImpact = High" switch. This will make it so that the user gets asked for confirmation without explicitly passing -Confirm. A user can set their $ConfirmPreference to High, Medium, Low, or None, to make sure they get asked for confirmation for any script that declare an impact at least as high as their preference.
[1]: https://docs.microsoft.com/en-us/powershell/scripting/learn/...
Cause if it is an entirely separate code path, doesn’t that introduce a case where what you say you’ll isn’t exactly what actually happens?
> because I didnt read the docs
Ouch.
> Or is it a separate routine that you have to write?
If you are writing a function or a module what would do something (eg API wrapper) then of course you need to write it yourself.
But if you are writing just a script for your mundade one-time/everyday tasks and call cmdlets what supports ShouldProcess then it works automagically. Issuing '-whatif' for the script would pass `-whatif` to any cmdlet what has 'ShouldProcess' in it's definition. Of course if someone made a cmdlet with a declared ShouldProcess but didn't write the logic to process it - you are out of luck.
But if have a spare couple of minutes check the docs in the link, it was originally a blog post by kevmarq, not a boring autodoc.
T'was a sad day.
This pattern has the added benefit that it makes it really easy to write unit tests, which is something often sorely lacking in these sorts of batch scripts. It also makes full automation down the line a breeze, since you have nice shearing layers between your components.
https://news.ycombinator.com/item?id=29083367
It defaults to not doing anything so you can gradually and selectively have it do something.
Learned about when I posted my command line checklist tool on HN: https://github.com/givemefoxes/sneklist
(https://news.ycombinator.com/item?id=25811276)
You could use it to summon up a checklist of to-dos like "make sure the collection in the dictionary has the expected number of values" before a "do you want to proceed? Y/n"
(step1) Download video from URL. Include the Id in the filename.
(step2) Grab the list of files that have been downloaded and parse to get the Id. Using the Id, delete the original file.As they say, measure twice, cut once.
Don't feel bad, I think every professional in IT goes through something similar at one time or another.
Decide, then act.
There's a whole menagerie of failure modes that come from trying to make decisions and actions at the same time. This is but one of them.Another of my favorites is egregious use of caching, because traversing a DAG can result in the same decision being made four or five times, and the 'obvious' solution is to just add caches and/or promises to fix the problem.
As near as I can tell, this dates back to a time when accumulating two copies of data into memory was considered a faux pas, and so we try to stream the data and work with it at the same time. We don't live there anymore, and because we don't live there anymore we are expected to handle bigger problems, like DAGs instead of lists or trees. These incremental solutions only work with streams and sometimes trees. They don't work with graphs.
Critically, if the reason you're creating duplicate work is because you're subconsciously trying to conserve memory by acting while traversing, then adding caches completely sabotages that goal (and a number of others). If you build the plan first, then executing it is effectively dynamic programming. Or as you've pointed out, you can just not execute it at all.
Plus the testing burden is so drastically reduced that I get super-frustrated having to have this conversation with people over and over again.
Automated tests are awesome :)
During buildup of the our_id list: assert (vimeoId not in our_ids).
After creating the list: assert len(set(our_ids)) > 10000 and assert len(set(our_ids)) == len(our_ids)
Before each final deletion: assert id not in hardcoded_list_of_golden_samples.
Depending on the speed required you could hit the api again here as an extra check.
But as always everything is obvious in hindsight. Even with the checks above, Plan+Apply is the safest approach.Yes, that can be a simple but powerful live on screen log. I developed a library to use an API from a SaaS vendor, in much the same way as the author. It was my first such project & I learned the hard way (wasted time, luckily no data loss or corruption) that print() was an excellent way to keep tabs on progress. On more than one occasion it saved me when the results started scrolling by and I did an oh sh*t! as I rushed to kill the job.
Before: there is chance there is a bug in my "delete" use case
Now: what we have before plus the change that there is a bug in my "--live-run" flag
If both paths are on the same disk moving files is a fast operation - and if you discover a screw up, you can easily undo it. On the other hand if everything still looks fine after a few days, you just `rm -rf` that folder and purge the files.
Instead of performing the dangerous action outright, just log a message to screen (or elsewhere) and watch what is happening.
Alternatively, or subsequently, chroot and try that stuff on some dummy data to see if it actually works.
Make sure you got everything out and off before you pull up your pants, or else you better be prepared to deal with all the shit that might follow!
SELECT COUNT(1) FROM table
-- UPDATE table SET col='val'
WHERE 1=1 BEGIN TRANSACTION
UPDATE table SET col='val' WHERE 1=1
ROLLBACKGood decisions come from experience. Experience comes from making bad decisions.
(Your description is so, so, spot on.)
While I didn't know what I was doing, I did manage to get the beeping to stop, and had to come in at 5 a.m. the next day to restripe the drive I'd yanked out.
Did I mention there were no backups? When I was a little bit more seasoned on the job, I raised a polite but persistent issue with management of the need for durable backups. Although I kept at it for months, they thought about it, talked about it, and ultimately did nothing. A few months after I left, the entire array failed. Since the group's work relied on the irreplaceable data, all work ground to a halt for the several months it took for an off-site company to recover the data.
That backups comment sounds very familiar.
I accidentally deleted a clients products table from the production database in my early years as a solo dev. There was only a production database. Luckily I had written a feature to export the products to an excel sheet a while before and happened to have an excel copy from the prior day. I managed to build an export to ingest the excel and repopulate the table in record speed while waiting for my phone to ring and the client to be furious. Luckily they never found out.
Thank God my automatic backups were so close to the mistake I made and I didn't lose 24 hours.
Haven't made a mistake like that since and I don't destroy DB records like that anymore.
One of the best things about HN is that so many incredible, talented people post. It's incredibly inspiring to raise your own game, to see what the best are doing. But sometimes it's equally important to realise we all fuck up, and for every unicorn dev there's another thousand of us grinding away.
OP - well done for sorting the problem and telling us all about it!
Vimeo OTT has a codebase written in Rails, whereas the main PHP application is written in PHP. At the time Vimeo acquired Vimeo OTT's codebase, the Vimeo OTT codebase was small — around 10,000 lines of Ruby. Rewriting that codebase inside the Vimeo PHP application would have been a tough technical challenge for the all-Ruby team, and they'd have likely lost some people along the way and missed out on some content deals, so they decided instead to maintain two separate codebases and two separate login systems.
The video-playback and video-storage infra has since been unified, but all the business logic is still siloed.
I know how these things happen. Support ticket queues and all. And while I don’t fully know the difference in cost, I would assume a customer upgrading to an Enterprise plan would get a better support experience.
Whoever within authors company negotiated the upgrade to Enterprise (or didn’t) and failed to embed some agreement around OTT to Enterprise transition assistance was the one who made the first mistake.
Yes and No. At the end of the day, you as a business have to insulate yourself from your infrastructure provider.
This story resonates with many people here because many experienced engineers had done something similar before. For me, destructive batch operations like this would be two distinct steps:
1. Identify files that need to be deleted; 2. Loop through the list and delete them one by one.
These steps are decoupled so that the list can be validated. Each step can be tested independently. And the scripts are idempotent and can be reused.
Production operations are always risky. A good practice is to always prepare an execution plan with detailed steps, a validation plan, and a rollback plan. And, review the plan with peers before the operation.
> These steps are decoupled so that the list can be validated. Each step can be tested independently. And the scripts are idempotent and can be reused.
This is the most underrated comment.
I'm saying it as someone who had the ultimate oversight of deleting hundreds of TBs per day spread of billions of files on different clouds and local storage.
If you need to do it once, it’s probably 2-3 hours of work? That is identifying a duplicate video and then clicking the button(s) to delete it once every 20 seconds.
Reminds me of https://xkcd.com/1205/
I have found working with Vimeo to be very frustrating, especially recently. They have a great video solution, especially for streaming, but they seem to put these unnecessary and frustrating roadblocks that make me constantly question my decision to use Vimeo. From in ability to move videos from one place to another, requiring complete uploads (resulting in problems like this post) to nonsensical limits and pricing, especially on their new webinar offering, which has a limit of 100 registered attendees. For anyone who has run webinars before, this makes no sense since 100 registered attendees usually means 20-30% of those people actually attend, so you're capped at 20-30 live attendees. They should price it like most event sites and charge per live attendance rather than registration.
Regardless, I've been very frustrated with Vimeo since it could be so much better if they didn't have these roadblocks in place. If they could have easily enabled moving videos from one product to another, the post (and 7TB of lost videos) would never have happened. It wasn't always this way with Vimeo, but they went IPO in May 2021 and it's no surprise they're turning the screws on their product offering and pricing now.
A friendly team will harness that enthusiasm and tame the quickness / encourage respect for production. We all made a massive doo doo and its how you proceed that'll define your career.
1) parsing a web page shouldn't be considered incredibly fraught with problems
2) that reloading web pages should be part of (1)
3) that this should ever possibly be run without validating the list of files that would be deleted
So forget the specifics. Where are people learning these things, and what do we do to teach them better things?The point is not to prevent these mistakes, but to keep the consequences low.
Have backups, have version control, etc.
Would that work? I don’t see a bear backing down and I don’t see the human winning either.
Learn to learn and learn to work carefully. It starts in school and should be part of a proper college/university education or vocational training.
There's several ways of learning the specifics: by experience on-the-job, which can be hard if mistakes can get you fired; or by putting in the work in your free time.
If your job is to work with certain web frameworks and you're not very experienced, either ask senior devs to assist/review before going live with critical changes. Alternatively, practice at home. Unpopular, but you need to get experience from somewhere. OSS projects are a great way to do that - be that by creating your own or by contributing to an existing one.
You will do it at least once in your career. If you're old enough you will do it twice. If you're really old, you get the joy of doing it a third time.
The subtlety increases each time because you do learn.
If someone delivers code that looks like that, especially if intended for a production system, I'm firing immediately.
It's a miracle nothing has happened sooner.
>I'm a Junior Developer with less than one year of actual experience.
>The bad news is that this was on Friday, and we needed to have the videos back up at most for Tuesday morning.
You say:
>If someone delivers code that looks like that, especially if intended for a production system, I'm firing immediately
Fire immediately? What a miserable sounding place to work.
It’s a miracle nothing happened sooner
Only recently they made some god aweful policy changes for content creators(1), but it looks like they treat their enterprise customers just the same.
Surely, there must be better alternatives for hosting videos than being at the mercy of a company who couldn't care less about big paying customers.
(1) https://www.theverge.com/2022/3/18/22985820/vimeo-bandwidth-...
It also makes it more testable. Instead of putting the delete call right in the loop, split it into four functions.
function getAllVimeoVideos()
function getAllDbVideos()
function getVideosToDelete(vimeo_videos, db_videos)
function deleteVideos(videos_to_delete)
Your core logic lives in getVideosToDelete which is simply a set difference.Given that there are only a few hundred videos, it is easy to run the getter functions above and quickly verify they are returning what you expect.
List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars)
function is the first time I thought about time complexity in my job.Say Foo and Bar have fields in common, such that you can say a Foo object "equals" or "matches to" a Bar object, like if they have name and dateOfBirth fields or something else that are the same (nothing like a common ID between the two). Now say there are some other fields too, like amountSpentThisYearOnDogFood that you know is always accurate for Bars, but might be out of date for Foos. How do you get the list of all the Foos to update?
Initially I did the nested for loop solution that's like
List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars)
{
List<Foo> returnList = new List<Foo>();
foreach (var foo in foos)
{
foreach (var bar in bars)
{
// check if "equal" or "matching" based on some criteria
// if equal, update foo dog food expenditure with bar dog food expenditure, add to returnList, and break
}
}
return returnList;
}
but that's O(n^2) right.The solution with a Dictionary is obviously better. All you need to ensure is that you have a method for both the Foo and Bar classes that will produce the equivalent hash for both, if they would be considered equal or matching by whatever criteria you are using.
So you could have something like
int GetHashOfFoo(Foo foo)
{
string firstName = foo.FirstName;
string lastName = foo.LastName;
DateTime dob = foo.Dob;
return (firstName, lastName, dob).GetHashCode(); // convenient c# method
}
int GetHashOfBar(Bar bar)
{
string firstName = bar.FirstName;
string lastName = bar.LastName;
DateTime dob = bar.Dob;
return (firstName, lastName, dob).GetHashCode();
}
These two functions will return the same value if those fields are the same. So then you can do something like List<Foo> getFoosToUpdate(List<Foo> foos, List<Bar> bars)
{
List<Foo> returnList = new List<Foo>();
Dictionary<int, Bar> barsByHash = new Dictionary<int, Bar>(bars.Count);
foreach (var bar in bars)
{
int barHash = GetHashOfBar(bar);
barsByHash[barHash] = bar;
}
foreach (var foo in foos)
{
int fooHash = GetHashOfFoo(foo);
if (barsByHash.ContainsKey(fooHash)
{
returnList.Add(foo.CopyWith(dogFoodExpenditure: barsByHash[fooHash].DogFoodExpenditure))
}
}
return returnList;
}
Which is faster cause you only have to go through the bars list once.I actually messed up something like OP with this, but with doing undesired additions instead of undesired deletions.
You can think of it as having two endpoints, both expecting a .csv with rows being the things you were updating/changing/deleting.
The problem was, there was a column to indicate (with a character) whether the row was for an edit, or addition, or deletion, but this was only with one of these endpoints. For the other, there was only addition functionality, but I thought changes and deletions were also options for the other kind of .csv due to some unwise assumptions on my part (thinking that the other .csv would have the same options as the other). That's how we accidentally put in over 100 additions that should have been changes that had to be manually deleted. Luckily I had a list of all the mistaken additions.
Don't write a blog post.
9 years ago I was working for a major broadcasting company in the arse end of London as a junior dev, building one of their Android apps.
We'd roll features out months before & enable them with feature flags via a json file we'd manually push to a prod server at a later date.
We'd just built a huge new feature letting you request content to be downloaded to your set top box remotely & it had a 250k marketing campaign to go along with the launch.
Senior dev trusted me with prod deployment rights.
I pushed the wrong json config to prod, launching the feature weeks before the marketing campaign.
Thank god I was a junior perm, that was definitely a firing offence.
That part's crazy! If you think it was a firing offence wouldn't they've been fired? (I don't think it is, but obviously requires system changes/explanation.)
> foreign to the "Silicon Valley" world but paints an accurate picture of what
> development is for small IT companies around the world
Everybody makes mistakes even in the "Silicon Valley" world, but such problems cloud be easily caught by testing (which he did but it was restricted to the first page) and performing a simple dry-run.
Things are complicated, people are human and forget things, there are pressures to "get it done" and override the guardrails. Everybody has horror stories. Some worse than others. Welcome to the OP's day of horror. I would think "Silicon Valley" dev-ops horror stories make this one seem like a triviality.
1. Vimeo responds to the original request with "will look into it", then... nothing happens? This may depend on culture, but at least from my experience in the UK, this is a very non-committal response, and if you really want them to do something, you'll need to chase them. Wait a few days and inquire if they have any estimate for when it might get done, or if they need more information. I find that the "looking into it" response is sometimes used to gauge how important the request is to you.
2. Once you go with your own solution, just drop a quick message to Vimeo: "Hey, just wanted to let you know we've found our own solution for this, and won't require your help any more. Sorry if you've already committed any resources for this task. Have a nice day, yada yada." This not just avoids what happened here, but is also a courtesy to them.
In my defence, we had to get 2 PR approvals before anything was merged! But I definitely learned a thing or two from that experience
(The problematic line lacks the closing ", probably a typo? I though it closed in an unexpected location)
Start by creating one empty page for every component of your system. You won't remember them all, but over time you can add missing ones. Each page is the authoritative source of info on that component. If you need more pages for one component, put them in a directory of the same name as the page and add ".d" to the directory name, and link to them from the first page. Finally, create a diagram (however you want) that includes every component you have a page for. Add the count of components to the top of the diagram. If the count on the diagram doesn't match the number of documents, time to update the diagram. If you ever add, remove or rename a page, time to update the diagram. If you do this the same way for every different system you have, you can link them all together and get both small and large scale diagrams. (p.s. don't waste time automating this unless you find the system changing constantly or you have a very big system)
We've all made big dumb mistakes. Recover and learn.
rm -rf /path/to/delete/ *
And realize it is taking too long... rm -rf /path/to/delete/ *
Note the space between the last / and the *This will recursively remove the directory /path/to/delete and remove every file/directory that matches * in the current directory where 'rm' is being run.
When what was most likely meant was:
rm -rf /path/to/delete/*
Note the lack of a space between the last / and . This will remove all files that match that reside in the /path/to/delete/ directory.Oh yes Vimeo, the crappy company that won't let you play videos unless you enable autoplay in your browser[1].
Selecting them as a provider was the actual mistake.
[1] https://askubuntu.com/questions/777489/vimeo-video-not-playi...
When I just started as a junior dev at a small company I made the classic mistake of emptying the prod db instead of my local dev db. This was a small and in hindsight insignificant project. But Google was our customer, so it didn't feel insignificant at the time.
In this case my inexperience was partly my savior. All the data was inputted by people via a web form. Normally you're supposed to use POST to submit a form. But I was quite clueless at the time, so I had used GET. This meant all requests were still in the Apache logs. I could simply replay all requests.
I still feel my hard pounding when I think about the moment I realized what had happened. I was really relieved when everything was back!
What I learned from this incident:
- make automated backups
- no access to prod db from anywhere but prod
This allows me to run a small sub-set of commands and test those under a live-environment before running all commands at once. In addition, this also functions as a complete log of what has been changed manually in production.
Not everyone is an innate rockstar developer who provisions k8s clusters for breakfasts and delivers features for lunch!
Being a developer is a really hard job and there are endless complexities and difficulties along the way and when we are more seasoned already.
Don't let any negative feedback deter you from keeping doing what you're doing: learning from your mistakes and improving along the way!
Means, if you delete one small file you need one confirmation, if you delete thousands, you need a intent stating i expect thousand files to be deleted. Same goes for size. So not a okay button, but instead a form allowing you to enter the dimension of the intented outcome. 100 files max, 1 gb max deleted.
If the request goves over the intent, the system aborts.
What happened is that no one knew how to react and I was probably the best suited for it, we don't really have seniority in office.
That said when I deleted the videos I immediately told my boss. He was kind of scared but his reaction was mostly "Well, now we have to re-upload them immediately, find a way. The people that uploaded them once won't be doing it twice". I was basically left on my own to find a solution (which I luckily did).
Please note that I'm in no way blaming my company or accusing it of something, this is the standard knowledge base and way of dealing with things in many places, contrary to what working in big tech or reading HN might make you believe!
> "HN is the top 1%" + "this is the standard knowledge base and way of dealing with things in many places, contrary to what working in big tech or reading HN might make you believe!"
I'm in fact from Spain and now live in Japan, and I believe the practices in Spain would be as bad as Italy, and in Japan they are def worse (great at hardware, horrible at software), so I do understand a lot of what you are saying. FWIW, in Spain I've seen whole dev teams composed only of interns!
> "we landed a big contract for one of the biggest gym companies in Italy, the UK and South Africa" + "we don't really have seniority in office"
Maybe now that seems like you have the budget it's a good time to go to management and suggest to hire some senior devs who can mentor the rest into learning best practices? You can sell it like a reinvestment in the company to management if they want to take it as pure profit. If Italy is like Spain, many devs won't really even want to learn these things, but some will and then those will become seniors at some point.
I think it also teaches us that adversity sometimes leads to better solutions. I love that the OP made a hacky script that did in 4 hours what a guy was paid to do manually over several months!
To rebillionizing!
https://www.youtube.com/watch?v=wGy5SGTuAGI&t=369s
...yeah, the Tres Commas bottle was on the DELETE key. The corner of it was just, it juuuust got on there...
I venture this kind of (misplaced) over-confidence is not atypical of many junior developers. As someone with a few years under my belt, I don't care how sure I was of the code I wrote that deletes important data, I would have gone through the code over and over again, and at least ran a simulation (by maybe logging the generated delete urls for manual verification).
It's a rite of passage and we all went through something like this. It's how you learn and grow.
>It also should probably teach something to Vimeo
No. Even if Vimeo could have made things better, it's still your fault. You have to take responsibility for your business. At the end of the day, if this causes the closure of your company, Vimeo is still fine.
So: I never feed the data straight from the gathering script into the modifying script, at least not in the first runs. Instead, I dump the whole list of items into a file, count them in there, gawk at them to see that they're right, and compare with the source data by hand until I begin to annoy myself. Then I feed that file to the second script.
I think I would reflect on why this is a script to begin with. It's run once and with only 500 items could be done manually, though 500 is certainly a bit much.
But it's not a massive time saver; the point of the script should be almost entirely to increase accuracy. I think I would write one script to generate the list of videos to delete; that's the part that's actually difficult, and a human can then verify the list. I would probably just delete them by hand after that, but if I really wanted a script for that part too, it would be a separate script that uses a list that has been vetted by a human even if initially created by the first script.
This is why I use so many print statements and comment out destructive actions! Lots of experience with these feelings!
The only concern is Google Drive as the only backup, please make sure you have a local copy on a local RAID drive and another one regularly archived and stored in a bank locker.
I am also a junior dev and completely empathize with being given a lot of responsibility and potentially messing up. I think for someone with < 1 year of experience, to solve the problems you created as fast as you did is really impressive. Thankfully your story ends well :)
Still things happen. Hopefully you have a large enough client base where some bad experience doesn't define the whole thing.
Even though you were fortunate not to lose any data, you gained a lot of experience!
I'm guessing users accidentally deleted multiple documents one too many times, and now it's baked in.
> Fri May 06 2022
> I'm currently working [...] in Italy
https://thenextweb.com/news/how-pixars-toy-story-2-was-delet...
> my mind thought that url would refresh itself as soon as the page variable changed
This is what I thought too when I read the code. I don't think it's obvious at all!
That's not what a physical backup means.
Everywhere I have to implement a delete operation, I never hard delete data on first call.
url = f"https://api.ourservice.com/media?page{page}&step=100 ?
It does the same thing as `https://api.ourservice.com/media?page${page}&step=100` [sic] in Javascript, or "https://api.ourservice.com/media?page$page&step=100" in Bash, PHP, Perl or Groovy (and other languages). It outs you into variable substitution / interpolation in the string literal.
In Python these string literals are called f-strings if you want to look it up. They are defined in PEP 498 - Literal String Interpolation [1] and available since Python 3.6.
[1] https://peps.python.org/pep-0498/
[sic] there probably would be a missing '=' in this url after "?page"
[0] https://docs.python.org/3/tutorial/inputoutput.html#tut-f-st...
- Captain Obvious
Previously used rclone for doing massive transfers between cloud providers using "cheap" on-demand servers which provide unlimited data transfer (the public clouds make this very expensive).
Spoiler:
But there was a backup that could be reuploaded in time and everything was fine in the end.
Also Junior Dev: "Here's my source code"
Even as a Sr Dev I'd share stuff like that, it's code that'd appear on a stack overflow post anyway.
At least store in large TB hard disks connected with a SATA adapter when needed, and put them in a case in a safe place (better: two copies, stored in two places). What is the HD + copy time price relatively to production work ?
It's not even necessary to the story.
Thanks for the comment :)
1) Your current blog has your current employer + client linked to it. 2) Your github has your real name. 3) All of these have been crawled/archived.
None of this bodes well for your career in the future. While I think your blog post is a great war story, it's really not a good idea to post it on your main account which can be traced back to your real name and CV because it will come up the next time you apply for a job.
Unfortunately, even if it illustrates a great deal of ingenuity and creativity on your part in fixing a mess you made, many folks will take one look at it and be judgmental. You have to manage your reputation online and be careful.
Your current job is linked in your CV.
>I also want to preface this whole post by saying that I'm a Junior Developer with less than one year of actual experience. Some of the things that might seem obvious to some might not be so for me, thanks!
It's just some kid sharing a mistake they made and owning up. Ease up on the "LOL what an idiot" attitude
And yes - I've literally done this exact same error (with TB of video data!). Spending the following week remediating all of that data loss was a great lesson in patience and attention to detail. :-)
OP: If you're ever looking for a job be sure to send me a message. Contact info in profile.
I do think Vimeo was irresponsible in the whine affair though.
I have made a couple of huge ones - luckily I kept my job
But nope, they went right back to scripting and got it done.
I have multiple years of experience than this man and still I could *very* *too* *easily* make a 7Tb mistake (or likely more :P )
I envy those who claim to do no mistakes at all.
So you wrote bad code, didn't test it properly, ran it on production on the Friday before a release and are blaming Vimeo and [name redacted]?
And your resolution was yet another cobbled together script that you probably didn't test?
This isn't a great article to have attached your name to
After deletion, what should he have done? Postpone the go-live? That's often not a a cost-effective option. As for a risk-analysis the worst what could happen was deletion of the remaining videos. I don't think that that makes big difference in this situation. And to do the right thing, you have to have the infrastructure in place, if you are in a hurry. I doubt that's the case for a 10 heads shop.
To me this indicates intelligence, competence, integrity, grit and generosity. TechnicL proficiency is much easier to come by than integrity, grit and generosity. I would trust the author to deliver on commitments.
I did a similar thing ~20 years ago when I first started my career, accidentally deleting a production database because I thought I was working on the test database.
I owned it, learned lessons from it, and it's never happened again.
‘Judgment comes from experience, and experience comes from poor judgment.’
:-)
The article hardly comes across as 'blaming' them for the core issue but they were definitely not helpful.
A million times better than your comment.
Also, would you rather everyone only ever posted about all the times they were successful?
(We do this sort of thing to protect users, usually as the result of an emailed request, and you can tell when we've done it because of the word 'redacted' in square brackets.)
Well that's a brave move...