A baby saved ‘Toy Story 2’ from near complete deletion (2021)
insidethemagic.net
insidethemagic.net
> The virus, once activated, propagated within just seven minutes. And for Maersk, as with other organizations, the level of destruction was enormous.
> As well as affecting print and file-sharing capabilities, laptops, apps and servers, NotPetya severely damaged both Maersk’s implementations of DHCP (Dynamic Host Configuration Protocol) and Active Directory, which, respectively, allow devices to participate in networks by allocating IP addresses and provide a directory lookup for valid addresses. Along with those, the technology controlling its access to cloud services was also damaged and became unstable while its service bus was completely lost.
> Banks is candid about the breadth of the impact: “There was 100% destruction of anything based on Microsoft that was attached to the network.”
> In a stroke of luck, they were able to retrieve an undamaged copy of its Active Directory from the Maersk office in Nigeria. It had been unaffected by NotPetya purely because of a power outage in Lagos that had taken the service offline while the virus was spreading. A rather nervous local Maersk employee was tasked with transporting the crucial data to Denmark.
What this says to me is that multiple offline backups (or failing that, copies) will save your bacon some day.
[1] https://www.i-cio.com/management/insight/item/maersk-springi...
TheBadGuy changes the encryption key used to make the backups, then waits %long enough%, you’re still boned.
But yes, this is hard to get right as I hinted by the "but not impossible" quip.
That will help.
https://en.wikipedia.org/wiki/Iron_Mountain_(company)#Data_l...
https://www.nytimes.com/2019/07/19/business/safe-deposit-box...
https://www.nbclosangeles.com/news/safe-deposit-box-theft-mi...
https://www.dailyrecord.co.uk/news/scottish-news/widow-heart...
Encrypting the complete backup is working against the purpose of backups, making it harder for a small, if existent, gain in security
And picking and choosing is a recipe for disaster when something inevitably slips through and is leaked. Encrypt it all from orbit and let god sort it out.
A financial org I worked for back in the day planned on slagging hardware in the event of a coloc breach of any sort including fire response. How else could you be sure?
I guess you could drop an emp pulse on my house and fry the drive, but absent that it's safe.
And for an org, there's still attacks in both cases. Do your backups get requested through an automated system? Well you've just made a very slow tape library with people rather than robotics. Are the tapes/USB sticks encrypted and the encryption keys stored in an online system? Then the hackers have a mechanism to deny you access to the contents of the tapes, which is probably their goal anyway.
Yes a USB drive with a text file is pretty safe from such attacks, but it also doesn't scale to the backup concerns of even relatively small orgs.
A drive in a closet doesn't get you there. The nigeria office offline due to power does.
I feel good with S3 with an object lock and retention policy for backups.
Unless you're practicing continuous restoration to ensure that your backups are actually backing up what you need, the odds that you got everything critical are not good.
I think this illustrates where folks can go wrong.
Many folks and business TODAY are backing up to NAS / RAID arrays etc. Many are easily corruptible by ransomware. We see this REPEATEDLY with many infections of large orgs.
The basic story here is that something not editable has great value even if not perfectly current. If you do a diff every 4 hours you have covered a LOT of ground already.
I've sat through continuous DR architecturing discussions - the massive cost / complexity / scale can just be out of control for many users.
In many use cases Active Directory for example behind by 4 hours is not game ending. With notice to users that recent edits may not be captured you may be able to get back up and running pretty quickly.
There are shades of gray, here: continuous restoration is a best practice, but backups which are tested even once, and perhaps incompletely, are better than backups which are never tested at all.
Verifying that restoration is possible once is better than not testing restoration at all. As for verifying restoration on an ongoing basis, if the attacker has root I suppose we can argue about the cost/benefit ratio but continuous backup validation doesn't hurt.
Where did I say that I don't test restores? I think you are misunderstanding the S3 object lock feature. With S3 object lock, you write your backup, then can read to verify or do a test recovery, but you cannot modify or delete for the retention period.
In the package I use it's called Veeam Ready-Object with Immutability and there are now a fair number of solutions that meet that target.
From the start, I've been trying to add the qualifier that regardless of whether you throw your backup into a USB drive or an S3 bucket, if you don't test restoration, your backups are probably incomplete.
"OK, but how about at least testing that that it's actually feasible to restore from that S3 bucket of yours?"
These comments have been both off topic and unhelpful. I gave an example of S3 with object lock (or veam equivalents) that I liked - to be helpful to others exploring this question.
Continuous type services are what this story is about - how things that are disconnected / not continuous may fare better in a hack / ransomware situation. Lecturing me that I need "continuous restoration" and or making random claims that I don't test restoration (huh?) is weird.
If I was to give some feedback -> focus on one thing. If you don't like S3 with object lock, say that. If you want to add some comment about restorations from S3 with object lock being tricky for some reason, say it and explain it. If you have something about "continuous restoration" helping in malware hacks explain why. In many cases this can pass corrupted data over to your backup site if you are doing HA/DR work where continuous stuff is most common.
The state of industry with regards to backups is wretched. If backups are done at all, they are often not tested — people just copy a bunch of files onto a RAID, or a USB drive, or an S3 bucket, or whatever, but never test restoration because that takes extra effort and may be inconvenient unless the system has been architected from day one to enable seamless restoration of live systems without obvious interruptions in service.
Specifically with regards to ransomware hacks...
The first thing that an attacker who gains root must do is corrupt production data. If there are no backups being created, or if the backups don't actually contain all critical data, then the hack is complete.
If there are backups being performed, then the attacker must also hack the backup system sufficiently that restoration from available backup data is infeasible (or at least more costly than just paying the ransom). This adds an addiional layer of complexity to the attack.
if restoration is being tested on an ongoing basis, then the attacker, in addition to corrupting the production data, must also spoof monitoring so that the the corruption of live and backup data doesn't get caught and interrupted before the corruption finishes. This adds yet another layer of complexity to the attack.
A sufficiently competent and thorough attacker may overcome all these difficulties, but criminals are not always perfectly competent — they don't need to be because there are so many target organizations that are not competent either.
It seems that your specific organization does thoroughly test backups to prove that restoration actually works as intended. Kudos! However, I wasn't evaluating your specific case — that would have been weird and inappropriate as you were just giving a (helpful!) example, not laying out a comprehensive system design and asking for review. My point was to add on the condition that in principle backed up data must be not only be 1) current and 2) not corrupted, but also 3) complete. And the way to prove completeness of backups is to test them, ideally on an ongoing basis.
This is why distributed version control system like Git are so nice.
If an attacker get root they can theoretically disable the continuous restoration, but the more moving parts they have to tweak the more likely that they'll be noticed while the org still has access to decent backups.
> Jacob admits that about 10% of the film’s files were never retrieved and are gone forever, but they were able to make the film work without the scenes.
Also, this seems to be a second hand rereporting, the linked article has a little more detail (https://thenextweb.com/news/how-pixars-toy-story-2-was-delet...)
It also has the real beef of the story:
> Pixar, at the time, did not continuously test their backups. This is where the trouble started, because the backups were stored on a tape drive and as the files hit 4 gigabytes in size, the maximum size of the file was met. The error log, which would have told the system administrators about the full drive, was also located on the full volume, and was zero bytes in size.
Offering Continuous Restoration as a feature should theoretically be a way of differentiating a development tool in the software marketplace. But short-termers hold power in so many organizations that it's not clear to me whether it would actually win. Even if ransomware kills the company, executives still keep their money.
It's not enough to confirm that the backup has actually been created you need also to have the procedures necessary to do the restore and the personnel and hardware to do it.
He wrote that everybody in the team had forgotten about Sussman's copy and she is the one that told in a meeting the copy existed, to the astonishment of everyone else. The van that went to pick up the computer on Sussman's house was equipped with pillows because they were scared even from the road's vibrations.
edit: Now Im really curious what this version was even like considering it was never released
Roughly speaking the book validates the whole idea of "to build awesome stuff, hire the most creative and talented people you can find and create an environment for their imagination to run free". If you're interested in this perspective, I do very much recommend.
But, obviously, he doesn't address issues like John Lasseter running free with sexual harassment and other uncomfortable issues.
https://www.courthousenews.com/disney-to-settle-animators-wa...
If you are trying to read it purely for business advice or anything of that nature, you might find it a bit eh.
It is very heavy on telling an interesting story, which should explain rather well why it might not be the best for the purpose of being an educational material. But if you want an interesting delivery of the story of Pixar, along with a look at behind the scenes in the industry (e.g., Steve Jobs was featured in there as one of the important personalities, given he was heavily involved with Pixar) and behind some of the decision-making/philosophy of Pixar as a company as it was growing, it is a great and entertaining read.
Quick reference to lots more detail posted here -> https://www.quora.com/Did-Pixar-accidentally-delete-Toy-Stor...
And then, some months later, Pixar rewrote the film from almost the ground up, and we made Toy Story 2 again. That rewritten film was the one you saw in theatres and that you can watch now on Blu-Ray.Pixar did recover the data from a home office setup and complete the movie, but Pixar is so perfectionist about story-telling that the creative heads decided the end result wasn't good enough.
The story was rewritten and the movie changed so much that almost none of the recovered data was in the version that was finally released.
Your standard "shot" will contain 10s or hundreds of versions of animation, lighting, texturing, modelling & lighting. Not to mentions compositing. They are often interdependent too.
A movie like shrek will have literal billions of versions of $things.
> Does anybody knows how that works?
This changes by company, but everyone tends to have an asset database. This is a thing that (should) give a UUID to each thing an artist is working on, and a unique location where it lives (that was/is an NFS file share, although covid might have changed that with WFH)
Where it differs from git et al is normally the folder tree is navigable by humans as well so the path for your "asset" will looks something like this:
/share/showname/shot_number/lighting/v33/
There are caveats, something like a model of a car, or anything else that gets re-used is put somewhere else in the "showname" folder.
Now, this is all very well, but how do artists(yes real non linux savvy artists) navigate and use this?
Thats where the "Pipeline" comes it. They will make use of the python API of the main programs that are used on the show. So when the coordinator assigns a shot to an artist, the artist presses a button saying "load shot" and the API will pull the correct paths, notify the coordination system (something like ftrack or shotgun) and open up the software they normally use (maya, zbrush, mari, nuke, etc etc) with the shot loaded.
Once the artist is happy, they'll hit publish.
The underlying system does the work of creating the new directory, copying the data and letting the rest of the artists know that there are new assets to be pulled into their scene as well.
Then there are backups. Again this is company dependent. Some places rely on hardware to figure it out. As in, they have a huge single isilon cluster, and they hook up the nearline and tape system to it and say: every hour sync changes to the nearline. every night, stream those to tape.
Others have wrapped /bin/rm to make sure that it just moves the directory, rather than actually deletes things.
Some companies have a massive nearline system that does a moving window type sync, so you have a 12 hourlys, 7 dailies and 1 monthly online at once. The rest is on tape. The bigger the company, the more often the fuckup, the better the backups are tested.
Sadly my animation knowledge has left me since I left the studio working gig three years ago. But we did create content for netflix and many animators came up sheepishly because they deleted a folder for it to be restored. It's not as uncommon as you think.
FWIW, the live archive/backup server was called Dumbo. Which were 3x 4u Supermacho chassis with over 1.2PB in drives served over 10Gbit to each workstation connected to 1Gbit running CentOS 5. I dropped the new chassis while racking once and is partly the reason to why I lost my job :/
That said, Toy Story 2 was developed in the late 90s, and while Perforce existed then I don't know how popular it was.
"Did Pixar accidentally delete Toy Story 2 during production?", Quora answer by Oren Jacob: https://www.quora.com/Pixar-company/Did-Pixar-accidentally-d...
Once those negotiations were over, Pixar started work on a sequel, and this version was never released.
The command that had been run was most likely ‘rm -r -f *’, which—roughly speaking—commands the system to begin removing every file below the current directory. This is commonly used to clear out a subset of unwanted files.
https://thenextweb.com/news/how-pixars-toy-story-2-was-delet...
I occasionally worked from home in 1998. The biggest impediment was network speed -- at work we had 10mbps and sometimes even 100mpbs. But at home all we had was dail up or if you were really lucky, DSL, which was 0.384mbps, so 30 times slower, and also latency was a lot higher.
So you were very limited in what you could do. Mostly it was just email and remote access ssh with tmux, where you sometimes had to deal with really high latency.
It only happened in this instance because the person in question had a newborn baby and I guess maternal leave wasn't a thing at Pixar back then.
The opposite. After she finished mat leave she was allowed the special exemption to work from home so she could still spend time with her baby.
We all treat backups like open loop: billion-dollar project? Meh, just throw it over the wall and hope it lands somewhere.
I harp on this every time it comes up. Ugh.
Kids these days who have branch protection turned on won't get this joke.
Whether this is done in a chain (backing up the backups) or as multiple transfers from the sources is an implementation decision: ideally the latter as otherwise a break in the first backup system can stop the second updating, but this may require notably more resources at the source end so compete for those resources with production systems so the former isn't always a worse plan.
More importantly: test your backups. Preferably in an automated manner to increase the change the tests are not skipped, but have humans responsible for checking the results. Ideally by doing a full restore onto alternate kit, but this is often not really practical. A backup of a backup is no extra help if the first backup is broken in a way you have not yet detected.
3 backups
2 media formats
1 off-site
LGR and RMC have both done videos on the SGI Octane and are worth watching: https://www.youtube.com/watch?v=ZDxLa6P6exc, https://www.youtube.com/watch?v=VHsA8iq4N0s
Edit: It was Retro Man Cave, not Adrian's Digital Basement as originally written.
Oren, who worked at Pixar recounts the story, and said “She has an SGI machine (Indigo? Indigo 2?) Those were the same machines we had at all our desktops to run the animation system and work on the film, which is what she was doing. Yes, it was against the rules, but we did it anyway, and it saved the movie in the end.”
https://www.quora.com/Did-Pixar-accidentally-delete-Toy-Stor...