What Happened: Adobe Creative Cloud Update Bug
backblaze.com
backblaze.com
Back when Flash was still a thing, they got so low as to try to sneak in more of their products when I just wanted to update Flash.
It's crappy behavior and should be called out for being what it is. We are too quiet about it and we got so used to taking crap from software producers that many people don't even complain.
http://community.eveonline.com/news/dev-blogs/about-the-boot...
That said I completely agree Adobe's software is disgustingly invasive, just not for this particular reason.
Honestly (and imnsho) the prevalence of that sort of thing makes it even less excusable for a large company like Adobe.
> Line 468: rm -rf "$STEAMROOT/"*
Without a check for $STEAMROOT validity. I'm not sure you couldn't sue Valve for the incredible recklessness. That is just beyond the pale.
No. It's an idiot mistake. It's an idiot mistake that should have been caught by code review. It should have been caught by release engineering. It never should be made in the first place because deleting files that your software didn't create in the first place is asinine.
If your software does this you should be fired.
End users should be free to install software without having to sort through their backup/recovery process.
Which is the problem: if you hire a builder or a plumber and they do something stupid which causes you significant harm, you can sue them. And you'll probably win.
If you buy a car and it's defective and causes you significant harm, you'll likely get compensation from a class action suit.
How did destructive software errors get an exemption from the legal principle of accountability? (I know the EULAs say whatever, but that doesn't mean they can't be challenged in court.)
Adobe, MS, Apple and the rest don't need to care about software quality because they don't need to worry about expensive legal action by customers.
Stupid crap like this will continue to happen until that changes.
Look, bugs in software are as old as software itself, comparing software to a builder is not helpful because it is simply not the same thing at all - software is a lot more abstract and the effects of doing certain things are far less measurable at composition time - also builders have several thousands of years of tradition, programmers have less than a hundred.
I don't take kindly to my system being commandeered, so I cancelled my CC subscription, removed their bloated, slimy POS software and bought Affinity Photo.
Have you ever attempted to analyze the periodic data transfer that happens from your computer to Adobe servers when you have one of their products running? They log every action that you perform: from opening a file to clicking a menu item.
Hint: The "encryption" scheme employed to _encrypt_ the user actions is a substitution cipher.
Under no situation should any software randomly delete folders and files it doesn't own, period, specially something so widely used by professionals. This is not just a little bug on a rarely used OS platform not limited to people using backblaze on the Mac, it should have been a show-stopper, and should have never left the test environment (do they even have one?). It reflects on the abysmal quality of Adobe's software and their QA.
There's no reason Adobe should even touch the root directory "/" on the Mac, leave alone delete anything from there. Why does Adobe CC need root permissions to install on the Mac in the first place?
I hope the people responsible for this disaster, their managers and the QA team were fired. In my business, this sort of mess can sink our company.
I'm just glad they didn't migrate to messing around with EFI variables for their copy protection.
1: https://bugs.launchpad.net/ubuntu/+source/grub2/+bug/441941/...
Doesn't really excuse Adobe.
The only thing I could find from Adobe was the basic release notes, which really don't say much. Has anyone seen more details on how this happened? If it is licensing, I doubt they'd admit it.
It reflects on the abysmal quality of Adobe's software
and their QA. [...] I hope the people responsible for
this disaster, their managers and the QA team were fired.
Eh, I can kind of see how it might happen myself. Remember that time Valve had a bug in Steam for Linux that would wipe the user's hard drive? [1]It's easy to imagine how this sort of thing would happen - manual test guys don't have hidden folders in their root directories because their boxes are locked down by corporate IT and they can't create any folders in the root directory. Developers don't have them because they understand the unix way and know they shouldn't be putting files there. CI servers don't have such folders because, well, why would they? None of the test environments include the specific condition needed to trigger this bug - a hidden folder in the root directory - so the bug never appears in testing.
Is it bad and serious? Absolutely. But I don't think the fact testing missed this is a sign of testing incompetence - you could have cutting edge testing following every industry best practice and still miss this bug.
[1] https://github.com/ValveSoftware/steam-for-linux/issues/3671
It really isn't hard to have a company policy about nounset, errexit, and, in the extreme, noclobber.
I am a hardcore fan of Macromedia Fireworks, and Adobe releases caused so much problems that I decided to use Macromedia Fireworks in the literal sense (ie: in my current machine, I have Macromedia Fireworks 8 installed).
When Adobe bought Macromedia, I installed their Adobe Fireworks thinking it would be just an updated Fireworks... it was right, and wrong, right because it barely had any new features or bug fixes, wrong because it install for some reason needed "adobe common files" anyway, and installed 3gb of crap on my disc (when discs were still 20gb big, thus it was a huge waste of space), AND somehow it conflicted with their own Adobe Reader and made it stop working.
When Adobe Fireworks CS3 was released, I was in university, I was excited to try it... and not only found again it had almost no new features, old features don't worked correctly either, most aggravatingly, the selection boxes would always render in the wrong place, and not just 1 pixel off, they would be completely off, when I tried to input text for example, the text box appeared 300 pixels away from the actual text.
When a friend installed it on his computer, the computer started misbehaving.
And the problems kept coming, to the point that eventually I gave up, I have no idea what version Fireworks is now, because I decided when CS6 came out, to install on my machine Macromedia Fireworks, and not install any Adobe product on my computer (eventually I found it was hard to avoid Acrobat Reader, thus now I install it, but only that).
EDIT: Also I always disable Adobe Updater, the first thing I do after installing any Adobe product on any computer for any reason, is shut down the updater process, and forbid it from running.
The bright side is that they haven't completely ruined it. I haven't noticed the text box problem, but I skipped many versions before I switched over to OSX and had to get the new one.
Counter-point:
* Let's be honest, Linux is not a first class gaming platform. Linux is not going to be Steam's #1 platform. Whereas here it's an Adobe product on OSX.
* The Steam bug was triggered by the user moving ~/.local/share/steam. That's not something that users will be doing regularly. Whereas this Adobe bug wasn't caused by any uncommon behaviour.
> Developers don't have them because they understand the unix way and know they shouldn't be putting files there.
Developers are very definitly going to have dot-files. .ssh, .vim, .bashrc, .local, .gnome, .mozilla, etc. etc. If something wiped my SSH private keys on my desktop, that would be a massive problem.
Developers are very definitly going to have dot-files. .ssh
The article only mentions the system root directory, never the user's home folder. If you're storing your ssh keys in your system root directory, you're doing something very unusual.The variable was empty when a user moved steamroot but that would have been fine with proper input checking - also there were probably other scenarios where that failure soul result in catastrophe too.
http://m.theregister.co.uk/2015/01/17/scary_code_of_the_week...
This is what QA is for. It should be thorough, more thorough than you imagine and performed by someone other than who wrote the code.
Exactly! Any programmer worth their salt should be able to see how it might happen... and prevent it!
> you could have cutting edge testing following every industry best practice and still miss this bug
I disagree. If you are following industry best practice (checking files/directories before you delete them) you would not have this bug.
There is no reason why an application should touch anything outside its own sandbox (unless requested by a user via a open/save dialog). Incidents like this show why developers should distribute their Mac applications sandboxed by default.
I hope something very different: I hope they're required to write a thorough post-mortem analysis of what went wrong, and what in their development & testing process led to the failure. And then required to implement improvements to both that would prevent this class of bugs from happening again.
Firing people for a mistake is rarely a good idea. It makes them not want to admit mistakes. Firing them for making the same mistake several times and not learning for it -- fire away.
c.f., http://itrevolution.com/uncovering-the-devops-improvement-pr... and the "blameless post-mortem"
If code like this escapes code review and testing and wacks actual customer data, firing is an appropriate response.
It's never been triggered except by my unit tests, but I'm sure glad it's there.
(a): Pull out first directory found in readdir("/"), system("rm -rf ${dirname}") ( in no real pseudocode).
(b) removeCacheDirs(); --> recursiveDelete(getLocalCacheDirList()); --> ReadConfig(GetUserPrefLocation(CACHE_DIR)), .... <some horribly nested mess of functionality created by 10 different programmers over as many years, where somehow the code misbehaves under a specific set of environment attributes, etc.>
Sure - (a) would be awesomely stupid, but it's probably not what happened. Oh, it might be. But (b) is more likely, and reflects as much an organizational and dev process mess as anything. And if (a) happened, you need to ask yourself how: Why was someone with so little clue allowed to implement something like this with no code review / why did it bypass code review? What's the longer-term policy and mechanism to prevent this?
When a three year old shoots someone with a gun, do you put the three year old in jail, or do you figure out how the $*! they got a gun in the first place?
And it's more than fair for the lead who approved the code change, and perhaps even the lead who approved the entire design to have their jobs on the line for this.
This is a actually a rather typical response, made with full awareness of the consequences of an action by people not under the same constraints as the participants in an incident. Diane Vaughan (an expert on risk, causes of incidents, and how organizations come to become error-prone) terms these actions 'ritualistic':
> Causes must be identified, and in order to move forward, the organization must appear to have resolved the problem. The direct benefit of identifying the cause of an accident as Operator Error is that the individual operator is the target of change. The responsible party can be transferred, demoted, retrained, or fired, all of which obscure flaws in the organization that may have contributed to the mistakes made by the individual who is making judgments about risk. The dark side of this ritualistic practice is that organizational and institutional sources of risk and error are not systematically subject either to investigation or other technologies of control.
To make it even better, you're now throwing away all the people most experienced in the conditions that lead to a specific incident. Unless very specific conditions are true, this is a sub-optimal recommendation. The evidence seems to be against building a culture of fear.
Yes, call out the lynch mob.
Coincidentally, this bug was released a day before a holiday weekend. It's practically guaranteed that people were crunched on their deadlines and faced with either a) working all weekend b) getting reprimanded for failing to meet their deadlines or c) calling it good, releasing with inadequate review, and enjoying the long weekend (or so they thought). Firing minions and writing technical post-mortems won't help with creating a sane work environment.
Surely there are others who didn't even know they lost data; when asked about how they'll notify those people, Adobe responds "We have advised customers to take local backup of their files just to be on a safe-side."
Some talk coming from a company that deleted user files with wanton disregard. I'm aware that legal culpability is a complicated topic, but Adobe is doing themselves no favors here.
[0] https://forums.adobe.com/thread/2089459?start=0&tstart=0
Absolutely 100% unacceptable. Massive props to the BackBlaze team for handling Adobe's shitstorm so well.
Meanwhile, I'd love to see the postmortem from Adobe. What file(s) were they expecting to delete and why did they feel it prudent to not do a strict/stricter check for the files/directories to delete? Presumably some sort of "adobe" directory, but I would hope to at least find out that they couldn't reliably expect the folder to be a specific name and the issue was that their searching was just too loose.
For example:
ls / | head -n 1 | xargs rm -rf rm -rf /.adobe
?My point being that things indicate a more interesting problem they were trying to solve than just deleting a directory. What that actually was, and how that lead to such an awful solution, would be interesting to know.
I'm the guy who decided to put things in a folder at the root level of drive. Originally this was a system for EXTERNAL drives that get plugged in then unplugged, then come back later. We need to know what state the backup is in related to that drive, and if the customer has several external drives that come and go we need a unique ID for each one. So I created a top level folder on the external drive called ".bzvol" (hidden) and then placed two files inside of it, a README that explained what the folder was and who created it and why you should not delete it, and a 100 byte XML file that had the drive's unique id AND ALSO the identifier of the backup that owns this drive (some customers have two computers both backed up by Backblaze and they carry a hard drive between them).
There are many designs that would work, in retrospect a better design might be to have used the drive's internal serial number (which is globally unique) and maybe a little mapping database either stored on the Backblaze website or somewhere down under /usr/local or /Library that maps the globally unique drive back to the backup it is associated with.
That's the problem with developing software. After you have worked on a problem for 8 solid years, you are finally qualified to BEGIN working on it and should rewrite it from scratch. :-)
- For production code, categorically reject any language that will happily treat an unset variable as an empty string. Or, if possible, configure the language implementation to run in a stricter mode that fails fast instead. Question: Is that possible with shell or PHP?
- Favor manipulation of data structures over concatenation of strings. See Glyph Lefkowitz's blog post "Data In, Garbage Out" [1]. When manipulation of strings is practically unavoidable, as with filesystem paths, use a well-tested library such as Python's os.path module or .NET's System.IO.Path class.
- Embrace the principle of least authority at every level. For example, when developing for the Mac, work within the Mac app sandbox as much as possible.
Any other ideas on how to prevent this kind of catastrophic bug?
[1]: https://glyph.twistedmatrix.com/2008/06/data-in-garbage-out....
#!/bin/bash
set -eu
<actual work>
Explained:
-e Exit immediately if a simple command exits with a non-zero status, unless
the command that fails is part of an until or while loop, part of an
if statement, part of a && or || list, or if the command's return status
is being inverted using !. -o errexit
-u Treat unset variables as an error when performing
parameter expansion. An error message will be written
to the standard error, and a non-interactive shell will exit. -o nounset
Worth a read: http://ss64.com/bash/set.html "${VAR:?Error message goes here}"
Will exit2 if VAR is unset or the empty string.Or, which I often use, build the path you want to delete, call realpath on it - then check if your base directory you never ever have any business leaving is a prefix string of the realpath resolved path. This still allows to symlink stuff into subdirectories but will catch if something resolves suddenly all over the place.
A long time ago, we had a Java web application that we'd deploy, and after a couple of days would stop working. We'd investigate and find it hadn't been deployed properly, yell at the person who installed it and go on our way.
Until the day data disappeared. That was weird.
Then we saw files physically disappearing while we had Windows Explorer open in front of us.
So it turned out that a library we were using was using exception handling for flow control, and in a limited set of circumstances this meant a file wasn't closed. That was fine. Usually.
Except that Java 1.4.1_01 had a bug on Windows where if you opened more than 2036 files open THE NEXT FILE THAT WAS OPENED WOULD BE DELETED[1].
I'll never forget the pain.
[1] http://bugs.java.com/bugdatabase/view_bug.do?bug_id=4779905
[1] https://blogs.adobe.com/adobecare/2016/02/12/creative-cloud-...
Such a shame that they're market-leader in so many ways, because once you get passed the pretty uis and impressive features, their software just a train wreck.
It's interesting, also, that Windows' WoW16, and then WoW64, both provide their own levels of filesystem virtualization for "messy" apps... but those same constraints aren't pushed on "native" apps.
I still don't really understand why no OS just virtualizes every app's filesystem, without having to opt into something like sandboxing. It'd actually be able to provide a much nicer programming model, a lot like Plan9: just spew all your program's files into the virtualized equivalents of system directories, because they're directories that are really just for you. No subdirectories; you just put configuration right in /etc, manuals right in /usr/share/doc, etc.
That could then be combined really well with a database-filesystem: going in the file manager to /usr/share/doc would display a "virtual library directory" with virtual subdirectories for each app-container that had made use of the directory. (Or you could skip the virtual subdirectories and get a merged view. Good for e.g. a Fonts directory.)
I was using Backblaze back then and still do.
A great company and well worth supporting.
Thank you for the analysis and sorry for the troubles you encountered due to Adobe -- a company that had multiple "we are doing better for the employee" HR moments which mean "you are screwed" over the last 5-7 years.
Giving random 3rd party applications access to your whole computer feels like a relic from another age.
There only a few ways they could have worked around this, and I don't blame them for their choice:
• Require apps to sandbox themselves to the degree they're able, and then request "entitlements" for anything they need to do outside of the sandbox. Use entitlement list on app submission to guide the app review process.
This is what Apple chose; it's not very costly for them, since they were already doing review on app submission, and the entitlement list can actually make review faster than before, since it gives each review-iteration more of a clear picture of the app's architecture.
• Divide apps into "system" vs. "pure-userland" apps. Run all "pure-userland" apps in containers. Optionally, require that "system" apps are as minimal as possible—more like "system extensions"—and set it up so that most "system apps" end up as "system extension" + "pure-userland app" pairs that use IPC to coordinate, where one has some sort of official UX to trigger the app store to download+install the other. Require, additionally, that the "pure-userland app" still retains some sort of usefulness even without the system-extension.
This is a cool solution, and the one I'd expect Linux (or Android) to use in the future, pairing "app" packages with "system" packages. It can be nearly completely automated with little-to-no review required. But it is very restrictive: either you're containerized or you're not. The components that require elevation can't benefit from isolation at all. There's no partial sandbox in the sense of Chrome's "even if you get write-memory access to the browser, you're not getting anywhere" sandbox.
• Do what iOS does: have only pure-userland "apps", but where some apps can prompt you to install configuration profiles, that then apply a layer of (reversible-by-uninstalling) changes to the OS, which the app can rely on. This is how e.g. the MaaS360 app "takes over" and manages corporate iOS devices.
I could see OSX being written this way. OSX has configuration profiles(!), though they're not used for nearly anything other than enterprise management (and, I think, the recent-ish OSX beta program.) Personally, I think they're a cool system, a bit like if the Windows Registry was composed in terms of immutable, introspectible patch-layers rather than being a single mutable tree. (By analogy: applying a configuration profile is like applying a filter in Flash or Fireworks, while installing a .reg file is like applying a filter in Photoshop.)
The flaw with configuration profiles is that you have to anticipate every way an app might want to modify the system, and provide an API to request that the system modify itself in that manner instead. You might be able to anticipate e.g. MacFUSE's desire to install filesystem drivers, or f.lux's desire to play with screen gamma, or even Synergy's desire to take over your input device handling. But you'll inevitably miss things, like e.g. Dropbox's desire to inject a shared object into the Finder that overrides some of its exported functions with new ones.
>At 12PM I received emails from Adobe’s PR who wanted to make sure we were on the same page with how we were addressing the issue
In a sandbox, by default your software can ONLY write to a specified place like the "/some-random-globally-unique-string/yourapp/yourcrap" folder. Then, guess what: the most that "rm -Rf /" can do is obliterate your own stupid app's files. And, if you further compartmentalize features, it may only damage a segment of your app.
Adobe's still stuck in the decades-ago software development model of "well, guess I need full root permission" with no thought invested beyond that. No wonder Apple further locked down root folders in El Capitán so that you now need to restart from a special partition and essentially invoke "sudo --act-from-god --yes-I-really-know-what-the-hell-I-am-doing --no-really-yes-I-do --yes-really" to remove the protections from most root folders.
Preventative steps: https://backblaze.zendesk.com/entries/98786348--bzvol-is-mis...
> In a small number of cases, the updater may incorrectly remove some files from the system root directory with user writeable permissions.
Ahh, much better.
You're suggesting people should switch to a GNU project program because GNU has a better track record for not doing counterintuitive insane, dangerous bullshit?
`strings` was exploitable. `strings`, for fuck's sake.
And the autotools are still a thing.
The day I assume GNU software is well written is the day GNU HURD ceases to be a punchline.