Dropbox confirms that a bug within Selective Sync may have caused data loss
gist.githubusercontent.com
gist.githubusercontent.com
We received several reports from users who used a Dropbox feature called Selective Sync and couldn’t locate certain files they’d saved in Dropbox.
When we took a closer look, we discovered that older versions of the Dropbox client had introduced an issue affecting a small number of users whose Dropbox application shut down or restarted while users were applying Selective Sync settings.
In light of all of this, we've taken the following steps to ensure the Selective Sync bug won’t affect anyone else going forward:
1) we've patched our desktop client so this issue doesn't exist in Dropbox anymore;
2) we've made sure all our users are running an updated version of the Dropbox client; and
3) we've retired all affected versions of the Dropbox client so no one can use them.
We've also put additional testing in place to prevent this from happening in the future.
We’re very sorry about this issue and the trouble it might have caused. We’ll keep doing our best to ensure our users' data is always safe and available to them.We received several reports from users who used a Dropbox feature called Selective Sync and couldn’t locate certain files they’d saved in Dropbox.
When we took a closer look, we discovered that older versions of the Dropbox client had introduced an issue affecting a small number of users whose Dropbox application shut down or restarted while users were applying Selective Sync settings.
In light of all of this, we've taken the following steps to ensure the Selective Sync bug won’t affect anyone else going forward:
1) we've patched our desktop client so this issue doesn't exist in Dropbox anymore; 2) we've made sure all our users are running an updated version of the Dropbox client; and 3) we've retired all affected versions of the Dropbox client so no one can use them.
We've also put additional testing in place to prevent this from happening in the future.
We’re very sorry about this issue and the trouble it might have caused. We’ll keep doing our best to ensure our users' data is always safe and available to them.
[1] http://www.bbc.co.uk/ethics/introduction/duty_1.shtml
(I happen to agree with HN’s ethical posisition, but I invite you to look at the section “Bad points of duty-based ethics” in the link).
It's not really is it? It's just something you believe and are justifying with obscure ethical arguments. Thanks for the reading :-D
I have an older laptop that I turned on. It was a work laptop a few years ago, linked to my dropbox account, etc. Since then I had added a bunch of things like a bunch of git repos to a folder included in dropbox.
I turned on that laptop and Dropbox started using 100% cpu after a few minutes. Then the fan kicked on and it was annoyingly loud so I looked at dropbox and saw it was chugging along in the repos directory. I went ahead and clicked on selective sync, unchecked repos, and left it alone for about 5 minutes.
It was still 100% cpu, so I killed the dropbox task and restarted it.
Minutes later, on another machine, I went to fetch from one of the repos and it had a gnarly error. So I went about investigating.
I found my way to the dropbox events tab (on the website - the desktop client doesn't have this feature) and saw an event where dropbox decided to delete 7,800 files.
I submitted a support request, but before they responded I had figured out it was (mostly) in the repos directory, which I fixed by simply deleting the repos and pulling from one of my servers.
Anyways. There's my real world run in with this bug.
Most medium-sized and large companies refuse to touch Dropbox due to these reasons, especially in the financial and medical space.
That said, I still employ it for personal use and like the product in general.
Nope, highly custom process that involves librsync very little. The sophistication of what has to be done to solve this problem well would probably surprise you.
A good example use of it is the rdiff tool that allows you to do exactly what I said before. A better real-world example would be duplicity [1] or rdiff-backup [2] that use librsync to de-duplicate backups, without needing access to the whole previous value, only a small signature of it.
[0] https://github.com/librsync/librsync
The very best? I use OneDrive across all of my Windows machines and I don't even notice it exists; never had any problems. I just access all my files everywhere. If you buy a windows phone you even get a decent amount of space for free (15GB). (Though I subscribe to Office 365 so I have virtually unlimited space.)
The bottom line is that errors happen. You should prepare for that and make backups.
Also, Dropbox are still among the very best when it comes to syncing. Many useful synchronization features are implemented by Dropbox, but not the competition. E.g., features that most competitors do not have:
- Modifying a large file on Dropbox will only resync modified chunks.
- DropBox avoids re-uploads, both when uploading identical files and moving files around:
http://macography.net/2013/05/speed-test-dropbox-google-driv...
- Dropbox does LAN sync. If a machine has to download a large file and another machine on the network has the same file, chunks are provided peer to peer. This makes using large files on multiple machines or in a team much faster.
- Dropbox does streaming sync. A machine can already download chunks when another machine is still uploading:
https://blog.dropbox.com/2014/07/introducing-streaming-sync-...
Sure, OneDrive and Google Drive do have many useful functions that Dropbox does not have, such as including complete office suites. But for the original task, file syncing, Dropbox is still pretty much unbeaten.
It's certain that Dropbox has a high quality syncing service, but there are other factors. Think, for example, how this case was handled: a fault in their core product, a breach of user trust in their service, and they understood that it needed more than a technical solution. None of this was part of their core sync reliability: it was part of a more broad quality, which is closer to their true reason for success.
I did not say anything about their reasons for success. Only what the technical advantages are compared to some of the other file sync services.
It really isn't that difficult a problem,
Difficult enough that some of its useful features are not matched by other services yet.
I have an rsync script that has been syncing my files reliably to an offsite location for 6 years.
That's great. But that is one-way sync and not something my parents could use. Dropbox is successful because they made sync technology that is relatively flawless to the average user. Also, there is a network effect.
In the longer term, it will be interesting to see if they survive, since Microsoft and Google have been undercutting prices heavily, and as far as I know there is no online Office suite on the horizon (only Microsoft Office integration for business users).
Doesn't the file have to exist on Dropbox's servers before it can be synced to another computer on the same LAN? The last time I looked into it, this was the case.
- There is no Linux/OSX client unlike DropBox (the OSX client only works for OneDrive and not OneDrive for business) so it's unusable with servers or environments with a lot of OSX machines. (so most enterprise usage)
- There is a list of approved file types and if your file is not on the list, it just refuses to sync it. This is really annoying because I need to create zip files all the time to bypass this bug.
- OneDrive modifies certain file types (like word document but also others) to add metadata on it so you never know if the file you are getting is exactly the same as the one you synchronized.
- We experienced bugs in the permission system which destroyed couple of files (thankfully we had backups).
The only positive thing with OneDrive is that it's integrated with Office 365 (it's the equivalent to Google Drive for Gmail) so you can preview files directly within your web browser on Office 365 (when it works because sometimes you just can't).
I would have a choice, I would return to Dropbox without any hesitation.
> This problem occurred when the Dropbox desktop application shut down or restarted while users were applying Selective Sync settings.
So, you must be in the midst of applying selective sync settings while the app shuts down or restarts. Although I'm not sure what they mean when they say, "while users were applying selective sync settings." I'm not sure if this means:
A) Changes made in the selection dialog box, but not committed (by clicking OK).
or
B) Changes committed, but still syncing.
The former is an edge case, the later, not so much.
Also, if I understand correctly, Google Drive has a better policy here: removed files are just placed in the trash until you remove them from the trash. Of course, trash takes space up as well, but it protects better against such cases.
I guess Dropbox is trying to maximize its profits with its 'remove after 30 days' policy.
As it turns out, I have other backups of most of the files, and the rest of them weren't important. So I was lucky. Still, my confidence in the product is unlikely to recover.
I want to note that I had been aware of the "dropbox is not backup" chorus, but that argument usually is just "sync is not backup", which is sort of obvious. The packrat feature pretty much addressed this issue, so dropbox with packrat WAS a backup solution. So the lesson here is never to rely on any ONE backup provider.
If dropbox kept backups like you're suggesting, people would be complaining about how you can't delete damning files from them and that law enforcement was abusing this.
A good backup policy uses a mixture of onsite and offsite, and Dropbox can be a (convenient) part of that. E.g., I store (non-sensitive) files in Dropbox, which gives me a certain period of undelete possibilities. My Dropbox folder is backed up on a local time machine backup. Critical parts of my Dropbox are also backed up using tarsnap, etc.
A good backup policy diversifies, and Dropbox can be part of that.
If you want to be picky you shouldn't rely on backups unless you have multiple independent backup systems, at least one offsite, at least one offline.
- convenient enough that you do it without thinking about it.
- technically as simple as possible, so it's easy to understand and review.
- secure.
Dropbox fullfills the first point, but not the second, and the third is debatable. Spideroak as a counterexample is just as convenient, has a pure incremental backup mode and is client-side encrypted, the gold standard of security.
Also from what I've seen spideroak is significantly more complex than dropbox.
Short story: if you plan to develop a sync product from scratch, be prepared to spend at least 2 years or hire core developers from Dropbox sync team. Eve now dropbox has issues with handling large number of small files. Try to stuff 200000 to 300000 files and see how it works.
The problem is that when I clicked 'restore all' from within the subfolder, Dropbox restored all 12,000 files rather than just the files within the folder.
Note to DB's UX team: when you place a Restore All checkbox above the lefthand file selection column, it means 'select and restore all files on the page', not 'lift the roof off my house and dump in all the shit I spent months decluttering.'
One of my greatest fears is that thousands of files might disappear without me noticing for years.
I use selective sync and twice I was looking for something that has disappeared and I have to restore it. I assumed maybe my wife accidentally deleted some files, but maybe it was dropbox?
What is the solution to this anxiety?
So even if you store data in Dropbox - it is smart to have one extra copy in some other cloud storage. Like Google Drive. Or Box. Or Egnyte. So if data is deleted in Dropbox (accidentally, maliciously, or due to a bug) you can restore it from other cloud.
Of course, cloudHQ is the system which can do that: http://chq.io/hnsc