Dropbox Bug Can Permanently Lose Your Files
konklone.com
konklone.com
It's important to have an automated one-way backup system that you can manually restore from. Something like Tarsnap [1] looks like a really good possibility (I haven't used it myself, but it seems solid)
"Even if your computer has a meltdown, your stuff is always safe in Dropbox and can be restored in a snap.
In fact, if you're using the Dropbox desktop application, your files are backed up several times. The primary copy on your computer's hard drive is synced online and that copy is then backed up again for safety (emphasis mine). If you are using Dropbox to sync files between multiple computers, your files are backed up on those computers as well. If that isn't enough, Dropbox also keeps backups of all of your deleted and changed files too.
...
It's hard to imagine a scenario where Dropbox could lose your files. Hypothetically, let's say a nuclear bomb blows up the data centers where your files are saved. Even then, your files are still safe and sound on your computer and any other computers linked to your Dropbox account."
Clearly, though, Dropbox appears to have lost files (at least if we take it on faith that "these are bugs in Dropbox's syncing logic"), despite the fact that I see no mushroom cloud nearby.
Unless, as happened here, Dropbox erased the file and synced a blank version across all your computers. They they're safe and sound on any computer linked to your dropbox connection that hasn't been connected to the internet since the file got corrupted.
If you catch it in the 30 day window when they keep old versions you're find, but there are files in my Dropbox that I don't use every 30 days.
On the other hand, I don't want to immediately blame Dropbox just yet either. If you backup garbage (say, because you have disk corruption), then you can't blame Dropbox for backing up exactly what you told it to.
And Dropbox does offer a premium Packrat service if you want file history indefinitely. Perhaps the user can be blamed for assuming that he/she would only need 30 days of history, but this is really contingent on who caused the corruption to happen in the first place -- and that's unknown at the moment. [1]
A person running Time Machine (or similar incremental backup system) is a lot safer from this sort of problem than a free tier Dropbox user. Free Dropbox is better than nothing, but people can't keep assuming their files are safe because "they're in the cloud" and get synced to a few places.
Please stop. Right there. Stop believing some vendor marketing blindly (even if you do come to another conclusion later), stop reassuring other people who do so and stop calling a sync a backup. There is a very important distinction: Sync has mechanisms in place that are capable to touch the files on your backup. At least in dropbox' case these mechanisms are not completely separated from the initial backup mechanism of each version. I've had two almost catastrophical data losses with Dropbox until I was able to make that distinction. I've come to the conclusion that IT professionals should never ever treat a sync system as a backup system, and if you still think so please don't spread that advice to others.
Most of us here are developers. How many of our products are perfect? Always assume something will fail in a new and interesting way in the future. I'm not letting Dropbox off the hook - their product shouldn't do this - but you'll be happier if you treat backups with the same level of redundancy and planning as the rest of your infrastructure.
So I think the underlaying problem here is that any backup/syncing system might have a bug (like this one) or there might be operator or user error (deleting your revision history is just a couple clicks away). Recovery oriented computing website has a lot good papers on this topic [1].
This is very similar to problems with outages on Amazon EC2 - yes Amazon cloud is great but in order to make your service highly available you do need to have standby system on some other cloud (for example, we run on Rackspace but our standbys are on Amazon).
One approach to protect yourself against problems like this is to replicate/sync all your files from one cloud storage (your primary one) to some other cloud service (GDrive, SugarSync, Box, etc.). So should Dropbox have a bug, then you still have everything in other cloud service: including all revisions.
Services like cloudHQ [2] (that is my baby) can replicate and sync all your files from Dropbox to, for example, GDrive. And of course cloudHQ has options like "two-way" sync, "don't replicate deletion", "backup" (weekly incremental are in folders - so your will be fine even if "revisions" feature fails), etc.
For services that purge old versions and deleted files at 30 days, you lose if you don't notice a problem promptly. You can't be expected to be watchful over gigabytes of data; that's the whole point of a backup service.
"I'm Alan Fairless, a co-founder at SpiderOak" [1]
[1] http://news.ycombinator.org/user?id=rarrrrrrAs an aside, it seems Dropbox is biding on the [SpiderOak] keyword on Adwords, you should probably at least outbid them on your brand terms to reduce confusion/ misdirection for potential customers.
So they are. Is it legally acceptable to bid on competitor trademarks? I thought that was regarded as being over the line these days - anyone know for sure?
If you look at the adwords link, it's clear that Dropbox is bidding on "competitor keywords" as a class.
In terms of a trademark in your keyword, this is only unacceptable in: Australia, Brazil, China, Hong Kong, Macau, New Zealand, North Korea, South Korea, or Taiwan, and only after the trademark holder files a complaint.
[1] http://support.google.com/adwordspolicy/bin/answer.py?hl=en&...
Thanks for the link.
Bugs happen. If a bug happens on the sync-ing service and it trashes your 'back-up'/history then syncs and trashes your primary copy, your toast.
Sync'ing != Backup. They are for different problems and have different restrictions/pitfalls.
I do agree with you but, I can tell you that selling backup service is harder than you think. Also as pointed by paper [2], the human error accounts for ~50% of all system failures. And the worst thing is that majority of users who accidentally delete data, don't even notice data loss until lost data is needed and they don't recollect doing something wrong.
What I found out interesting that people (i.e., small business owners) will are scared of losing a credit card (even though you can call the bank and cancel your lost credit card and get a new one - inconvenience but not a big deal), but they will not backup critical company documents and data (even if they lose them the company will be pretty much closed - there is no "bank" to go to and get data back).
[1] http://blog.cloudhq.net/post/33844549768/the-difference-in-d... [2] http://roc.cs.berkeley.edu/talks/pdf/HP.pdf
I have a friend that strongly recommends CrashPlan, but I haven't tried it out yet on my Mac. I'm curious to though.
The setting is controlled thru the CrashPlanService.ini file.
Even if I get hit by lightning tomorrow, the service runs perfectly fine on its own for months at a time, so you'd have plenty of time to get your data back.
p.s. Please avoid golf courses this weekend!
>> is there a dead man switch or notification procedure in place?
The absence of weekly HN posts.There are people who should send out that notification if needed, yes.
p.s. Please avoid golf courses this weekend!
Don't worry, I don't play golf. :-)
I very commonly hear why people are using Tarsnap, and from time to time I hear why people are no longer using Tarsnap, but I very rarely hear why people never started using Tarsnap, so I really appreciate you taking the time to comment.
My earlier point was that I think data is stored on Colin Percival's S3 account (he is the creator of Tarsnap) and therefore you might lose access to the data (if he couldn't pay the bills or got hit by a bus) even though S3 itself is fine.
Maybe I could plug my new app here as well, tidy.io[1] lets you archive or backup your files directly to and from your Dropbox. Feedback is always appreciated!
[1]: https://www.tidy.io/
It has a lot of options for retaining old versions, too: http://support.crashplan.com/doku.php/reference/version_rete...
My current scheme is to make rolling snapshots of my Dropbox folder backed up to a local RAID array which backs up to a separate RAID array nightly. More info: http://aaronparecki.com/2010/190/article/1/how-to-back-up-dr...
Hackernews link for that post if you're in to that sort of thing: http://news.ycombinator.com/item?id=4704667
You could also add a bit to the script explicitly looking for zero byte files and alert yourself to their formation.
Sync doesn't necessarily have the capability to destroy files, rsync has a switch to delete files that are locally deleted. However Dropbox is supposed to be rsync + rcs so this kind of problem is supposedly easy to fix by simply reverting to a good previous version.
That's simply not true (I've built sync systems that are incapable of destroying files).
IMO, syncing designs that do have unrevokable overwrites are inherently brittle. (I don't know if Dropbox is built that way, but AFIAK, iCloud (and MobileMe before it) is -- and it sucks.)
It's also important to remember the meaning of the word "permanently", while this is a UX disaster, the title is misleading.
I just wanted to let you all know that we take any claims like this really seriously. There aren't any known bugs on the Dropbox side that would cause this, and unfortunately there are potential causes such as hardware errors, filesystem corruption, and other OS issues (including those like http://www.phoronix.com/scan.php?page=news_item&px=MTIxN... which another poster pointed out) that can corrupt data or create zero byte files.
Nevertheless we will continue to look into this just to be sure, and we also work hard to find ways for Dropbox to shield users even when the OS, disk or other components fail (our undelete/file revisions and Packrat are among these).
Thanks for posting here. The reason I'm confident that this bug is not due to filesystem errors on my machine's part is because Dropbox's version history for these files has become corrupted. If you read over the details and correspondence, you'll see that Dropbox engineers were able to recover two of my files that Dropbox's version history had reported as only having a 0-byte version. These files were edited recently, within the 30-day window that all Dropbox users, Packrat or not, have version control for.
The files had clearly been whole when first synced to Dropbox, but that version was not listed in Dropbox's history, and so I had no power to restore them. Even if my own disk had spontaneously 0-byte'd those files, this should not have caused Dropbox to lose the ability to restore it to its original version.
More circumstantially, others are reporting similar issues:
https://twitter.com/dangillmor/status/261921738441515009 - https://twitter.com/pc1oad1etter/status/261957001234505728 - https://twitter.com/frr149/status/261957708746469378 - http://news.ycombinator.com/item?id=4704236 - http://news.ycombinator.com/item?id=4704178 - http://news.ycombinator.com/item?id=4704063 - http://news.ycombinator.com/item?id=4704485
It's tough to tell from these small updates whether these users had their files 0-byte'd only locally, possibly by the ext4 bug, or whether they've also verified that it's not recoverable from Dropbox.
In addressing this bug report, please specifically address Dropbox's loss of version history. The two files that Dropbox engineers recovered had their version history wiped within the 30 day window.
I ran the find command suggested by the OP,
and it came up with a long list of files -- all
names that I had intentionally and manually deleted.
It seems when one of my "other" machines booted up,
it put 0 byte files back in their place. A review of
the file's history, by clicking Dropbox -> Browse on
Dropbox Website shows the original file, the day I deleted
it, then a few minutes later, a 0 byte file added back.
Edit: posted because it seemed like the sort of information I'd want if I worked at Dropbox and were looking for clues to the nature of the bug. If this isn't kosher let me know (rather than simply downvoting) and I'll delete the reposted comment, if possible.I use Dropbox primarily as a tool to synch content between my desktop/laptop/phone, but any significant change I make to those files gets saved locally 100% of the time.
I am not a very trusting man.
I've confirmed files that should contain something (nav images for a website) and that do contain something in the original source folders stored elsewhere are empty in Dropbox. Interestingly, Ubuntu is one of the clients syncing to my Dropbox folders.
I'm turning Dropbox off now and disabling it on startup.
I'm thankful for this post, now to figure out what's going on.
That said, I'm definitely writing a script to do nightly backups of the contents of my dropbox folder going forward.
Crashplan for backing up.
That combo hasn't failed me yet.
http://jeffreydonenfeld.com/blog/2011/12/crashplan-online-ba...
Dropbox is only the former.
As jpadvo said, I might look into making some one way backups to S3 or something.
I backup to an external HDD and to the cloud and still have the originals (as well as having extra copies again of my music and photos synced across my computers) - the more redundancy you have the better.
It sucks that so many people need bad stuff to happen to them to do something about it - I'm so thankful that storage became so cheap before anything really catastrophic happened to me. I've lost data in the past but it was back when so many things were offline, nowadays it's CRITICAL to have a good backup plan.
To be fair I never would have considered such a bug when using dropbox - I would probably have considered it safe considering you have a local copy of your data, especially since as other pointed out they present themselves as a backup service.
I almost got hit by this locally actually, I synchronize folders against multiple PCs, and the exact same thing happened, and a number of files had their bytes zeroed out and then this was propagated through the network. Thankfully I spotted it before all copies were overwritten and fixed it, but that's where you also want something like Time Machine.
Ugh, so many ways to lose data, even when you're doing the 'right' thing!
[1] http://www.phoronix.com/scan.php?page=news_item&px=MTIxN...
I do Dropbox manually, I rsync a list of folders to a linode instance running mercurial. It's simpler, scriptable, more flexible and as fail safe as Dropbox. If something does get corrupted its really simple to go back in the version history. I can't remember the last time I had a file system corruption with ext3, I suppose they still happen, but not to me in solid use for years. Obviously my mercurial repository is also backed up regularly.
I don't agree that a backup system has to be restore tested periodically, instead I believe that the restore process has to be an integral part of the workflow. In my scenario, I rsync to hg, commit and push, then from other machines in my workflow (or more likely vm instances) I pull+rsync back. This way the backup and restore cycles are just part if the workflow and everything is version controlled at the same time.
gci $env:USERPROFILE\Dropbox -r | where { $_.Length -eq 0 }15 Files affected here. I haven't checked to see if any of them are unrecoverable (none of them are vital), but this does seem like a very bad bug.
https://support.google.com/drive/bin/answer.py?hl=en&ans...
For RAID like protection with Dropbox and other providers, you can roll your own BRIC (Redundant Bunch of Independent Clouds). I did this using Tahoe-LAFS to stripe data across storage providers. Requires a bit of set up, has some caveats, but does work. If you use with duplicity you have versioning on top of a distributed, encrypted, redundant store.
To that extent I keep my important files backed up not just in Dropbox, but also to Crashplan, and to a spinning-rust hard drive especially kept for backups that I protect in my home. That's three points of failure I can recover from if something goes wrong; and if all three fail at once, then I probably have worse things to worry about, like the zombie apocalypse.
I am a bad sysadmin, bad at back-ups, bad at security, bad at redundancy. And I would guess that describes 99.99% of people who care about family photos.
I don't know, doing basic backups aren't super-hard for someone who's reading HN. My setup is really just a regular Linux box with a 1TB HDD sitting in a closet of mine, with dynamic dns pointing to my own domain (not even a strict necessity), and I rsync it whenever I have new data like photos. That box in turn syncs to Crashplan and Dropbox, which is automatic.
Yeah rsync and the concept of a backup PC is beyond mere mortals, but for those of us here--if Dropbox (or any service) is your only backup, I think you can only blame yourself.
Also, I had several occurrences throughout 2011 but so far hadn't lost anything during 2012, so either something has been fixed or I've just been lucky. :)
So, if you've noticed zero length files lately, can you check the timestamp on the last update? Is it recent or over a year ago? You can also go to that date in your event log in the Dropbox web UI and see what was happening around that time.
1) Process which syncs files to the server fails to get access to a local fils, sees it as length 0
2) Process proceeds to tell the server file is length 0 and the server updates the file to be 0 length.
3) Access to the local file is restored and client notices that the server file is 'newer' than the local client version and it was length zero so it truncates the client copy to length 0.
The thing is, I can imagine a number of scenarios where the local file might seem to be zero length (oplocks on NTFS volumes being one)
The trust my mother had in dropbox is now gone, and probably will remain so for the next couple years.
The command
find /home/keith/Dropbox -size 0
shows only cache files for deleted files, and some backup files that I saved while empty (I know those should be zero bytes).A personal work around is a simple bash script to copy Dropbox directory to another with the date as directory name, I'm running this once a day or so. Then my normal old-school backup onto an external drive will catch each day's dropbox.
Surprising how convenient I found automatic file sync, and how quickly I came to trust the dropbox daemon running in the background on 3 computers!
If you'd like a batch file to check this, that you can run periodically (perhaps programmatically), try this:
@echo off
for /r C:\Users\username\Dropbox %%F in (*.*) do (if %%~zF LSS 1 echo %%F)
pause
If you'd like to echo that list to a file, change the middle line to: (for /r C:\Users\username\Dropbox %%F in (*.*) do (if %%~zF LSS 1 echo %%F)) > C:\Temp\emptyfiles.txtThis isn't a case of dropbox losing your files, it's a case of you not understanding how the tool was designed and how it works.
There is really no way for it to not sync file deletions without making a massive clutter.
# Look for zero byte files in Dropbox
MAILTO=my.email.address@example.com
@reboot /usr/bin/find ~/Dropbox/ -type f -size 0
I also have full system backups going back three months, with two hourly incremental backups for the last two. So if I get an email about any zero byte files, I should be fine restoring them from my own backups.If you use plain git, you'll probably need twice the space. Files stored in git are compressed, but music is already compressed, so there will be little savings. You'll also need a lot of memory if you want to copy files VIA git.
If git doesn't keep its own copy of the data, a Dropbox "failure" would still be unrecoverable, right?
Plain regular git does keep a copy of it's data in your .git folder, and every checkout/clone/copy of that repository stores the data, and it stores all old copies of all files. That's how git works. It also makes it a bit unweildy for large files like that.
What's cool (about git) is that the hash revisions (i.e. what git uses instead of version numbers) is basically a checksum of every file and every old version of every file. So if an old version of a file changed, the checksum would change and you'd be on a different branch!
It's awful that it had to come to that, but it's reassuring that they will be willing to work with you on that level.
I read "hopefully the zero-size test is reliable" as "is a reliable way to detect if there is a problem with my files" and that is what I was trying to comment on. Apparently I misunderstood.