Flickr Accidentally Deletes a User's 4,000 Photos and Can't Get Them Back
observer.com
observer.com
What it implies is that the internal tool(s) that the support staff uses was designed to do deletion instead of status changes. I imagine they have a web interface that fires off some backend script which a) runs sql deletes on an account ID and all tables and b) fires off scripts that physically deletes all photos across all servers.
Taking those implications in hand, it points to an incredibly poorly designed system for account management. Some years ago, there was an internal planning meeting where a decision was made to do this. Some people objected and were overruled. The ramifications of that project manager are still sticking it to people to this day.
I have a hard time believing that Flickr would have purposely engineered their systems in the manner I've described. Not impossible, but unlikely.
So I think its more likely that the data is still there, but that customer service has no method to resurrect it. Its probably a mix of internal policies and the lack of devoting a couple of engineers to fix the problem that is causing all of this.
Neither should happen. You never actually delete rows from your tables. You mark them as deleted, sure, but the data needs to be able to come back. Space is cheap, there are simply very few cases where outright deletion from a DB table is warranted.
I have a hard time believing Flickr would make a mistake like this. I agree with you - odds are if they actually let an engineer loose on this problem it'll be licked with a DB backfill in 20 minutes.
Consider how deep this rabbit hole goes. If expunging any copies of copyrighted data were strictly necessary, then:
What of data existing on backup media? Should those be re-hydrated, the target data expunged, and then recreated in abridged form?
What of data existing in various content caches (such as memcached)? Should all of those systems be flushed of any possible contaminant?
What of the actual data on physical media? Deleting a file on any modern form of media does not expunge the data, it merely unlinks the location of the data on disk from the file system directory. It might be necessary, depending on filesystem and drive type, to scan all of the unused sectors on an entire disk to find out if the target data or any part of it existed.
As you see, quite quickly you get into absurdities. Nobody goes to that much trouble to delete merely copyright infringing data.
I should note that this not true in the case of child porn. Once the feds give them the signal that is completely purged.
This comes with a cost, and it's not a small one. Every single query that ever touches this table, in every single piece of code owned by the company, written or maintained by every single developer (consultant or transient or offshore, none of which ever got training sessions on this requirement) needs to be enhanced to respect "where isDeleted <> true".
It's well understandable if the expected cost to recover after the fact from an occasional issue is smaller than the cost of implementing and maintaining that in the first place.
A sane approach would be to do all your table access via views, and the view definitions have the 'where is_deleted = 0' or whatever. Then your database handles it.
Updates to the view work as well, in pretty much every database system, including MySQL even.
Using views adds a layer of abstraction. That's often a good idea and a step in the right direction, but does come at the cost of increasing the overall complexity of the system. It's an incremental step to the positive, but no revolutionary solution.
The design for a system where you had multiple redundant copies of all the data and could roll back any changes to the data set is different - and would cost a lot more.
EDIT: I don't believe that Yahoo! Flickr does not have a way to recover these things. Or I'm completely wrong.
For a single account it's more of a pita.
Retrieving all the datas associated with one account, put it in a format readable from the production environment and recommit all the changes while dealing with the potential errors and inconstencies seems quite a huge deal.
Messing unrelated data in the process would also be a nightmare.
Of course, not deleting anything from the start wouod be the way.
The lack of a retrieval system could be from a lack of Yahoo wishing to invest in upgrades to the platform in general. The last few years have seen few changes for users beyond geotagging and tagging persons, and competitors (free and otherwise) have stepped far ahead in terms of personalisation and galleries.
I'm very curious to see if Flickr turns into another 'Delicious' situation in a couple years.
http://www.observer.com/2011/tech/flickr-accidentally-delete...
Having seen this kind of effect from a brain drain in the past, I think you're exactly right. Software needs love to survive and this is what it looks like when software and a company becomes unloved.
I can see why Flickr might engineer things this way. Most of their users don't pay (granted, this guy was a pro user). Deleting images when an account is deleted saves money.
As we start to rely on cloud providers to look after our data, we need to either become more educated (and proficient) by creating regular backups ourselves (which kind of makes the idea of managed cloud services defunct imo), or be in a position where we can individually sue for damages incurred by negligence on the part of the service company in question.
So many of these companies have liability clauses which negate all responsibility in situations where negligence leads to loss.
I think the current situation is crazy, considering the amount of responsibility and trust that's involved in using web-services which store and manage irreplaceable data.
Indemnification is likely going to push prices up further, making cloud non-competitive compared to traditional storage.
Still, I feel a lot of old technology is merely being relabeled. Today everything that stores data remotely is called a 'cloud'. We may yet see small NAS units relabeled as 'private micro-clouds".
However, I'd argue the ability to take legal action, would actually encourage service providers to diversify their own backup strategies and procedures. If a mistake becomes too costly to consider - I think less mistakes would be likely to occur.
This conclusion assumes that the people running the business are capable of resolving such issues. If they are not, then all it does is put people out of business. Not everyone grows and adapts when faced with such pressures. Some of them just go extinct.
If a company can't be relied upon to look after my data - I want them to go out of business.
And I have a feeling that this money does not necessarily go into making the product more reliable, but that it is used to buy a better insurance.
My camera holds about 2,000 pictures on the highest compression setting on an 8GB card. It's not like that amount of data is precious. I have hundreds of movies and thousands of TV episodes stored, why is it that this person couldn't keep their photos?
I hear a lot of stories like this, and being a writer I can only think "are they stupid?" Until recently everything I'd ever written, if stored in RTF and rar'd would still fit on a floppy disk. I think I've expanded to two.
I now have a flash drive for storage. However my basic method is to simply download and rar all my files from Docs, email it to myself through gmail and load it onto the USB drive. It leaves me backed up in 3 places and the original copy.
Beyond that, I have well over 2000 photos in my photo library, and I'm by no means a serious photographer. A professional might take several hundred photos in a single day. (Not all will be useable or worth publishing, but the idea that you might just never delete any photos off your memory card just doesn't work for a lot of people.)
The issue though is the links, comments, viewer history and comments that you cannot backup or restore - but which have real value.
There is a joke in there about putting a cloud in your cloud, but I am not sure how that meme goes.
What I've done in the short term is to use FlickrTouchr (https://github.com/tominsam/flickrtouchr) to retrieve my ~7 GB of photos from Flickr and upload them to Dropbox. The Dropbox gallery feature works quite well and has the advantage that you can easily backup/modify the underlying files.
In the long term I'd like to move the pictures to my own server (allowing me to use my own domain, etc.) and hack together a simple frontend similar to Dropbox' gallery.
If you decide to move away from a Flickr pro account (and if you also have non-public photos on your account) don't forget that after your account is switched back to a normal account you can only access the 200 most recent photos. So if you want to delete your photos from Flickr it's best to do so before the pro account expires.
I think Opera's personal browser based server ("Unite") is going to have it's day (or something similar will) - http://www.opera.com/press/releases/2009/06/16/; has a "photo server" built in.
That's the biggest point I took from this. Removing the wrong account can always happen, but actually wiping data instantly without an invisible "grace period", or retrieval from backup, or anything else to get it back, that's just poor. Especially if it's a pro account.
It sounds to me like they didn't invest much time in the internal admin section. They probably delete a lot of accounts every day, and they've never had an issue with the wrong account being deleted? Or an after-the-fact clearing up of a misunderstanding? They don't start a dialog, at the very least only with pro-users, to get the other side of the story?
I hope something else went wrong here as well, and it's not the way they work normally.
A few days ago we had a big thread on Undo in software development, and a lot of people were suggesting Undo is very very difficult to implement with little benefit to most users.
Then you have articles like this where people are shocked that undos weren't in place.
If I were Yahoo, I dont care how much it would cost, I would have someone, or someones, recreate his account as it was, even if that meant calling in a data forensics team.
Doing a quick google search turns up this gem:
Staff - heather says: "When an account is deleted, photos and any metadata associated with account is queued and deleted. If not instantaneous, it's removed within minutes."
RAID is there for system interruption protection. RAID is there to keep flickr cruising along without interruption.
If they delete content from a volume covered by RAID, then all the mirrored copies go poof simultaneously.
What you are suggesting is a higher level content management system. Apparently (from the drift of the comments) that is something flickr doesn't want to get involved with.
They need to learn the mantra: "Replication is not backup, DROP DATABASE replicates pretty well.".
Of course, deleting a complete account without any means of rectifying it sounds like a very bad idea and shouldn't have been possible. At the very least there should have been red warning signs flashing when the support person deleted a pro account with 4K+ pictures in it.
And if you delete the data, then $publication will write articles about you being unable to restore it.
Take the Facebook "account disable" feature as an example. In the context of this article here, it was sold as a "good" feature (so as something that Flickr should have too).
But this exact same feature was all over the press a year ago as something really bad and just showing how Facebook just doesn't want to let you go.
So. What is it now?
Personally, just flagging for deleted and living with the privacy concerns might be the better option because there are potentially more people deleting data by accident (or have it accidentally deleted) than there are people too concerned about privacy to use the service.
Also, users concerned about privacy just won't be users whereas users who just lost their data were users and now aren't, so that's a net loss compared to the users not joining.
Maybe they'll do the latter after some rethinking!
Most other services like Facebook, GMail are explicitly designed to keep your data forever, and if you don't want that, well, you should have thought of that before using the service. I would actually be happier and feel safer if we heard more complaints from Facebook users that accidentally deleted their account and couldn't get it back. It would be a clear sign that their system was designed with privacy and "data hygiene" in mind.
There's been quite a few well constructed and thoughtful apologies circulating HN lately, one of which was Andrew Mason's apology for Groupon's Japan blunder. Hopefully, Flickr will pick up the torch, and issue a more official, and hopefully more well thought out, apology. After all, it's been less than a day since the incident.
Of course, the deletion in itself is inexcusable. I just hope Flickr steps up.
(Not to mention the Flickr API has a KNOWN ISSUE with the search that has a fix ETA of several WEEKS! http://www.flickr.com/help/forum/en-us/72157625560721827/)
what else could flickr have offered to compensate?
It's obviously expensive for Flickr to offer something like this but I mean they accidentally deleted 5 years of a paying customers work. I think a refund is the bare minimum.
Now they've made a mistake, and have refunded this years payment and added future access for free. They have done more than 'the bare minimum'. It might be good PR to refund previous years' fees, but they were at the time provided in good faith and competant manner; there is no 'fair exchange' reason to refund them.
i've seen some flickr-specific backup scripts that just dump all of the images to a directory, but they don't preserve titles, dates, sets, collections, tags, etc.
Why?
The difference between backing up your Flickr data to an encrypted drive at, say, Amazon and backing it up to an encrypted drive under your bed is: Amazon has better physical security, and far better redundancy, because (assuming you use S3) the "drive" you save it on there is actually a replicated data store that spans multiple regions and datacenters.
The point of backups is to provide redundant storage that is relatively uncorrelated to the original. Once you move the data to different disks at a different company in (if you like) a different country, you've done a lot to solve that problem.
also, many of the reasons backupify.com cites for needing backups would also apply to their own s3-backed service. "hackers", storage failure, human error, TOS violations, accidental account lapse/termination, etc.
Anyway, the artist has originals backed up, but the article points out that all of the links to those images, and embeds, are now broken.
If URL is constructed from the data somehow (auto-slug from the title, hash or something), such restoration becomes way simpler
I am forced to conclude that few read the article, since most of the comments have focused on the - erroneous - idea that he depended on Flickr to store his data and had no backups.
It's a paid service[1]. So yes.
[1] he had a pro account.
(Point being, nothing site-wide in my quick glance seems to indicate a promise of the integrity and forever-ness of your data.)
That said, I can't seem to find it now. Weird.
1. people who've been in the situation of lost data when somebody else responsible for its backup hadn't backed the data up.
2. people who don't belong to the first group.
After graduating from the second group into the first, i've always been classifying my data as either "can loose, fnack it" or "i'm backing it up myself, right now"
What isn't backupable, and thus is lost forever is all of the comments he had on the various pictures, all of the groups his pictures belonged to, all of the embed's of pictures on various sites, all of the social aspect of Flickr (such as friends, family, all of those things).
I use Flickr to share my photos with the world, but I keep a local copy of all of my pictures on my local file server and on my laptop and all of the backups I have of that laptop...