Instagram kept deleted photos and messages on its servers for more than a year
theverge.com
theverge.com
These people imagine you just rm -rf a file and run a few SQL queries, you literally can't just delete data. I've worked at the scale of Instagram before. Literally nothing works like that at that scale, the person who wrote that article should have called up someone from Instagram and had them explain at a really high level how complicated deleting data is at large scale.
Also, you're saying they intentionally built a software stack without even thinking about how to obey the law (the 30-day deletion requirement existed even before the GDPR, since 1996 even, just with lower fines)
I hope in a few years this case will be taught in school just like the Therac case is being taught right now.
So you don't know anything about Instagram but at the same time know for certain that they have legacy infrastructure that keeps them from meeting legal requirements? How about we don't give them a pass just because it requires a bit of work to comply with laws and regulations.
Laws and rules need to be followed no matter what.
It's the same with GDPR.
And the fines for not doing that need to be big enough to motivate them to do that.
Regardless, in practice, billion dollar companies have the resources to ensure 'delete' isn't actually "partially hide for years/forever". With proper surrogate IDs they could even preserve a minimal amount of meta such as tombstone IDs to ensure all first and third party systems are in compliance.
Also, if it not works at scale, don't tell your users that it works and pretend it works as you communicated.
I find really bothering, when engineers trying to hide behind the "it's a very complicated process in the background, so we just say we have done something, when the truth is, it is in progress. It's not a user error, when you see the word "deleted" you assume, your picture is gone, it's your error, if you state it's gone, while it is not.
You literally can.
You can't remove data that you don't control though e.g. Removing embarrassing pictures from the internet. However if you control the server you control the data. I can appreciate that it might not be as simple as a unix command but it is a single letter in CRUD. It's fundamental.
'Your delete request is being processed. No one will be able to see this item though it may take up to three days for the data to be completely removed from our servers'
This is not a technical challenge, it's an issue of priorities and Facebook doesn't prioritize privacy. They were able to release a TikTok clone in a few months, but can't solve deleting data?
Imagine some engineer writing some feature you've never heard of and they decide they need data in a different format. They write a job to do this.
Now imagine that job somehow has a bug which prevents the deletion system from understanding that copy exists. That bug could be non-obvious to find.
It does if you have cold backups that take a year to cycle out. They're often offsite, compressed, and incrementally hashed so finding individual items and removing them is really hard -- you're much better off just waiting for them to expire, and a year isn't an unreasonable amount of time.
Backups are covered under GDPR, although when a user requests erasure you can say "your data will be rotated out in X months/years". Not sure how this applies to access, but I assume it's similar: https://ico.org.uk/for-organisations/guide-to-data-protectio...
Before MVPs were championed, companies would create a full-fledged project management suite with 50 database models and release it to crickets.
An MBA and a junior engineer define an MVP as “You can log in, and add todo lists and check them off.”
A senior engineer defines that same MVP as “You can log in, add todo lists, check them off, delete your account, and if the server crashes or we’re hacked, we have backups for 30 days.”
Or in other words with an MVP you still need to complete the 80% of the iceberg the MBA and junior engineer can’t see, but you don’t build the whole iceberg at once.
Sure, eventual consistency and replication and distributed storage is hard.
Alternative approach: encrypt at rest, delete encryption key. It would require extra resources for sure, but if privacy was a concern... ;-) (and then have a reasonable expiry on your CDN if you use one).
(of course then we can argue that deleting the encryption key may be difficult too, for the same reasons)
People that suggest using encryption and throwing away the key have not thought through the implications of that approach. It scales very poorly and therefore is not suitable for most practical systems. There are good reasons “obvious” solutions like this are not used.
If you change the definition of “delete” to mean unrecoverable physical deletion, then sure, it scales poorly. But it is a bit like redefining “fast car” to “can travel faster than Mach 5” — technically a valid as a redefinition while completely ignoring the engineering realities of what a car can do.
Actual deletion is what Facebook told us they did; that's what everyone assumed when they said they were complying with GDPR. This is like advertising a car as 'faster than Mach 5', and then when people call you out for lying, saying, "well, 'faster' is a relative term."
This kind of crap is exactly why people don't trust Facebook. It's not because people are paranoid, it's because Facebook systematically creates expectations in their advertising and public releases that they're acting responsibly, and then acts like they're the victim of unfortunate circumstances and misunderstanding whenever they get called out.
If three weeks ago a tech journalist had written an article saying, "Facebook will fully delete your data when you ask", no one at Facebook would have been reaching out to that journalist saying, "Oh, in the interest of preventing misunderstanding, we don't actually delete the data, we just mark it to be ignored." But now that they've been caught, now it's just a big misunderstanding by people who don't understand database architecture.
When companies tell the public that they're doing something, it is reasonable for the general public to assume that they're referring to the commonly understood definitions of the words they use.
Deletion is pretty much the hardest thing to do in data management, way harder than inserting or retrieving data. I know GDPR says you need to be able to delete a user's data, but how reasonable is it to trawl through your offsite backups to find their items and remove them?
No one expects you to just rm -rf. They just expect you to expect the same effort to remove the file that they put in to create the file. It's not like the file got propogated to 10 different data centres by accident. There was a decision that an uploda needed to be everywhere in 5 seocnds and a deletion was an advisory hint that could be ignored for a year.
If you can't reliably get rid of the data you administer, don't collect it in the first place.
And if they don't, they're mismanaged and need to allocate more resources to it, because it's illegal not to, for good reason.
This makes it sound like they consider "you could see it" the issue, not "we were still keeping it". In other words, the fix was to hide it, not to delete it.
If I were the Irish DPA (and actually wanted to do my job and had the resources to, instead of being intentionally lazy/crippled to attract tech firm headquarters), I'd definitely be asking for retention plans and evidence that the data is now being removed in a timely manner, and start issuing fines (small ones for past transgressions, big ones if they keep doing it or don't have a decent plan how to make sure to get rid of data they shouldn't be having).
For comparison: Deutsche Wohnen (large real estate) got slapped [1] with a 14.5 million EUR fine for over-retaining sensitive tenant data and not having an automated system to delete it.
[1] https://www.dataprotectionreport.com/2019/11/first-multi-mil...
From the moment the problem is brought to their attention, then it becomes intentional and willful, so the fines should be much bigger.
It's simple: change your ways and use the initial fines as a wake up call. If you then do not wake up and persist the fines will get heavier and heavier until you will pay attention.
A Dutch hospital managed to get to the third round of fines and they weren't all that happy afterwards. 460K Euro fine for a single instance of ignoring the regulators on a single individual.
Believe me when I tell you they have understood now.
The initial fine was zero, just a warning to improve.
The case revolved around a very minor dutch celebrity whose data was reviewed by hospital employees that should not have had that access.
There's middle ground between "we take 100% of your revenue" and "we take 0.001% of your revenue". Given that we're not this lenient with private citizens and small companies, why should we be with international corporations?
My second argument is the efficiency of using capitalism to fight capitalism. Take money from companies who make mistakes with personal information. Be that out of ignorance, malice or bad luck. Why does it need to be fair and just - make money from it. Make it a risk to capture personal information in the first place. If those companies go out of business then so be it. Others will take their place. It’s not impossible to do business without storing personal information, there is just not enough incentive to bother.
(2) small fines do not play down the seriousness of the transgression (which is the word I think you should be using for instances like this). They merely indicate that you should clean up your act assuming no real harm has been done. In some cases the regulators have immediately resorted to fines, and quite large ones as well if they felt that the case warranted it. They do have that option.
But putting companies out of business was never the goal, contrary to what a lot of alarmist people were screaming when the law went into force. Also, over time as more and more companies have been fined I would expect that the initial fines will go up because claiming ignorance really isn't an option any more. Some comments in this thread are particularly worrisome in that light, it appears that some people still don't get it and they are in positions where they really should know better.
Turning data into a liability rather than an asset is the long term outcome. This will take time, and when it happens I'll be that much happier. Every company will have to seriously weigh the price of holding on to some datum vs the price of losing it.
We do live in that society. For example, the first few time you get caught speeding, you get a warning or a ticket with a fine. Keep getting caught speeding, and you lose your license and/or go to jail.
These companies profit from personal info whether one is a customer or not. I would imagine everything exists with a "user de-activated since $DATE" entry or similar in their database.
Perhaps I'm too cynical?
If they can't show the pictures that they have claimed to have deleted then they can't use it to bring in eyes for advertisers, which means that the data is a net loss for them, causing them to have to buy more storage sooner.
They should want to delete data as soon as possible
You see a similar pattern with respect to signups vs account cancellation.
The weird thing to me is that it is usually the marketing department and not the legal or the compliance department that has the upper hand in these data retention discussions. Fortunately thanks to the GDPR this is now changing and slowly companies are coming around on this.
Data is funny that way. It is easy to acquire, easy to lose if you want to keep it and devilishly hard to get rid of for real if that is what you want to do.
One day, they will be burned when some investigation finds out that those rows are still sitting in uncompacted tables, or on now-unallocated disk blocks, or on now-remapped SSD sectors.
The only way to be sure the data is deleted is to copy all the data you want to keep to a new drive and burn the old one. Anything less and there is a reasonable chance an expert could recover at least some of the deleted info.
Secure erase is a contractual requirement for many relationships, it is interesting that none of the major db vendors as far as I know have a secure delete option.
There are lots of tricks and attempts to work around it but no official support for such functionality afaik. And that's before we get into VM snapshots, database snapshots, backup copies, copies that were made by developers or data scientists for test purposes (of course, nobody ever does that) and so on. Nasty little problem.
[+] or some employee shooting their mouth off in an online forum.
Key issue: you can't do any operations on encrypted data, essentially you're killing off your database.
Homomorphic encryption is academic research, not something that is widely available and supported in common open and closed source databases. The best you can get is a database that encrypts the on disk data (Oracle TDE), but that only protects against a server being stolen or hacked on the OS level.
Also, for purposes of operations on data it all depends on what the column holds. For instance, you don't actually need access to fields such as names and dates of birth if you have a client ID and that is yours, not the customers and so you could leave that field in plain text. Any operation would then need that client ID but that's workable.
You could even say that if you would need that decryption key for anything other than user or controller directed computation that you are probably doing something you shouldn't be doing. In all other cases the context is clear, consent has been obtained and the data can be decrypted if required.
I’ve thought a lot about what it would take to design a “delete optimized” database kernel every since GDPR became a thing. In principle it is possible, but I seriously doubt anyone would use a database that is literally orders of magnitude slower and less scalable for everything except delete operations. Would it be acceptable to increase the resource intensity of databases by 10x (and environmental footprint implied) to get “real” deletes? That is the tradeoff here.
This has precedent in SQL databases designed for high-assurance applications with ultra-fine access and visibility controls. When the average software engineer understands how they work, it sounds like a great idea for having more secure data and they wonder why it doesn’t seem to exist. The reality is that they do exist but they are so abysmally slow for even elementary things that no one would ever dream of using it unless there is a narrow government requirement.
But you're totally right that this is a tricky problem and hard to do properly. On the plus side, field level encryption is something that we've been doing for ages and that already works quite well. If you design carefully you can even get some processing done on those fields though it will take some major breakthroughs in DB engine design before you can have your cake and eat it too, in the sense that you can be both legally compliant and do all the kinds of processing you can do today. I even doubt if that is desirable, lots of those examples of processing should probably not be done in the first place.
First, encryption has a block size. Data field storage in databases is typically measured in bits, as closed to the information theoretic limit as practical. Storing an 11-bit datum for some column as a 256-bit AES block represents a 20x expansion in storage cost. We could go back to the old row storage model, which would allow the record to be encrypted as contiguous memory, but that would both bloat storage (for different reasons) and we get to relive the golden age of very poor query performance. Modern databases are built on succinct representations because throughput is memory-bandwidth bound. Even in conventional databases, ignoring this design detail will cost you 100x in throughput.
Second, keys and key schedules thoroughly thrash the CPU cache for scan operators. In a typical data model, a single row will be much smaller than the key schedule for decrypting that row. Every single row, several thousand per page, will require an unpredictable cache line fill to access the required key. Then you have two choices: compute a new key schedule for each row, which will be computationally expensive, or precompute the key schedule and take the even larger RAM hit. In large scale-out databases, many gigabytes of key infrastructure will need to be locally cached in RAM on each server -- you can't afford a network hop or page fault -- to decrypt each row in a page scan. Key management state consumes most of your runtime resources and crowds out the data model. I've worked on the design of such schemes in real systems, you end up devoting almost all of your cache/RAM to key state to the exclusion of the actual data model. Also burns up quite a bit of precious memory bandwidth without doing any real work.
Any database built on encryption for record-level physical deletion will be unusable for almost any modern application for well-studied technical reasons. It will work, in theory, if you can run your business on a database that performs and scales like it is 1995.
The best technical solution for physical deletion today is to rewrite cold storage, which is still extremely expensive and has extremely low delete operation throughput if you do it synchronously but at least it doesn't break database computer science. The only high-throughput and economical way to implement rewriting is asynchronously with very long deferrals e.g. over 30 days. Which is how databases have always worked, but with the deferral being indefinite.
Perfection, the way you describe it is for now unattainable. But the bigger problem as far as my practice tells me is that people simply don't care and set the 'deleted' bit and leave the plain text records + backups + log files all untouched.
The low hanging fruit is pretty much dragging the ground.
The “erase-in-place” operations you mention were common in databases a few decades ago and abandoned because the typical performance was terrible compared to the alternative. I am old enough I even implemented a few. It isn’t like these designs didn’t exist, they were deeply flawed for reasons that apply today.
This has nothing to do with perfection. Database engineers would love if deletes where inexpensive even in the absence of hard delete requirements, as that would make update operations, which people want, dramatically cheaper. But that isn’t the reality. You are essentially re-litigating settled database kernel engineering without understanding why they are designed the way they are.
If you thinking there is “low hanging fruit dragging on the ground” then I encourage you to prove it by designing a useful database kernel that can scale deletes while preserving insert and query performance, the two operations that drive all economics in databases. You’ll be instantly famous as a computer scientist because there are some nasty theoretical computer science problems in there.
You can't do effective per-user encryption on columns the database software needs to read (things you'll query or join on), but the database rarely needs message/post content and image content (often not even stored in the database). So encrypting those could be privacy helpful, if you can make the per-user encrypted store better at deletes than in general.
For person to person messages, if you have a separate record for the sender and the receiver, you can do some per-user transformation on the other correspondent, but that might be indexed, so data without keys would still show messaging patterns.
The second starts way further back: encrypt sensitive fields at rest, drop the keys when the user requests a deletion.
That way you never even have to touch the backups in order to have the right end result.
Do you back them up?
I'm wondering, more or less, whether the key management is another case of backups and deletion being hard.
Key management is indeed just another - hopefully simpler - version of the same problem. The reason why it simplifies things is because a single key can invalidate a lot of data stored in places that are out of reach such as cold backups.
But you still have the same essential problem.
I understand the user utility of a brief soft-delete period, but that hard-delete sweep should be performed on a fairly tight delay.
Or invoke COPPA and say your 11 year old cousin used your account, can’t remember what they did, but wants it all removed now.
Feature broke then companies replication.
After an extremely confusing conversation with the database team, i realized that they only replicated writes and updates.
I was the first in the history of the company to delete records.
They couldn’t imagine deleting anything. I couldn’t imagine keeping stale user data.
Ohh do I have some bad news for you.
Seemingly everyone in this industry can’t or won’t design databases and applications to actually allow for data to be deleted.
Arguably a well designed system should be built to handle having certain user data deleted without interrupting or breaking anything.
Can some of you weigh in (anonymously?) on this topic? Do you guys do hard deletes of user data instead of just soft deletes? If so, are logs or backups kept? For how long?
In other words: if I'm a user of $POPULAR_SERVICE and I delete my account at time t0, is there a t1 > t0 after which every trace of my data is gone from the platform?
My (cynical) guess is no, but I hope I'm wrong :)
The typical reasoning is that marketing wants to hold on to the data, they will never ever say 'ok, enough, you may delete it' because there is this infinitely small chance that they can re-activate an account, market to it for some other product (no matter that that is against the GDPR) or to sell the data to some third party if there ever is a cash crunch or panic. They see data as having positive value no matter what, whereas data that you shouldn't be holding on to is actually a liability.
I wonder when popular databases will have some first class support for soft deletion built in.
OP asked respondents to weigh in anonymously. That’s a call for anecdotes.
here is one not particularly recent discussion https://news.ycombinator.com/item?id=954393
Now that I think of it I would probably have done soft delete if I got around to automating deleting accounts at my old startup for lots of reasons.
on edit: another more recent article https://news.ycombinator.com/item?id=23005060
Imagine how hard that is when a datacenter is switched off for 14 days for maintenance, and then a fire breaks out and takes it offline for a further 20 days... When something is powered off, it's very hard to do those deletions... Yet misses of the deadline are exceedingly rare, even in cases like the above.
Sometimes disks are crushed in a crusher to meet the deadline if software approaches to deletion can't be done in time.
In Google, a user clicking delete is treated exactly the same as a written deletion request.
In fact, the law requires that they be the same - "Therefore, an individual can make a request for erasure verbally or in writing. It can also be made to any part of your organisation and does not have to be to a specific person or contact point.". (https://ico.org.uk/)
Only stating you want your data to be deleted leaves the counterpart wiggle room.
A person who may decide to bring suit however should always cite chapter and verse to lay down the line and to indicate that they are very serious about it. Just the fact that you would be citing that article will likely give you a better chance of seeing your request honored. But if your request is refused and you decide to tip off a regulator it won't make all that much of a difference, they will do their own investigation outside of the particular case and may broaden/narrow the scope of that investigation as they see fit.
This has already surprised more than one company by the way, they decided to play fast and loose with a single individual and as a result found their whole infra and processes under review with plenty of things found out of order. Fines were handed out that were higher than what it would have cost to arrange things properly in the first place.
My understanding is that deletion is a hard problem with entire teams working on it (imagine how many different random systems data flows to...) but that yes, the intended behavior is for deleted data to really be gone after 30 days. This is necessary to comply with GDPR and various other laws.
Of course, if two people each own a copy of a piece of data (for example, messages person A sent to person B), then person A deleting their copy won’t affect person B’s copy (just like how emails work).
Contrary to popular belief, Facebook doesn’t actually have anything to gain from nefariously storing data you delete. Ad targeting has plenty of non-deleted data to train on; Facebook has no incentive to break the law to keep tiny amounts of dubiously useful extra data on the margins. I’m almost certain the issue described here was genuinely a bug.
The rest of what you write is an open book to anybody in tech. And whether Facebook has anything to gain or not from nefariously storing data you delete is a much lighter shade of gray than building up shadow profiles, profiles on people without an account.
https://theconversation.com/shadow-profiles-facebook-knows-a...
However, there's usually some large slice of user data under legal hold, which legally can't be dropped as it's pertinent to some random long running court case, so not every trace of your data is gone.
You're probably starting to see the problem here. It's that the data actually exists in many systems, big and small, all of them with different processes and staffed by different teams. So deletion is really not a single operation but a coordination of many actions, relying heavily on a complex system of attribution and provenance to find all the places that each piece of data (among literally trillions) went.
All of this infrastructure is huge and it's active. I've been pinged many times while oncall to provide information or take actions in support of it. Every log stream, every database table, has to be carefully scrutinized to see if it could possibly contain user data, no matter how remote that possibility might be. It really is something we work hard at, and I know we're not perfect but anyone who says it's because we don't care is talking out of their ass. We're merely human.
That said, I'm hard pressed to explain the particular scenario in the OP. It seems to me that, no matter what other mechanisms are in place, there should be an egress filter to provide that One Last Check on data leaving our custody, and that should have kicked in here. But I have almost no interaction with Instagram from where I sit, so I can't speak for them any more than anyone else here can. Nor should I try. Probably said too much already.
Some of the engineers at these companies forgot their morals as soon as something got hard I guess.
If you ever downloaded your own account data, I think OP and other concerned posters would understand how much data companies retain and this wouldn't come off as a surprise.
Snapchat data for example, has chat logs, snap history, what accounts you've added/requested as a friend, and friends that have added you all retained. I bring this up as an example because the idea behind this application was to send a message that would disappear after a variable amount of time :)
Time to dig into it.
Shouldn't there be fines for this?
Took almost an year after being reported.
So removed from other person's chat. Still associated with your account
I have my account data from last year. Time to dig into it and sue instagram.
Oh wait. I'm not in the EU. I'm in India, where they can even sell my info. So they definitely didn't delete it
> > But clicking delete or unsend on a photo is not that.
> In Google, a user clicking delete is treated exactly the same as a written deletion request.
> In fact, the law requires that they be the same - "Therefore, an individual can make a request for erasure verbally or in writing. It can also be made to any part of your organisation and does not have to be to a specific person or contact point.". (https://ico.org.uk/)
Cars with defects don't have to be recalled.
Cars with defects that kill or maim people have to be recalled.
Software that does the same, also is recalled.
The bug was showing users photos that were internally marked as deleted. Not that the photos were not in fact removed from Instagrams servers.
No, there should be this expectation. If I have a photo of myself that I'm not comfortable being saved somewhere, there should be an expectation that when I delete this from a service, that service will actually delete it.
The expectation should be that the service will do what it said, not that it hid everything very well.
This is grounds for sending a strong message to Instagram for not doing what they are legally bound to doing.
Here's an answer I pulled from reddit:
If you are appropriately using TRIM the data will be obliterated when 'garbage collection' occurs -- forensics will not be able to recover deleted data. This is going to be largely dependent on your OS and Hardware, but if you confirm TRIM is working correctly in your setup, shortly after you permanently delete something (not recycle bin or trash) it will be gone for good.
The answer is that simple.
Supporting data: http://forensic.belkasoft.com/en/why-ssd-destroy-court-evide... http://digitalforensicsmagazine.com/blogs/?p=271 http://www.mcgoverngreene.com/advoedges/AE_pdfs/2011_pdfs/AE... http://www.forensicswiki.org/wiki/Solid_State_Drive_(SSD)_Fo...
There's also this as a less reliable added measure:
http://cmrr.ucsd.edu/people/Hughes/SecureErase.shtml
The main takeaway is that people vastly exaggerate how recoverable disks are, making any effort at all will stop 99.9% of hackers trying to recover data.
The GDPR does not allow for your data to selectively be forgotten either. You can request that they forget all of you, but not individual pieces of data.
https://en.wikipedia.org/wiki/Right_to_be_forgotten
vs
On the other hand, storage prices seem to be low enough for all these companies with bulk, long term contracts that developers wouldn’t bother doing real deletes of data.
I suspect that big companies in the market help justify manufacturer R&D into larger and faster and helps justify production capacity. I think we'd have smaller, slower, and slightly more expensive drives without their demand. But, speculation only.
A flag to mark data for deletion makes sense at scale.. given the number of other automated processes that run more often than a few times a year.. the user should be in control of their information and intent.
Deleted files hang out longer but are gone within a month.
Let me test this.
"Changing the rules" can reduce data availability, probably well enough for most purposes, and that's good. But it's simply strictly true that once you publish something, you cannot assure it's unpublished. And everyone should know this and act that way.
Opting out of the camera mesh being built in SF is impossible without moving. ISPs monetize our browsing habits, and a typical individual has no recourse -- common suggestions like a VPN just kick the can of trust down the road, and logless VPNs are usually proven to actually track you. Even if you didn't use any Google services, the adwords network is on the bulk of websites, and even with an ad blocker is a typical user expected to know that being logged in to Facebook means opting in to being similarly spied on nearly anywhere you visit on the web?
IMO it's reasonable to expect uploaded photos to live forever somewhere (even if GDPR obligates companies to do otherwise), but even learning how to opt out of the rest of what's collected where it's even possible would be a daunting task when undertaken from scratch that inevitably missed some key component and still resulted in being spied on and tracked. With that in mind, and given the scale of the problem, regulation preventing the most egregious of data-related offenses seems prudent.
Tomorrow at slate "How instagram permanent delete perpetuates white supremacy"
There is no policy that could satisfy all.