This is actually how the GDPR is designed: it's not about user data it's about personal data, that is data that can be associated or related back to an "identifiable person." The problem is not that you need to delete all the user's data across all your systems, the problem is that you need to break any associations that would allow you to identify a natural person who has asked to be "forgotten." The funny thing is that while some people make lots of noises about GDPR being some huge "burden," the reality is that any architect worth her salt should've been designing systems like this from the very start rather than letting personal data be replicated en mass from one system to another. It is a basic normalization of data. All the GDPR is doing, like most regulations, is requiring businesses follow best practices and not cut corners that might harm end users. The great thing about this in the long run is that this fixes a huge problem with the internet. Today users are reluctant to sign up to a service precisely because they don't want to surrender their data because once it's gone it's gone forever. I suspect users will be more open to trying new services if they can be assured that it's possible to "un-sign up", that is be "forgotten" by a service.
The problem is that as you collect more "impersonal" data, the probability that your collective data can still be used to identify someone approaches 100%.
"We don't know their name, but Deleted User 4510 was friends with every single member of the John Doe family except for John Doe himself..."
If all you do is rename someone to DeletedUser4510, you pretty obviously haven't deleted all the data you hold on them.
That is what this article is about, and methods to deal with that.
You erase the friend-linkages...
... Then you find that Jane Doe made a post in response to a blanked DeletedUser4510 post, and responses can only be made by friends, so therefore Jane was at some point friends with DeletedUser4510.
So you put in tombstones for all posts and all post-to-post causation links.
... Then you find that the entire Doe family is tagged in a photo by Jebediah Doe, except for one guy, and the comments are every family member and somebody named DeletedUser4510...
Anyway, my point is that it's really easy to get caught in such a fog of relational data that gaps merely change a certainty of identification down to an extremely high probability, and an event-sourced system -- by design -- makes it extra difficult to remove data or to even plan for its removal.
For instance, if someone posted on a message board, is it enough to rename their user to anonymous. Or do you have to go back and delete their user, leaving orphaned records? Or do you have to delete all of their postings, which could leave discussion history in disarray.
What about something like a phone service? Erasing a lines recent history is easy enough, but going back years to delete records from archival systems that weren't designed to handle it could be problematic. For instance, in very large data tables, deletes can be very slow. Call records are often stored in compressed flat files. Which would mean searching through tens of thousands flat files for lines to delete. And some of that data would have been processed through a system like splunk or logstash that isn't particularly friendly to deletes and would require a massive re-indexing operation to flush the necessary records. And some of those systems probably have tiered storage that includes offline, slow to recover archives built with cost assessments that did not account for frequent data removal. (Think about how much it would cost to download 500TB from glacier, decompress it, decrypt it, put it on an active system, remove a few records, re-encrypt, re-compress and re-upload it 4 or 5 times a month). And think of how that cost compares with "we generally need to reference maybe 1gb of data every three or four months".
The law is quite broad and fuzzy, which is what makes technical people uneasy:
‘personal data’ means any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person;
It is entirely possible that in your message board example, you've deleted the user, you've deleted their postings, and are still in violation. For one, the person requesting deletion doesn't even have to be your user. Someone else may have posted such information about them on your board.
People misconstrue GDPR to be only about databases and unlinking users' records from tables (mainly because it's the easiest thing to do). But it really is about all and any personal and sensitive information of natural people, full stop.
If what you wrote is true, then organizations are liable for data that someone may, possibly even just theoretically, be able to use to identify someone else.
Given that the law is not 'technical', maybe it'll be interpreted much more leniently than a straightforward reading would lead one to expect.
Elizabeth Denham, UK's information commissioner in charge of data protection enforcement, had this to say:
"Having larger fines is useful but I think fundamentally what I'm saying is it's scaremongering to suggest that we're going to be making early examples of organisations that breach the law or that fining a top whack is going to become the norm. Our office will be more lenient on companies that have shown awareness of the GDPR and tried to implement it, when compared to those that haven't made any effort."
In reality, nobody knows how GDPR will pan out exactly (including the authorities).
But if someone was a user of your service and your services included say photo or video hosting, and you delete their name, their phone number and such, but all of the data you don’t delete still reference a common user ID, then that all of your data can still be related back to an identifiable person.
If instead of holding reference to the user id on the data themselves you have separate table for data ownership that is one step better.
I.e you have all data stored separately and each piece of data has a guid, and you have the user profile separate with a guid and then you have separate tables tying data guid to user guid and when you delete a profile you delete both the user profile and the tables that tie data to that user profile guid.
However I think even that is not enough.
Yes it’s hard to implement full deletion, but you are the one that chose to accept data in the first place. If you can’t implement a system where the data can be deleted you shouldn’t be accepting that data in the first place IMO.
Because even if you delete the references to the data, the individual pieces of data can still be tied back together with data analysis. For example by looking at meta data in pictures, or looking for artifacts that come from lens scratches.
Also, even if you delete references of ownership of an image, probably you aren’t deleting records of comments that other people made on the photo, because if you went to that extent you probably could have just deleted the data proper, so then with access to the image and knowing what other people commented it will likely be possible to determine who originally uploaded that photo.
[1] https://www.quora.com/Is-a-photo-of-a-person-considered-PII-...
It does seem that interpretations handling user-generated content will be tricky, and likely would require additional legislation to clarify that, which is a common scenario for such laws.
7. ‘controller’ means the natural or legal person, public authority, agency or other body which, alone or jointly with others, determines the purposes and means of the processing of personal data; where the purposes and means of such processing are determined by Union or Member State law, the controller or the specific criteria for its nomination may be provided for by Union or Member State law;
8. ‘processor’ means a natural or legal person, public authority, agency or other body which processes personal data on behalf of the controller;
You are right though that GDPR doesn't apply to non-business individuals, but that doesn't make the image host the data controller. Maybe the user and the host are data controllers 'jointly'? Dropbox is quite clear that at least in their Business and Education accounts, the account owners are data controllers for all their data.
For example you could store the body in a content addressed store (a key value store whose keys are hashes of the value). This preserves the ideal model of immutable events as you cannot accidentally change the content of the event since that would violate the property of the CAS. But nothing prevents you from deleting an entry in the CAS.
However you still have to find all the records you need to delete!
The trick with forgetting the encryption key let makes this step a O(1) operation.
On the other hand it forces you to decide upfront the granularity of ownership, i.e. if you treat payload deletion as a batch process you can deal with mistakes, change your mind about which data should or should not be deleted. While with the encryption technique, changing your mind would require you to perform a full history replay and reencrypt data in a new stream (and throw away the old one) with the new granularity model.
https://blog.varonis.com/gdpr-requirements-list-in-plain-eng...