The OKCupid data release fiasco
socialmediacollective.org
socialmediacollective.org
The other important point is, that there is a shift to questions instead of answers. Technology moves just too fast to give good guidelines in terms of rules. It would perhaps help to define a questionnaire that researches have to fill out, just to demonstrate that they thought about the question. For example, the question "Which groups are especially vulnerable to a leak of the data?" may prompt filtering of said groups and prevent some damage. (That obviously does not help against black hats, but would probably helped in the current OKCupid case.)
And lastly I think that there is a mismatch what "reasonable expectation of privacy" means, most people do not have any idea what "big data" means, they believe that Facebook reviews flagged comments, but do not understand that they can (and I would guess do) analyze each and every interaction with Facebook. (I don't mean to single out FB, it is just they are the most obvious example due to their size and the sort of private nature of the network.)
First, I don't get why sitting in one's bedroom and writing in a public forum illustrate the clear distinction between public and private spaces. You already noted the forum is public. IT IS PUBLIC. If you put up your web page on the Internet and only search engine would ever come to your page, well, your page is still the public domain. Being in a small community, take a cult, well, if you show up in the public and doing your activity in the public, I would expect that to be a public demonstration. In fact, if you try to run naked with your cult in the middle of the street, you are going to get arrested and charged for public misconduct. If you are running naked in a cult house, well, that's private. The distinction is very clear in these examples. I am not arguing there aren't examples that would illustrate a blur, but your example is too weak.
My response to the second suggestion. Well, leak of data is a big deal. So if my research somehow harms people with AIDS, am I allowed to collect the non-AID patients? Another way to put it, filtering doesn't do any good and provides no meaning. From an ethic standpoint, people aren't concerned about what groups will be harmed from the outcome of the research - because the whole point of research is to either validate or invalidate some hypothesis. Research ethic should be about the treatment of data and the process of conducting experiment.
However, I do agree that some users aren't aware of what online service do with their application.
The problem here is that the data included in the release (answers to questions) can only be accessed by logged-in users, and the scraper they used utilizes a login, in which case the researchers blatantly violated the ToS and a DMCA takedown of the dataset was a valid response by OKC.
However, almost everything else in an OKC profile, including username, city, and sexual orientation, is public to logged out users by default. (see this analysis of mass-scraped OKC data: http://yakamo.org/?p=112). That makes things a grayer area than usual, and serves as a reminder that nothing online is private.
"I setup 7 servers with mpi4py and it took 48 hours to scrape 3.2million profiles. Now i was ready to get the answer and find out how many profiles where public."
I think that works out to around 20 requests/sec. At a certain point it becomes borderline DoS?
I only divided by 7 because I thought he wasn't running all the 7 servers simultaneously, come to think of it, I was wrong.
Or at least anyones liking and anyones doing things non the less (to not become too personal here).
No, this is an abuse of the DMCA. The DMCA only applies where there is copyright. Databases of facts are not copyrightable in the US. See Feist vs. Rural Telephone.[1]
Violation of the terms of service may be grounds for a civil suit, but it is not a criminal offense. That was settled a few years ago.[2]
[1] https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R....
[2] https://www.eff.org/deeplinks/2010/07/court-violating-terms-...
OKCupid is a collector, not an creator.
I wondered, wasn't the data still technically private data that was just stolen and released? And wouldn't it be against the law to then download a copy of that data yourself? But it was actually the legal department that downloaded the dump, so I didn't have much to say.
Someone explained this to me by saying "If my email signature said 'by receiving this email, you agree to pay me $5', would that be considered a legal contract?"
(By reading this comment you agree to pay me $100.)
It does not generalize to other professions/pursuits; it would be like saying "well, the person who reads the parking meters is allowed to issue parking tickets, so anyone else should be too" -- the person issuing the tickets has specific-to-them legal permission to do so.
It was certainly unethical for the researcher to release the data, but the public really really really need to be educated about putting information on a publicly available and searchable database.
It seems, though, that this discussion is overlooking the fact that there are companies out there which are gathering this exact same data(And then some) to sell to other organizations who want to identify their users.
In your first point you mention adjusting volume to restrict communication to your desired audience. OKCupid has privacy settings. Users adjusting their privacy settings would be analogous with adjusting volume in the restaurant. The people collecting data in this story created a fake account and used it to bypass the privacy settings. For your second point, it would be more like if you filled out a restaurant survey and saw your answers on an unrelated website like Reddit the next day.
Anyone accessing websites with their names, emails attached to their names, emails attached to their phone numbers attached to their names, birthdays in their handles, handles that aren't random strings, identifying passwords, avatars in use on other accounts, images also hosted on flickr, g+, insta, etc. should take it for fucking granted that there is now zero reasonable expectance of privacy.
We had the chance for the internet to be a beautiful, anonymous playground, but a few too many people showed up dead on the chan and now the crybabies with body image issues have managed to make all of the above borderline required in the name of stopping "cyber-bullying" and "terrorism."
(try to sign up for an account, say a meetup account -- requires an email. Burn emails are blocked for "spam." So you need a gmail. Gmail requires phone number. SIM requires passport identification in europe... > meetup.com requires passport identification... a few alternatives still exist, mail.com for example doesn't require phone verification... yet. Facebook requires an actual scan of the passport if you put in estimably fake information.)
And with that thought paradigm, that your ID is a requirement of internetting, comes stupid people getting butthurt by shit like this "oh, but I put forward all this identifying information with the expectation that I would be anonymous, oh, oh, oh".
All this because some bitch on youtube can't handle the fact that obesity is not in fact a disease, and because something that kills less people than accidentally suffocating in bed has become an excuse to usher in the police states of america.
There was a better way, where this this would never happen. But we as a people opted to go this new direction. Wear a fucking jacket if it's cold and remember _you_ wanted to vacation on the baltic beaches in march, not me.
there is data that is only available if you are a member of the site and answer questions on the site. for example if a question asks you, "would you date a smoker?" if you answer you can post it so other members can see what you answered to gauge whether what you picked matches with what they liked. what this guy did was go through all the questions and answer them at random then scraped the data that was available only to actual members of the site.
as for profiles being "public". they can be "public" in the sense that you can go to a url and see their profile, yes. but most people set it so only people who have signed up can see them. so by doing this he most likely violated the TOS of the site and scraped the data without permission.
edit: it is not as though it was public, the guy who scraped the data made effort to log into an account, answer all the questions and then _scrape_ the data. it's not like that data is plainly visible if you go to someone's profile.
* Requires effort to access
* Is in out-of-the-way places many people don't know about
* Is in a form that's difficult to consume
Even though it's "public", in many ways it behaves as if it isn't, since until one of those conditions is removed people don't really bother to access it, pay attention to it, or try to consume it.
In other words, it's one of those "if a tree falls in the forest and nobody's around to hear it" situations.
So information which is only available to someone with an account, but on a system which grants accounts for free to anyone who wants one, typically is thus seen as "private" or at least in a "not really public" state until such time as someone massively aggregates it and makes it available to people without accounts, as happened with OKCupid.
Another issue has to do with the expected lifetime of the information. Users on OKCupid expect that when they remove their accounts the information is no longer available to anyone.
This is all about what tools enable.
This issue is critically important to people who have a much higher privacy need than the average person. Consider an abusive ex tracking down their victim in an unknown new city.
If the profile is indexable, the abuser is trying to either use OKCupid's search (which returns results with a high degree of randomness), or using things like "site:okcupid.com $interest1 $interest2", which is noisy, to say the least. This task looks hopeless quickly.
But if someone provides a database of profiles, the job gets much more tractable. Suddenly, you've enabled the attacker to filter by people who have 1,2,n interests of their ex, search with photographs of their ex, hell, even use facial recognition. Some of these things were kind of possible before, but with a database they are likely. If someone builds a "find your ex" web frontend on this database that uses facial recognition technology, it becomes accessible to everyone, and maybe even popular.
Saying that this lesson should be educational is theoretically nice, but doesn't match how users interact with products in the real world. Anyone who has worked on a product with millions of users knows that the average user (a) doesn't know that there are settings (b) they don't know which things are controlled by settings (c) they don't know how to find the settings, and (d) they don't know how to change them. This is so common that I no longer believe that users are responsible for the implications of default settings (and OKCupid's defaults are, inconveniently enough, Public).
Tools usually are capable of both good and bad things, but a database of profiles has limited upside, and a huge downside. Like, this thing had better be curing a minor fatal disease for the amount of trouble that it will cause.
To take an analogy: if a child plays with matches, you teach the child not to do that, explain how fire is dangerous, etc. You don't light a fire and hold the child's face down on it to burn them and make the lesson "painful". Yet when it comes to deliberately making sensitive/intimate data about people public that basically is what you're advocating for: the potential emotional harm that can be done to the people involved is off the charts.