these researchers should be ashamed
these researchers should be ashamed
These are social spaces where a lot of young people essentially grew up. An important part of social development is making mistakes and learning from them. How can you make mistakes when those mistakes are archived for all time for everyone to see?
Similarly, we have a huge problem right now with massive partisanship in the west. Changing one's opinion should be viewed in a positive light, but unfortunately our society doesn't seem to see it that way. Someone with an odious opinion when they were young, who changes that opinion to something more moderate when they grow up, should be viewed as a positive change. But increasingly what we're seeing is people going through data sets to find those old odious opinions of somebody when they were 14 years old and using that as proof that they must still be a terrible person at 24. It's a paranoid, completely self-defeating worldview, but unfortunately it's all too common right now, and I think it's honestly a huge reason for the massive political polarization we're seeing in the moment.
So yes, shame on these researchers. I know they claim to have anonymized the data set, but let's be honest, that never works. It's always easy to find common threads and identify someone, and it's particularly easy now when we have access to all sorts of machine learning models that can really do very effective denonymization.
You'd have a more convincing argument if you said something like "oh these servers have the implication of semi-private chats so people may be more inclined to share personal information" or something.
Otherwise, let me play on the worlds smallest violin for the poor massive corpo when people dont obey their 500+ page long legalese ToS designed to maximise ownership over each user.
I do.
>Do you similarly get upset when someone violates the Facebook ToS?
I don't get upset, but I recognize that they would be breaking the rules.
>Otherwise, let me play on the worlds smallest violin for the poor massive corpo
Remember the golden rule. If you want agreements you make with others to be upheld then you should respect agreements others make with you.
Besides, most of the salacious talk happens in DMs and private channels, which weren’t scraped.
If you can find evidence of some way to use this data harmfully, I’ll agree with you. Till then you’re making a fuss possibly without merit.
In general I’m skeptical of the power of data alone to meaningfully harm someone, except in obvious cases like private health info, exposing affairs, financial documents, and so on.
Public information is public, and it’s arguably wrong to keep it locked up rather than the other way around. Datasets like this are the only way open source ML has any chance against the big players.
Bluesky had a similar freakout when HuggingFace packaged up a snapshot of the firehose and this discussion was had ad infinitum - in Bluesky's face every post is explicitly public -- but a very vocal minority of users still felt they had been wronged because no one asked if they were OK with being a part of a research project. There's definitely a gap between what users /should/ be aware of and what impression they actually get using a service, and you have to keep in mind most people are not techies that have spent years of their life wrestling with issues of privacy and data ownership, so announcements like this can come as a surprise, the first time they considered that what they said in one context can be moved to another.
Plus, even if the server is in Discovery (and thus really publicly advertised), they’re still mostly sorted into 10+ rooms. Just because a library is public, doesn’t mean I should expect Study Room B will be recorded by John Jameson and entered into a public dataset.
Just because something is public means that someone is actually going to save it for later en-masse.
Educating 14-year-old kids so that they don't post public chat with their real name is more important than shaming the researchers.
"It's easy to find common threads and identify someone?" Prove it because I don't think it's that easy.
Through the rest of us abandoning the, let’s call it, presumption of endorsement: that you having said or even done something stupid ten or five years ago (or even less, if you’re young) means you still endorse it now. Right now it feels like things are moving in the opposite direction, extending that presumption from things you said to things you allowed others to say, to things others said elsewhere that are entirely unrelated to things you allowed them to say in your presence.
I have to seriously question the actual prevalence of this. If someone posts and says "Look at what this person posted 12 years ago!" I'm not going to take them very seriously. I don't know anyone who would. This sounds like the usual "cancel culture" stuff that mostly boils down to "people face consequences for their present shitty behavior"
Anyway, as an avid Discord user, manager of a Discord community, and privacy advocate, I sympathize with this position, but as a software user and developer, and advocate for accessible FOSS (and accompanying information), I unfortunately have to side with the release of this and similar datasets. I would much rather this than the inevitable loss of so much valuable information.
> These are social spaces where a lot of young people essentially grew up. An important part of social development is making mistakes and learning from them. How can you make mistakes when those mistakes are archived for all time for everyone to see?
Understandable, but really, if you want to solve that, then you're up against all of social media. The only difference is that Discord wants you to make a free 2 minutes account so you can join the public server to look at what they said, instead of putting it on Google.
Let's bring back the concept of reputation.
the point isn't to make an artifact, like stack overflow. and certainly not be be experimented on.
It very clearly does not.
The fact remains that if you post something in a private discord channel that someone takes public, it's now public information. All these researchers did is expose that fact—the fact that private messages can later be made public is Discord's responsibility and Discord's shame.
Unlike Reddit and Twitter, Discord was never meant to be a space where your contributions are intended to be publicly viewable. People forget that while Discord is oftentimes used as a replacement for forums, it is actually the spiritual successor to IRC, AIM, and similar chat services, where the data is typically ephemeral. Message history in these services is still a fairly recent addition, and what we're seeing now is the consequence of that data being available.
Some people think that availability of historical data inherently puts it in the same camp as forums, Reddit, and Twitter, but I don't share that view. Discord data is still intended for present members of the given servers, not the public at large, even if anyone in the public space is able to become a present member. The distinction there is a meaningful one.
Is it? Discord's been around for a decade; as far as I'm aware, it's always had message history, unlike IRC.
> Some people think that availability of historical data inherently puts it in the same camp as forums, Reddit, and Twitter, but I don't share that view. Discord data is still intended for present members of the given servers, not the public at large, even if anyone in the public space is able to become a present member. The distinction there is a meaningful one.
I definitely agree with this, although the fact that anyone joining can scroll back does sort of mean it's available to anyone who wants to look for it - it's just difficult to do so.
Almost a decade, yeah. And Slack had similar functionality in the few years before Discord arrived. But chat services have been around since, what, the late 80s or early 90s? Services like Discord are still the new kid on the block. Now we're seeing it being taken advantage of (or at least the first instance of someone publicly and proudly taking advantage of it), prompting curiosity of what comes next: Will this disregard for privacy simply become the default expectation, will people shift to more private spaces less susceptible to abuse, or privacy regulations protecting people?
I can't speak for AIM, but I've never assumed IRC was ephemeral. Most IRC servers had and have numerous people idling in the channel with their client just archiving everything that happened. Many of those archives ended up published on the internet and that usually surprised no one.
Discord is even worse because it does the archiving for you and grants new users immediate access to the entire archive.
As a rule, each and every recipient of every message that you send has the option to archive or forward that message. That's a simple fact of information flow that everyone in the modern world must wrap their heads around. In this case, when you send a message to Discord you're transitively sending it to everyone who is currently or ever will be on a given server. That has always been true, and that fact has been exploited ever since Discord came about.
All these researchers have done is expose the kind of archiving was always possible and always happening.
> All these researchers have done is expose the kind of archiving was always possible and always happening.
What they've done is brought forward another instance of "this is why we can't have nice things". A useful feature being used in unintended ways without regard for ethics and privacy. The authors of this paper did not have respect for people's data and privacy. They did the bare minimum to make the claim they took privacy measures, regardless of whether the claim is true or not. I guess that's enough for them.
All these researchers have done is admit to doing what many less well-meaning people and organizations are already doing on a regular basis. If that admission leads to more people realizing that public Discord servers are public, then the researchers have done everyone a service.
Blaming the researchers in this case is like blaming LockPickingLawyer for locks being pickable. The bad guys were already picking the locks before he started his channel, all he did was shine a light on just how bad most locks are.
On the contrary, kudos to these researchers for bursting the illusion that people previously had that things said on public servers somehow would be ephemeral. That's not how the internet works, that's not how it's ever worked. If you send something to someone else's computer you have always had to assume that every recipient could have made a copy of it, and when the recipient list is "everyone who ever joins a public Discord server from now until the end of time" that makes it public information.
Better that it be widely recognized and talked about as these researchers are doing than have accessing public data remain a dark art that lay people mistakenly believe can't happen.
but people like these researchers who only want to exploit are making it necessary
1. The illusion of intimacy
2. The illusion of ephemerality
People think they chat to a small number of people, and people think that it's going to go away at some point.
So they think they're in a private conversation when they're not. They wouldn't behave the same if they realized what they say is being written down and stored forever in some database.
And yes, it sucks that technology doesn't allow you to have it both ways => public because you want to reach far and wide, and private cause you don't want it recorded.
Look at it like this: There's no chance in hell that intelligence agencies, hacker groups, and whatever other nasties you care to worry about haven't already been using archives just like this for all their nefarious purposes. They just didn't make their usage public because why break the honey pot?
What these researchers did is show what was possible and make their efforts public. Now you are better informed of what was always possible. It's always been necessary to think carefully before putting stuff on the internet, it's just now your bubble is burst.