-Riley Goodside
https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldm...
The idea that tokens found closest to the centroid are those that have moved the least from their initialisations during their training (because whatever it was that caused them to be tokens was curated out of their training corpus) was originally suggested to us by Stuart Armstrong. He suggested we might be seeing something analogous to "divide-by-zero" errors with these glitches. However, we've ruled that out.
The counting Reddit usernames are clearly a major source of glitch tokens, there's something about that sub that screws up the model, maybe the unusually predictable/repetitive nature of what goes on there or the sheer obsessiveness of the posters.
The fact that the only non-counting Reddit usernames are both people at the center of massive Bitcoin psychodramas is suggestive however, especially given the associations with the petertodd token. My guess is that these names appear much more frequently in other people's posts than is normal for forum usernames (especially if it read bitcointalk.org and twitter too), and that maybe this has some effect.
To get an idea of just how frequently this token will crop up in the reddit corpus, just do a search for it:
https://www.reddit.com/search/?q=%22petertodd%22&sort=commen...
I'm not sure how many posts match but you can keep scrolling a long time.
Having been more exposed than normal to what happened back then I wasn't hugely surprised by what concepts clustered there, and sadly don't think it's actually random or unexplainable. Flick through discussions pre-2015 (when the Bitcoin forums started to be heavily censored) and you'll see a lot of very similar words as what GPT spits out being used in connection. It was all very nasty.
Do you have a source for that or was it just an assumption?
Also, that Reddit is frequently used to train LLMs is widely known. It's an unusually clean source of conversational text because you can slice threads (i.e. pick a root comment, then pick a child, then a child of the child etc and then concatenate the results), and you'll get a coherent conversation. There are relatively few places on the internet where that is true. For example most phpBB forums conflate many different conversations into single threads, with ad-hoc quoting being used to disambiguate which post is replying to which. That makes it a lot harder to generate sample conversations from.
Imageboards.
DailyMail.
Slashdot.
Even a somethingawful dump would have been superior.
The Daily Mail (the newspaper) has been used for training LLMs in the past, yes. I don't know if it still is.
listen, some of the niche corners of that world aren't so bad, but it ain't the place to be training AI to do something, unless that something is a hate crime
It's very much "he who has the power to destroy a thing controls that thing"; so long as there's only one canonical copy of a post and it's on Reddit's servers, that's vulnerable to actions by the subreddit mods or Reddit themselves.