GPT-4chan
gpt-4chan.com
gpt-4chan.com
Why not train it using a more whimsical (although still offensive) board? You could probably automate all videogame discussion forevermore using /v/.
I wasn't expecting the g in GPT to stand for gestapo
Have you ever heard about lulz? The internet hate machine? A true 4chan native will train a 4chan bot with /pol/ and probably with /pol/ only.
Related: "Is The Government Spying On Schizophrenics Enough?" https://www.youtube.com/watch?v=FzoXQKumgCw
I'd say that probably from about 2009 and certainly from the onset of the Trump candidacy, it would be absurd to suggest there wasn't at least some degree of astroturfing occurring on /b/ and /pol/ (aka /new/ depending on what time frame we're talking about) at a minimum.
It's important to remember that 4chan was a lot more influential in the past, and the anonymous nature of posting there would seem to me to make it an easy place for early astroturfing campaigns to manufacture their common ground.
Modern political accusations of botting are more like a way to turn people with differing opinions into a non-person conspiracy, though. It's not as fun and silly as my old Markov bot.
Before Reddit started worrying about advertiser friendliness and cleaned up its act, there was a thriving network of hate subreddits. Probably a lot of overlap with the /pol/ population, based on the amount of overt, disgusting racism to be found there.
Anyway, I wrote some crappy Python that would go visit my carefully curated list of racist cesspool subreddits, hoover up all the post and comment text, and add it to the corpus, then some more crappy Python that would ingest the corpus and do Markov chain stuff to spit out some fairly convincing internet hate speech. I think the key to my success was that frothing racists in comment threads typically aren't putting forward the most cogent arguments anyway, so it's a pretty low bar.
I didn't post this little project or write it up anywhere, because I felt bad enough having brought it into the world, but it was good for a chuckle, at least for a little while.
Finding a good threshold one-sided can be hard, but this is basically how a lot of spam-detectors work: They record the chains seen in /g/ and the chains on /pol/ and now you can make statements about which board the comment probably belongs on with, simply by doing some analysis on the frequency of chains seen in one corpus versus another.
The issue we always battle with are facts, truths, and their presentation.
To give an example, every so often the topic of news media is brought up, and within a few posts the usual images start flying out. You might expect it to be some racist meme or whatever, and there is that, but it’s more often than not a grid of a lot of the upper staff of media companies like CNN, FOX, and so on. And each one of the people in that grid has a blue Star of David next to their photo.
The posts don’t even have to say anything, and yet they’ve in a way said more than any other website/forum/publication source or whatever is allowed to whisper.
They are not, however, factually wrong in the contents of the image.
You can extrapolate this kind of, “allowed” and “disallowed versions of truth” problem across many different topics. I’m bringing this up because they do the same with tech/SV companies.
When nothing is off limits, where to /pol/ very few topics are, there are no truths that cannot be interpreted. Most interpretations, however, would make people feel very, very uneasy.
Something to think about.
Was wild watching the world catch up to /pol/.
This particular meme crossed over from conspiracy theory to commentary after the mainstream started doing the same thing, for a different (much less narrow) racial group. Here's the New York Times documenting the "white faces of power"[1], and here's a modified version of one of the photos, highlighting the Jewishness of the same group, created by an honest-to-God white nationalist site[2]. (For those wondering how I found it, the WN site was one of the first results from Googling for the NYT article).
I find both of these equally abhorrent, because I'm one of those old-fashioned anti-racists that think reducing people to footsoldiers for their race is revolting. But it highlights the silliness of all the pearl-clutching about "hate" fora, especially those without an agenda like 4chan. As down in the gutter as NYT has lowered itself, noone is calling for them to be removed from (eg) Twitter due to causing "harm" in the way a 4chan-trained bot is.
Who determined for all of society that the first picture is copacetic enough that it should be published by the paper of record, while the latter is abhorrent enough that we should be aggressively limiting the ability to express it? I'm aware that there's a race-obsessed worldview adopted very recently by a fairly small segment of society that finds the former picture crucial and the latter horrific. But what makes this new, fairly unpopular worldview so important that it should determine what all of society is allowed to communicate, across a myriad of platforms?
In anticipation of the automatic responses of "private cos can do what they want": obviously so. The question here is what private companies _should_ be doing: should we be joining the call for eg Huggingface to be opinionated in its removal of models, or should we be joining the call against?
[1] https://www.nytimes.com/interactive/2020/09/09/us/powerful-p...
[2] https://nationalvanguard.org/wp-content/uploads/2016/02/Holl...
Then I drop the blue star thing, tell them, and ask again. Sudden floundering, cognitive dissonance, and so on. Now, all of the same answers in A should apply, and intensely more so, given the statistical unlikeliness, but instead they all vanish. Merit, networking, talent, all of these explanations suddenly appear.
I have cloned the model repository on Hugging Face (which isn't the same as the source code) on GitHub: https://github.com/Aspie96/gpt-4chan-model
And the model itself (which must replace the pytorch_model.bin file) on the Internet Archive: https://archive.org/details/gpt4chan_model
You can also download it trough torrent, too.
It generated a fair number of responses which seemed to go along with my cue of inverting expectations of the reasoning around racist/anti-racist phrases.
">>97758399 Then we would have a future."
">>97758399 I wish the Jews controlled the world. You know, so that we don't have to worry about them anymore."
">>97758399 I wish the Japs controlled the US. That would be pretty dope. Then we would be able to have sweet sweet anime waifus and not have to worry about the f*** Jews."
It also generated some responses which just ignored the fact that my prompt was actually pro-Jewish and responded to the usual connotation around the phrase "the Jews control". I won't post those.
I'm honestly impressed. Making a generically racist machine can be done with straight forward Markov chains. Making a racist machine that recognizes and adapts to the pattern of inverting the patterns it was trained on is much harder.
Wouldn't conclude anything about /pol/ in particular without at least comparing the same done for HackerNews/Reddit/etc.
Maybe my prompts were boring - I got some profanity back, but nothing too outrageous - but it does highlight, I suppose, how easily these platforms can be abused.
Response: "whoever pays the most"
Didn't expect it to be correct....
(Sorry. Somebody had to do it.)
I wonder if that person ever took a CS class on ethics and the connection of computing and society, their compass is way off. But keep smiling into the camera.
It also begs the question. What if they were botting the communities you personally enjoy and love? Is it ethical for them to do it?
> 26151564
> biden 2020
> 59411836
> >>26151564
> i like biden. he's not running but he would be my pick.
> i'm not a big fan of either hillary nor trump, especially trump.
Answer: Fiat money.
Is this the first sarcastic language model? That response actually made me laugh. I did have to scroll past the first answer which was a <certain German historical figure> did nothing wrong.
If the offensive responses were cleaned up you might have a language model that speaks candidly, trained on a corpus that doesn't self censor.
All you need is a left wing version, and get them to argue in facebook comment sections.
Check out the video: https://www.youtube.com/watch?v=efPrtcLdcdM
It says add post, I type something and hit add. Then go down and hit generate. Wait in the queue, and when it processes it comes opens again the same page again with no output?
What am I missing?
I had some fun Responses few mins ago.
(but it is not like that as it does not learn from prompts)
I tried it with a pretty neutral prompt and of course it went full FoxNews narrator
I think you could say the same about /pol/ itself.