How to Make Custom AI-Generated Text with GPT-2
minimaxir.com
minimaxir.com
"Each text is analyzed by how likely each word would be the predicted word given the context to the left. If the actual used word would be in the Top 10 predicted words the background is colored green, for Top 100 in yellow, Top 1000 red, otherwise violet."
Generally, whenever your detection strategy is "a spam generator would never do X", I simply update my spam generator to do X. (Note that "X" must be something that is relatively easy to calculate, not things like "this actually makes sense", because it must be something your detector can recognize.)
Also, if Google suddenly started penalizing all text that doesn't "seem natural", there would be tons of false positives: languages other than English, jargon-heavy websites, dyslectic people, etc.
I do have ideas for more practical content generation but those do not currently have plans to be made public.
They are grammatically correct but logically nonsense. The examples shown are also curated from thousands of generated text.
The only use for these is the seo landfill to get in top of Google
In part because we're trying to probe what it even means to "understand" something. If an AI is capable of carrying on a cogent conversation with you, does it understand what it's saying? How do you know?
It is slightly better now than it used to be, i.e. you now have to read two sentences to realize it's incoherent gibberish, rather than just one. But thinking you're getting closer to understanding this way is like the metaphor of climbing a tree to reach the moon: you can keep reporting progress until you run out of tree.
As for it being entertaining, sure, the first attempts were entertaining and instructive. Now, years later, reading more gibberish from a neural net on HN, not so much.
Hugging Face is another such company making massive impact among non-NLP developers to use such resources.
Kudos to these guys!
I have no advance knowledge of any GPT-2 related happenings, which has become inconvenient as OpenAI likes to release things when I am on vacation. :P
I quickly glanced through it but will do a deeper dive later on.
The outputted paragraphs don’t seem to realistic. However you point out that the text needs to be longer for a more a natural output.
I am genuinely curious how relalistic it can get.
In my personal experience I get fooled by the bots far too often, it's actually really scary, I can imagine it being used nefariously to push a specific political agenda, spam places, do SEO fraud etc. etc.