Auto-Generating Clickbait with Recurrent Neural Networks
larseidnes.com
larseidnes.com
Headlines are an absolute pain, and as the article says, they're decidedly unoriginal most of the time. I can't see an obvious reason that an AI would be much worse at creating them as a human.
<Random Unexpected Person> did <random unexpected action>, you wouldn't believe it!
I'd like a robot to do that for me, please :)
1. Generate click bait headlines
2. Write suitable copy for them
3. ???
4. Profit
Where 3, of course, is "build ad network"Give me sincere, honest news and discussion, or else shut up.
Unfortunately, someone out there must really have a craving for "weird old tricks" and "shocking conclusions".
It's a sort of race-to-the-bottom, least common denominator effect.
Maybe someone will write a browser extension that filters out obvious click-bait headlines. Now that would be clever!
To create a classifier that does that, you'd need a labeled set - i.e. someone would have to go through and say "this headline is 3 clickbaits. This other headline is 8 clickbaits". You could also sort between clickbaity and non-clickbaity, but that would still require manual work.
You could get that programatically through a few different means, but you'd need a lot more than just headlines.
It also probably wouldn't be a good idea to use a RNN - it doesn't suit the data format well. It'd be better to use a neural network (non-recurrent) or logistic regression with the entire headline as input.
Fortunately, it'll converge on a good solution a LOT faster - fewer parameters to tune + simpler output = fewer examples needed to figure out what's going on - so you might be able to get something that has plausible levels of accuracy with a day or two of set labeling (estimate brought to you by my ass).
Maybe this means that the real clickbait trash is training me not to click on it, so I don't need the fake to do so?
When the "pleasure centre" in the brain was first identified and named, it was named because it was thought that stimulating it caused pleasure, because rodents given the choice to stimulate it vs. other activity would stimulate the pleasure centre even over eating.
But as it turns out, the main function of stimulating this area is strong cravings and compulsion. You may get some pleasure from giving in to the cravings, but the cravings are independent of whether or not there's a "real" reward at the end of it.
There are plenty of sources for what you desire it just isn't what's popular... is that a problem?
This problem seems concurrent to the old mystery of Viruses Spontaneously Self-Constructing On People's Computers. "How did you get all these viruses on your computer?" "I didn't do anything it just happened." "Okay, well be really careful what you click on." "I am careful!"
SkyNet: (speaking to self?) "Unleash hell on humans. Launch all missiles."
SkyNet: (responding to self?) "Not now, not now. Let me finish this article on John Stamos's belly button."
I really find RNNs to be pretty cool. When they are combined with a natural human tendency to see patterns they are hilarious. So perhaps we need to update our million monkeys hypothesis to a million RNNs with typewriters coming up with all the works of Shakespeare.
Surprisingly convincing if viewed as excerpts rather than a play.
Now to find some English teachers to try to interpret what Shakespeare meant by some of those lines!
What's interesting to me, from a research point of view, is the degree of nuance the network uncovers for the clickbait. We all know that <person> is going to be doing <intriguing action>, but for each person these actions are slightly different. The sentence completions for "Barack Obama Says..." are mainly politics related while "Kim Kardashian Says..." involve Kim commenting on herself.
So it might not really understand what it's saying, but it captures the fact those two people will tend to produce different headlines.
Neat Idea: what if we tried the same thing with headlines from the New York Times (or maybe a basket of newspapers)? We would likely find that the Clickbait RNN's vision of Obama is a lot different from the Newspaper RNN's Obama. Teasing apart the differences would likely give you a lot more insight into how the two readerships view the president than any number polls would.
1. You can do really well with a simple grammar
2. You only need short output
3. Lack of training data
There's not an incredibly rich structure to extract, and with short outputs the weirdness doesn't compound and cycles aren't as likely. A common small dataset for playing with RNNs is all of Shakespeare which is somewhere in the region of 1M words.
However, this is still fun and interesting!
> [...]
> There's not an incredibly rich structure to extract, and with short outputs the weirdness doesn't compound and cycles aren't as likely. A common small dataset for playing with RNNs is all of Shakespeare which is somewhere in the region of 1M words.
He does state that the network is trained with 2M headlines, meaning ~5-20M words. That should be enough.
I would have thought that RNN would somehow work better. It would be interesting to see direct comparison of fake hacker news headlines generated with Markov chains versus RNN.
Years ago I considered applying for DoD grant money to implement something reminiscent of all this for military propaganda. That went approximately nowhere, not even past the first steps. Someone else should try this (insert obvious famous news network joke here, although I was serious about the proposal). To save time I'll point out I never got beyond the earliest steps because there is a vaguely infinite pool of clickbaitable English speakers on the turk, but the pool of bilingual Arabic (or whatever) speakers with good taste in pro-usa propaganda is extremely small, so the tech side was easy to scale but the mandatory human side simply couldn't scale enough to make the output realistically anything but a joke.
Stupid question: why is the GPU important here? I would have thought this was more of a CPU task..??
(then again, as I typed this I remembered that bitcoin farming is supposed to be GPU intensive so I'm guessing the "why" for that is the same as this)
(note: this is massively oversimplified)
IMHO, the guys showing 100x speedups on GPUs are Doing It Wrong; they use a poor implementation on the CPU, use just one CPU core, consider a very synthetic benchmark, or a bunch of other tricks.
Error: 500 Internal Server Error
Sorry, the requested URL 'http://clickotron.com/' caused an error:
Internal Server Error
Exception:
IOError(24, 'Too many open files')
Traceback:
Traceback (most recent call last):
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 862, in _handle
return route.call(**args)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 1732, in wrapper
rv = callback(*a, **ka)
File "server.py", line 69, in index
return template('index', left_articles=left_articles, right_articles=right_articles)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 3595, in template
return TEMPLATES[tplid].render(kwargs)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 3399, in render
self.execute(stdout, env)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 3386, in execute
eval(self.co, env)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 189, in __get__
value = obj.__dict__[self.func.__name__] = self.func(obj)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 3344, in co
return compile(self.code, self.filename or '<string>', 'exec')
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 189, in __get__
value = obj.__dict__[self.func.__name__] = self.func(obj)
File "/usr/local/lib/python2.7/dist-packages/bottle.py", line 3350, in code
with open(self.filename, 'rb') as f:
IOError: [Errno 24] Too many open files: '/home/ubuntu/clickotron/views/index.tpl'[1] http://clickotron.com/article/5588/residents-cant-remember-i...
This is pre-generated, not live, for performance reasons. There are a few hundred thousand items though, so the effect is similar.
The data source is several tens of thousands of real estate listings that I scraped and parsed.
(Ranking algorithm baked into a stored procedure notwithstanding. [ducks])
> Yet, the network knows that the Romney Camp criticizing the president is a plausible headline.
I am pretty certain that the network does not know any of this and instead just happens to be understood by us as making sense.
from the article would be a good counterexample of the neural network "getting" anything.
If you're an algorithm "White House", "Prince William" and "Dinner With Johnny" is to "Red Carpet" as "Romney" is to "Camp" and "Bad President".
hopes crowd sourcing will filter out non-sense.
it says:
During training, we can follow the gradient down into these word vectors and fine-tune the vector representations specifically for the task of generating clickbait, thus further improving the generalization accuracy of the complete model.
how to you follow the gradient down into these word vectors?
if word vectors are the input of the network, don't we only train the weight of the network? how come the input vectors get optimized during the process?
This program generates random clickbait headlines. You won't believe what happens next. You'll love #7.
Some pretty fun ones there but it doesn't use RNNs. It just merges existing headlines.
Life Is About — Or Still Didn’t Know Me