What a soul-crushing job that would be.
What a soul-crushing job that would be.
Fun fact:
https://medium.com/i-data/29-reasons-youre-reading-this-arti... "29 reasons you’re reading this article or why odd-length BuzzFeed listicles perform better than even ones"
They published at least 10,000 listicles in 3 months! now, that is what I call a soul-crushing job (writing those articles)
- Engineers are shocked with these results
- Here is the top 10 headlines they found. #6 will make you cringe
- The CEO wrote the sweetest message to them
Perhaps they should simply ban buzzfeed or use all of their headlines as examples.
What a kind, just world that would be.
https://www.buzzfeed.com/andrewrice/the-fall-of-intrade-and-...
"Roman bookies ran numbers on the election of Renaissance popes until Gregory XIV banned the practice on penalty of excommunication. Around the turn of the 20th century, Wall Street brokers openly traded election futures and newspapers quoted their prices like modern opinion polls. Strumpf estimates that at the peak of this practice, in the election of 1916, around $10 million was bet on these markets — more than $200 million in today’s dollars. By the end of the New Deal era, though, the electoral markets had all but disappeared, due to both competition — modern polling pushed the betting lines out of the newspapers — and legal crackdowns."
> Most clickbait is disappointing because it’s a promise of value that isn’t met — the payoff isn’t nearly as good as what the reader imagines,” Patel said. “BuzzFeed headlines pay off particularly well because they actually make fairly small promises and then overdeliver.
https://www.buzzfeed.com/bensmith/why-buzzfeed-doesnt-do-cli...
If you believe what you're doing is important and beneficial, it doesn't feel like too much of a grind. Ultimately this was probably about a week of work, if every labeling participant independently reviewed ~3-5K articles.
Some people might have chosen to mechanical turk this, or to gather feedback from a subset of users and use those labels. Doing that might be a better approach as it might not only be less labor-intensive up-front but also allow the system to be regularly retrained easily (ie by gathering new label information as clickbaiters inevitably try to adapt their headlines). I imagine there were reasons for starting off this way, though.
Paper: 'Antisocial Behaviour In Online Discussion Communities' (Cheng et al., 2015)
Link: http://arxiv.org/pdf/1504.00680v1.pdf
(Search for 'Mechanical Turk' / section Data Preparation -> Measuring text quality.)
I frequently have to read the randomly-sampled tweets for debugging purposes. And, yes, random tweets are often so dumb that my brain slightly regrets the time it spent reading them. But that is far outweighed by the benefit that the Internet is delivering me fresh test data all the time.
On the whole, I enjoy the notion that I am converting stupid babble into something somewhat useful.
You're making a big change to a system used by many millions of people. There's probably no easier way to validate your system is working as expected.
In fairness, the linkbait game has changed since then, with Medium posts using calls-to-actions and "just" needlessly in their headlines. Then again, with FB's machine learning expertise, I'm surprised they need a team to manually classify linkbait posts at all.