Insights into a corpus of 2.5M news headlines
freedom-to-tinker.com
freedom-to-tinker.com
"The classifier was trained by the developer using Buzzfeed headlines as clickbait and New York Times headlines as non-clickbait."
I've thought that Times article headlines have been more clickbaity over time. Some articles, especially if they aren't related to current events, are becoming more casual in their voice as well.
If they can maintain quality journalism by only sacrificing headline quality, that is probably a good and necessary tradeoff to make.
Allowed you to pull in each article on the various news sites and stick them into git where you could then see the difference in the articles over time.
A Phd student founds some interesting insights particularly in the Irish media where content would often be changed sometimes weeks / months / years after the original story was placed online.
https://github.com/johnl/news-sniffer
Often though; I found it showed how little time a journalist often had time to work on revisions to any one story....the world should lament the death of journalism.
> finding bias in headlines is a more subjective exercise than finding it in Wikipedia articles...
Probably because bias from new Wikipedia users is less careful. And absent smoking guns, once you get into the gray areas, several studies have shown bias is often in the eye of the beholder:
I'm on the other end of things though, I would really like a tool which can generate clickbait titles based on the article body, or even better, generate the whole fluff piece based on some keywords. Something similar to https://pdos.csail.mit.edu/archive/scigen/ maybe
I think the research may need to take a second look at the algos used. Would have been interesting to have the Economist in there as well.