Headlines
avc.com
avc.com
Headlines are not one monolithic thing, any more than publishers are.
There are giant worldwide news organizations, there are personal blog publishers, and everything in between.
They all have different constraints and different motivations for their headlines.
Even if you look at just one type of publisher -- a traditional print newspaper with an online presence -- you are likely to find that they write multiple headlines for each story, depending on where that headline will be seen.
For instance, they might write one headline for the print version of the story that's constrained by page layout requirements.
Then then might write another headline that gets displayed on their homepage. It has to be short, punchy, and eye-catching.
Then they might write another longer, more complete, SEO-friendly headline that gets displayed when you actually click on the link to the story.
There might also be an alternate headline that gets displayed as the HTML/meta title of the page, which can be useful for when the link is shared via social media.
And all of those have their specific purposes and limitations.
Another wrinkle: In the case of news organizations, it's highly likely that the person who writes the headline is not the person who writes the article, whereas typically with personal blogs the same person writes both the story text and the title.
So if inaccurate or sensationalistic headlines are one problem, then another problem is treating headlines as if they are all the same. They're not.
What you indicated was that headlines are written in different circumstances. But it doesn't mean that they're not all, roughly, similar.
I betting there exists a strong correlation between clickbait-y-ness and pageviews.
So,
1. Gather all articles with a lot of pageviews
2. Do some NLP to get the really bad ones, i.e. "Look at what this <noun> did after <predicate>"
3. Remove outliers
And bang, you mark all of the sucky articles. That could be included in a Chrome extension, much like uBlock. Hell, why not just have it in uBlock?
Incentives matter, and the underlying problem that gives us clickbait headlines (and fake news, and the rest of our current journalistic ills) is that sites are primarily incentivized to maximize the number of people who show up on the site (so that they load ads and/or can be surveilled) -- hence the chase for "clicks at any cost".
There are two ways to fix this:
1. Somehow change adtech's incentives from "I get paid in proportion to how many clicks I get" to "I get paid in proportion to some other combination of metrics that's a better proxy for quality content (and quality engagement) than 'someone clicked on this and my ads loaded.'"
2. Break the link between advertising and some worthwhile subset of "content that we want to exist in the world" by coming up with a scheme to entice readers to fund the content directly.
I think there's a whole universe of untried startup ideas in both areas. I've been noodling around with an idea for #2 as a side project, and maybe at some point this spring I'll attempt a "Show HN".
Anyway, my ultimate point is that attacking this problem at the level of the content itself -- verifying headlines or rating the "fakeness" of news, etc. -- is the wrong approach, and I think most of the smart people who toss these ideas out know on some level that the proposed cure may end up being worse than the disease.
The fundamental problem is the busted incentive structure. You have to find ways to incentivize the creation of quality content by either rethinking the relationship between advertising and users, or finding a way to get users to pay. There is no third option that's market-based and sustainable.
I don't think there's any going back.
Traffic that's worthless to advertisers -- i.e. I clicked to your page and bounced immediately without actually engaging or really "seeing" any of the ads that loaded -- should be priced at $0.
Traffic that's worthwhile to advertisers -- i.e. I clicked to your page and stayed there for a bit, engrossed in the content and noticing the ads and getting a positive feeling from both -- should be priced somewhere north of $0.
I've been in this business on the content and publishing side for almost 20 years now, and I've seen very many attempts to distinguish worthless traffic from worthwhile traffic. As with the fake news/headlines issue, the above problem persists and is not amenable to a quick technical fix because it's ultimately a problem of incentives: publishers and ad agencies are strongly incentivized to maximize the percentage of a publisher's traffic that (by hook or by crook) can be classified as worthwhile to the advertiser, whether it really is or not.
Given that the ad agency supposedly represents the advertiser in the negotiation with the publisher over campaign pricing and value but is nonetheless incentivized to collude with the publisher because the agency gets paid based on the size of the ad spend, you can see how the advertiser is left with no one looking out for his/her interests and is at the mercy of an industry that pitches it "metric of the month" as a proxy for value.
You can also see why performance advertising platforms that offer advertisers tons of real-world data and transparency, i.e. FB and Google, are eating up the entire ad industry. It's just a better deal for advertisers than the publisher + agency model. At least, we thought it was, but now that FB has admitted to "accidentally" screwing advertisers with bad metrics, who knows...
Anyway, yeah, traffic will always matter. My only point is that in a world where revenue scales with any and all forms of traffic, no matter how worthless to the advertiser (and the user, in many cases), is a world where people will pursue traffic by any means.
Edit: Ultimately, though, if you're optimizing for "user got some sort of potentially actionable informational value from this" instead of "user got momentary emotional charge (i.e. the classic 'surprise and delight') and the advertiser got positive/useful vibes in the process", the answer has always been user-funded content. This was true in the era of print/TV/radio, and it's true today. Every ad-supported model will always optimize for emotional manipulation over any other form of utility to the consumer. For some types of content (i.e. fiction and pure entertainment) this works out great, and for others (i.e. news, reviews, investigative reporting) it will always eventually lead to disaster.
It was brilliant work, and exposed deep flaws in many customer's traffic data. But no one cared all that much and it became a buried tool in a much bigger startup idea.
I'm sure a truly great way of measuring a user session's real value to the advertiser has been invented and discarded hundreds of thousands of times over the past two decades -- invented because it seems needed, and discarded because it actually worked and holy crap most traffic is garbage.
Most advertisers do track and measure actual revenue to them. I think your mistake is assuming junk traffic and bounced eyeballs dont convert or contribute to sales...They do. If you have a decent product with decent margins, each sale can pay for thousands of useless eyeballs.
If it didnt work, advertisers would stop.
This decision has helped keep the ad quality high, keep more revenue in my pocket, and allows me to have a relationship with the advertiser that isn't solely based on clicks and data. Edit: And this is important because I can have a conversation with them to make sure the traffic they receive is the type of traffic they want and not just immediate bounces.
Regarding headlines, my main content consists of office design image tours and my headlines are always basically "Company Name's Offices - Location'. I used to try and post catchy headlines, but after doing it for several years it lost its appeal and ultimately didn't match with the type of website I wanted to create.
First, what problem are you trying to solve? In this case, it's "How can I find good articles even with bad headlines?" So while the approach addresses headlines, the interest is in the content. So I'm not sure the proposed solution solves the perceived problem.
Second, what are the current solutions/workarounds to the problem? In my case, at least, the solution is blanket rejection of certain sites. I assume certain sites are so full of clickbait nonsense and/or partisan propaganda that I won't read them at all. The probably works better than some software that will consistently rate The Economist as good and anything from Infowars as nonsense (or worse, think the nonsense headline and the nonsense content are sympatico, so it's fine).
Third, what is the root of the problem? And the root is largely that people like their nonsense. People consistently read bad headlines and bad stories, often preferring them over respectable mainstream news.
And finally, how do you implement this? You clearly don't want something that can be gamed by crowdsourced campaigns, or it will be gamed. So you're either somehow relying on deep learning automation, or you're relying on human editorial effort. The former is unreliable, the latter is expensive, and itself prone to both bias and rejection (consider how many people consider Snopes to be untrustworthy).
I dunno. Maybe there's a great business or social idea here. But it's going to take some deeper thinking.
If this is a real problem for USV, they can financially back a team to do this.
While I know that some subsequent stories can do original reporting, too often sites with better SEO just republish stories without adding much and, whether intentional or not, often distorting some part of the actual story
If you search for a news story, below some results you'll see a "Related stories" box.
Some of those stories have labels like:
* In-Depth: a longer article about the story
* Local Source: an article from a source local to the story
* Highly Cited: the article that appears to be most frequently cited by other articles
* Most Referenced: web content that appears to be linked to from other articles the most frequently
* Preferred source: an article from a source you've marked as a favorite
I also have seen hundreds of stories written about me,
USV, and our portfolio companies that have sensational
and often inaccurate headlines followed by stories that
are essentially correct and well reported. It drives me
nuts but I don’t often do much about it.
Subjectively, this is not what I see. Instead I find that junky headlines go with junky articles. That would still be an interesting thing to try to objectively quantify, but different from what the author has observed.Reuters and other publishers have rss feeds however they are split accross many categories and also have strict ToU. I have been trying to find news & event feeds that are free to consume; ideally with a headline and article, but simply a blast like "3 trapped in hiking incident in montana cavern" would be useful.
Any resources or experiences would be helpful.
It would increase page load time of course.
Actually maybe this is a job for a browser plugin