Investigating how the New York Times A/B tests their headlines
blog.tjcx.me
blog.tjcx.me
But the quote included suggests the NYT does do the latter: "Half of readers will see one headline, and the other half will see an alternative headline, for about half an hour."
So given that the author mostly observes long consistent blocks of time with the same headline, that suggests the NYT is allocating them to a subgroup in a persistent way (by IP or whatever). Then perhaps the cases where they didn't observe A/B testing were just cases where they were randomized into the optimal (hence final) headline subgroup by accident at the beginning, and never saw any different.
Edit: The only example of clickbait here is OP's title.
Assignments are also almost certainly sticky based on a browser cookie.
I’m sure you would agree that would be a good representation.
There's a peak 30-min in the morning just before work starts for most people, for example.
Imagine you're running an AB test on a website, where the number of customers you get increases linearly over time. If you test two features sequentially, there is no duration you could pick that is small enough to where you would not expect the second option to win.
It would have been more interesting to see the title flipping back and forth, because that would reveal how long those tests last. Less neat though.
1. My scraper runs every 5 minutes, with a randomly-generated user-agent and never sends cookie headers
2. The charts are bucketed by half-hour periods, so even if the headline flips back and forth many times in half an hour, the colors are grouped together
3. Agreed that in the SpaceX situation (and maybe the Cuomo situation) the headlines change because the stories change. But, e.g., in the Meghan Markle situation, the first headline appears _after_ the interview is over. But that's something to watch out for!
And charts like this one[0] look (to me) like a clear example of A/B testing. But would be interested to hear other explanations!
[0] https://nyt.tjcx.me/articles/1163e0c4-e609-5cfa-aff1-b0945f1...
They may also conduct certain A/B tests only on cookied users, in which case your cookieless bot would not have seen everything.
I might also argue that online news drives this more so than print - with print you tend to buy the paper based on the totality of their reporting and reputation versus a zero-cost click of a headline on your screen.
Article about a bank robbery:
"First National Bank Robbed"
"Gunmen Rob First National Bank in Daring Robbery"
"Coup Klux Klan: Don triggers mob & rob bid"
[1] https://thepostmillennial.com/nyt-makes-terrorism-story-into...
They're very rarely "accurate" the way stories are intended to be for these reasons.
It's chosen for Wednesday in the budget meeting. It's 1A material. A cursory headline is written. Then a big news event happens: now it's 2A material. That means less space. The copy needs to be cut, the headline needs to be one line shorter, etc. This happens at, say, 9 pm and the journalist and editor in chief have been at home for hours.
My point, rather, was that headlines slip past the journalistic barriers applied to copy. Often the people on the desk are making a headline fit at 11:45 pm or later and have no real oversight.
A journal is a business, with something to sell, the news, and the attention of their readers to advertisers as well.
> The New York Times is a big deal. As they tell their advertisers, the NYT is the #1 news source for young, rich thought leaders:
Obviously ultimately it reflects badly on the profession, since all these news sites are using the same clickbait techniques, from Breibart to the Dailymail to NYT, since they probably hire the same consultancies when it comes to clickbaiting design, or at the very least, people who come from the same marketing circles/education.
Also "young, rich thought leaders", right?
C-Span is a good example.
PBS, NPR, and, via the Corporation for Public Broadcasting[2], lots of individual, independent local radio and television stations.
[1]: https://en.wikipedia.org/wiki/C-SPAN [2]: https://en.wikipedia.org/wiki/Corporation_for_Public_Broadca...
The article defending looting was absolutely a breaking point.
I know the CBC in Canada is absolute crap - 80% of the reporting is sensationalistic and for the past 4 years has been focused on the nuances of Trump - not exactly relevant to Canada. At least not deserving of more coverage than domestic issues.
https://www.icaew.com/insights/viewpoints-on-the-news/2021/j...
I often find that New York Times and other media put important details that sometimes contradict the headline or challenge the main idea of the article somewhere close to the end of the article.
This lets them still claim objectivity while preserving a highly polarizing message.
Burying the lead takes advance if this and puts the important but inconvenient information at the bottom.
That article links to this[1], by Howard Owens:
"It was then I realized, there is no historic basis for the spelling of a lead as “lede.” “Lede” is an invention of linotype romanticists, not something used in newsrooms of the linotype era."
He says he looked at books from the 1940s - 1980s, "The fact is, in none of the dozens of old journalism books that I have examined — none of them — spell it “lede.” I can’t find the definitive first reference to “lede” but it doesn’t start appearing in journalism books until the 1980s. .. It wasn’t until linotype was disappearing from newsrooms across the nation (late 1970s and into the 1980s), that we start seeing the spelling “lede.”
The safest conclusion, then, is that “lede” is a romantic fiction invented by those who were nostalgic for the passing of the linotype era."
[0] https://www.merriam-webster.com/words-at-play/bury-the-lede-...
Sensationalist junk.
The other problem I have, and why I stopped subscribing to them was their app was terrible and mostly only promoted opinion pieces to me. In fact, I'd love the OP of the story to study how many of the top 10 articles are opinion pieces at any piece of time.
[1]: https://www.theinformation.com/articles/meredith-levien-want...
With quite a few of the posts I'd just present them side by side, with the differences highlighted.
The integrity issues come up when even though the facts reported are real, the conflict that frames them is manufactured - and this is why people reject news. They don't reject it because of fake facts, they reject it because of fake conflict. When I want the real news, I go to fringe websites, because they get the real conflict right, and if I need details and facts, I can look those up. The reason they get the real conflict right is because by definition the fringe lives in that conflict, and they are the real anti-establishment that creates a counter balance narrative, whereas a conflict produced by setting mundane events and facts against the backdrop of an ideology designed to manufacture conflict is unreadable tripe.
If I needed to sustain the dissonance of establishment narratives, I would read the NYTimes to keep up the appearances, but since my livelihood and aspirations do not depend on that, I have the freedom not to engage it. If you think you are being played and manipulated, watching these A/B tests should be enlightening.
For multiple counter-examples, see pretty much everything on https://www.msn.com/en-us/news/good-news
Arguably, newspapers as a concern are a post-Marx phenomenon, and it's what people get taught in j-school, and what editors accept. Students of both are peddlers of pernicious nonsense, but their students also produced some entertaining and powerful things, notably recycled recipes for accelerating and reverting democracies into their inevitable tyrannies.
Anyway, the NYTimes doesn't register as meaningful to me, and if they were looking for a reason why people are leaving them and other mainstream outlets behind, it's becuase their conflicts are contrived.
Read local news.
In fact that gives me an idea, a local news aggregator which ignores redundant news and can find the most impactful ones.
If the narrative is entirely that if we dont actively consider capturing interest, we’d be doddering and hard to track, if we do we’re abusive, then media is forever doomed to be unsatisfactory. We all hope to improve.
In the world of headlines, the “spiciness” that’s been advanced as a function of engagement hunting is something that’s currently contended with through human intervention. All headlines are human created and the outcomes of AB tests are more about improving the understanding between author/editor and captured audience than manipulation or future interest conditioning.
[disclosure: i lead ML platforms and the algo related eng products at NYT. all thoughts are my own presentation of what i have experienced, and not company opinions]
It would be nice if we had a news source with a feature (opt in) where the reader got an email quiz the next day, asking just a few questions about the facts in the article. The feedback loop would tell us which articles actually make the reader better informed.
If this level of desperation is actually seen as acceptable for a media house, then journalism or whatever is left of it, is in dire need of help.
A number of times I have seen on some place like Facebook where the initial article has some extreme headline, and then hours after when engagement is up, the headline is swapped to one that is less inflammatory. A few more extreme-perspective friends will send me an article saying "see?!?!" - and by the time I see it it's already been rewritten.
A few times I would click on an article going viral and find the headline in the article itself doesn't match the one cached by Facebook. Remember that most people aren't even reading the headlines and just assume that NYT are trustworthy. The first headline is the one they end up internalizing.
Bare in mind, there is zero consequence for doing this either. The newspaper and claim they were "correcting an editorial error", whilst openly spreading misinformation about some hot topic.
I think at the very least it should be mandatory to maintain a list of edits to an article once published - and to indicate to the reader clearly that the article has received a number of edits.
It may not be "fake news" in the Fox srtain, but the OP's point is that it is not for the good of the reader or society.
I think, when you're in the day to day of a web platform, you really just simplified everything from your hiring to your board meeting to one metric: click.
Everything is number of click, it becomes no more evil than a Lion in the Savanna, you just must increase your number of clicks somehow and you forgot why.
It can lead to interesting "let's remove the legal text away and see what gives", "let's color ever buttons blue vs pink and see who wins", let's reorder the search result independently of what a normal human would expect (as in... cheapest first?)".
I see what you mean, but it's not evil anymore. It's now part of nature, they must hunt your click. They can't stop and they can't explain why. They could put one headline "Myanmar military coup" then replace it with "Myanmar cabinet reshuffle" to see which one you click on with little input from any writer and no philosophical debate whatsoever.
It's definitely evil in the parent commenter's and the article's cases, but if it's testing color schemes for your blog or a signup flow for your productivity app, there's a pretty good chance it's fine (depending on the underlying ethics of the blog/app).
But I overwhelmingly agree that ubiquitous or bureaucratized doesn't mean "not/no longer evil". If anything, there's possibly a light positive correlation. That comment's espousing one of the worst and most cynical philosophies I've ever encountered.
There might even be some natural/game theoretic phenomena at play here, too. Information wants to be free, and evil wants to be normalized.
Here's one: https://nypost.com/2020/06/02/new-york-times-changes-headlin...
Before: “As Chaos Spreads, Trump Vows to ‘End It Now"
Edited: "Trump Threatens to Send Troops into States"
Another one:
https://thehill.com/homenews/media/489013-trump-rips-ny-time...
Before: "Democrats Block Action on $1.8 Trillion Stimulus"
After: "Democrats Block Action on Stimulus Plans, Seeking Worker Protections"
And changed yet again: "Partisan Divide Threatens Deal on Rescue Bill"
https://www.techdirt.com/articles/20200413/11142144293/lessi...
Before: "A Harvard Professor Doubles Down: If You Take Epstein's Money, Do It in Secret"
After: "What Are the Ethics of Taking Tainted Funds?"
To be fair, its not only the NYTimes that does this. One that was circulating around Twitter yesterday was: "How I got COVID after taking the vaccine" with a footnote in the story: "Author did not get COVID". The headline was changed to "Why could you still get COVID-19 after receiving a vaccine?"
https://github.com/ecprice/newsdiffs
This alternative looks more recent
An interesting reactionary adaptation to this "temporary sensational title" exploit.
I think more needs to be done for publishing transparency.
I find it comical that any write-up critical of a left wing organization needs to include something like this to avoid being flagged, down-voted, whatever, on the platform its being posted on.
Whenever I'd load the website, sometimes all the headlines would suddenly switch to different wordings, once sufficiently loaded.
As frequent Hacker News submitters know, headlines alone can determine whether a post gets upvoted.
Selected excerpts:
"I actually watched this interview—all two hours of it—and I can tell you that the first two headlines are a much better summary of what went down. Yes, Meghan does reveal that she contemplated suicide, but it’s a five-minute interlude in an interview that has a lot more going on."
"Trump starts off addressing conservatives and claiming leadership of G.O.P. but in the final headline Trump has a hit list and is firing a warning shot. And sure enough, the bombastic rhetoric propels this article onto the “most viewed” list."
In other words, while you're correct that titles are important for clarity, the claim in the article is that "clarity" can be at odds with what gets clicks (or gets upvoted).
I'm able to viscerally internalize this because, as a tech worker, I've typically used A/B testing as a tool to maximize engagement, but I've never in my career used it as a tool to maximize correctness or "clarity".
I'm pretty sure engagement is more important than clarity. I don't see how A/B testing helps you optimize for clarity. What are they measuring to determine a clarity score?
Sadly, most of the time engagement is highest when something is more outrage inducing.
Clarity? Or click-thrus? It's definitely not the former. Which in turn discounts the NYT as a beacon for journalisms and associated standards. They're desperate for traffic. Clarity isn't a metric in the playbook.
Witin the world of high volume publishing, this type of A/B testing was popularized BuzzFeed.
It seems like MAB would be the ideal approach for something like headlines.
Service that scrapes websites known to A/B test things, caches the A/B test content, and presents all of the variants at once.
In this context, yes, I want all the headlines. Maybe they could be presented as a diff. With timestamps!
I saw one the other day from Nick Kristof: "America is not designed for the anatomically correct". It was about the need for public restrooms, but the headline didn't make me read it. It sounded goofy and, other than knowing Kristof writes decent articles, I passed it over.
Later, the headline was "America isn't designed for those who pee". A bit coarse, but very much more clear as to what the topic of the editorial was.