Things will get even worse as scammy companies start flooding the web with LLM generated content pushing their products to bias LLMs to increase the probability of outputing their name for keywords related to their business.
Things will get even worse as scammy companies start flooding the web with LLM generated content pushing their products to bias LLMs to increase the probability of outputing their name for keywords related to their business.
It's not a coincidence that the solution to this problem is exactly the organizations that are being systematically undermined and dismantled.
Libraries themselves are a special thing - they've been invented before intellectual property was a thing, and have been attacked by IP proponents ever since. In this way, they're very much like LLMs, and many of the arguments against LLMs trained on copyrighted material apply just as much to public libraries. Oh the irony.
I've occasionally thought to myself "once an idea is created it wants to be free," but in the back of my mind I've always had some sympathy for artist of any medium trying to profit off their talents or anyone trying to create anything and profit from their efforts.
The library comparison is notable, but a library has a limited, physical nature. They can only hold a finite amount of books, while anything digital can be replicated indefinitely.
Whenever I read or hear anything from the medias now, I'm now always asking myself "what are their political inclinations? who is owning them ? what do they want me to believe? how much of a blind spot do they got ? how lazy or ignorant they are in that context ? etc."
They killed the trust I had in them so many times I can't get any the benefit of the doubt anymore.
It's exhausting.
Except given the noise/signal ratio and the sheer mass of information we have today, the workload is much higher than training for a 42 km run.
The economics of just giving the news with little bias just aren't there anymore.
What's happened is that the income of media outlets has declined to the point that most can't get factual accuracy even if they want it.
Instead of facts being unaffordable, it seems that lies and bias simply pay more (or at least the media outlets seem to think so).
Or to put it another way, the media's accuracy rate has stayed consistent at some value less than 100%, but if all three TV channels reported the same information then it looked like they had 100% accuracy. Once there were more sources of information then it became apparent that the media's accuracy was less than 100% despite their protests to the contrary.
The result is that the media landscape is fractured. A person can live in a bubble where all of their news sources (eg NYT, WaPo, and Bluesky for one bubble; Fox, Newsmax, and Truth Social for another bubble) all report the same information, making their accuracy appear to be 100%, while any single source of information outside the bubble that disagrees with the bubble is disagreeing with a bunch of apparently 100% accurate sources and so can safely be discarded.
The solution is to realize that no source is 100% accurate or unbiased even despite genuine efforts to be. That isn't to say that some sources aren't more accurate or unbiased than others, but you should apply some base level of skepticism to any and every source
News is leaning more and more into entertainment.
You did have all of this before, but 24h news channel with empty content are reaching new magnitude, fox news types of outlet are getting bolder and bolder, manufacturing facts is now automated and mass-produced, consequences for scandals are at an all time low, concentration of power at an all time high, etc.
It was bad.
It is getting worse.
There's a simplified page for CNN news at <https://lite.cnn.com>.
I've found that frustrating as all the stories are jumbled together with little rhyme or reason (though they seem to be roughly date-ordered).
Ironically, the story URLs themselves include both date and news-section coding, as with:
https://lite.cnn.com/2024/12/28/us/patrick-thomas-egan-accused-tv-reporter-attack/index.html
That's a US story dated 2024-12-28.It's possible to extract these and write a restructured page grouped by subject, which I've recently done. One work product is an archive of downloaded front-page views, which I've collected over about the past 5 days. Extracting unique news URLs from that and counting by classification we get a sense of what CNN considers "news":
Stories: 486
Sections: 27
76 (15.64%) US News
67 (13.79%) US Politics
9 (1.85%) World
8 (1.65%) World -- Americas
6 (1.23%) World -- Africa
15 (3.09%) World -- Asia
4 (0.82%) World -- Australia
5 (1.03%) World -- China
2 (0.41%) World -- India
37 (7.61%) World -- Europe
21 (4.32%) World -- MidEast
2 (0.41%) World -- UK
8 (1.65%) Economy
45 (9.26%) Business
4 (0.82%) Tech
3 (0.62%) Investing
8 (1.65%) Media
8 (1.65%) Science
7 (1.44%) Weather
4 (0.82%) Climate
22 (4.53%) Health
2 (0.41%) Food
1 (0.21%) Homes
39 (8.02%) Entertainment
52 (10.70%) Sport
22 (4.53%) Travel
9 (1.85%) Style
The ordering here is how I display sections within the rendered page, by my own assigned significance.One element which had inspired this was that so much of CNN's "news" seemed entertainment-related. That's not just "Entertainment", but also much of Health, Food, Homes, Sport, Travel, and Style, which are collectively 147 of 486 stories, or about 1/3 of the total.
Further, much if not most of the "US-News" category is ... relatively mundane crime coverage. It's attention-grabbing, but not particularly significant. Stories in other sections (politics, business, investing, media) can also be markedly trivial.
Ballparking half of US news as non-trivial crime, at best about 60% of the headlines are what I'd consider to be actual journalistic news, and probably less than that.
On the one hand, I now have a tool which gives me a far more organised view of CNN headlines. On the other ... the actual content isn't especially significant.
I'm looking at similar tools for other news sites, though I'm limited to those which will serve JS-free content. Many sites have exceedingly complex page layouts, and some (e.g., the Financial Times don't encode date or section clearly in the story URLs themselves, e.g.:
https://www.ft.com/content/d85f3f2d-9e9d-4d92-a851-64480e56a248
That's a presently current story "Putin apologises to Azerbaijan for Kazakhstan air crash", classified as "Aviation accidents and safety".-------------------------------
Notes:
1. For those interested, most readily accessed and parsed, the Vanderbilt TV News Archive (<https://tvnews.vanderbilt.edu/>), which has rundowns of US natinoal news beginning 5 August 1968, to present (ABC, CBS, and NBC from inception, with CNN since 1995 and Fox News since 2004). It's not the most rigorous archive, but it's one that could probably be analysed more reasonably than others.
We just saw this with ABC News’s settlement with Trump because its owner Disney wanted to stay in his good graces.
We also saw this with Bezos owner Washington Post
It's not easy for a truly creative, new and unique content to get into your local library.
And journalism has been gutted, more gutted than is obvious. Especially, with mainstream journalists having few "feet on the ground" a lot can sneak by (what happened in East Palestine, for example, can be found on Youtube's Status Coup new but not the mainstream).
The real thing that ruined the open web and viability of search was, ironically, when Google killed display advertising by cutting Adsense payouts to near zero.
Now publishers monetize via the much more sinister “affiliate” marketing. You know, when you search for “Best [X]” and get assaulted with 1,000 listicles packed with affiliate links for junk the author has never even seen in person.
At least in the old system, you knew that an ad was an ad! Now the content itself is corrupted to the core.
Google is machine-gunning its foot since 2021, it’s really unclear to me whether they’re killing their baby just to make the job harder for competitors or something. For now… I open the Google Search results with a machete, and often don’t find any answer.
Talk about severing your own foot to avoid gangrene.
I just don't think Google cares enough about the web as a whole to make strategic decisions for content quality in aggregate.
Sure it cares about geeky nuances and standards (e.g. page structure / load times), but Pichai isn't considering the impact on web content quality when debating an algorithm change or feature.
If Google continues driving web quality off the cliff? Well, the business KPIs stayed green.
The only thing they care about is ad revenue. Google created Chrome which vastly improved browser user experience. Google is a major participant in web standard & JavaScript language evolution, among other work. That's all true, but not necessarily because they "care about the web", but rather it helps their ad business. If people put the entire world's information on websites, and people spend more time in browsers, Google ends up earning more money from ads.
I disagree. Any prescription for what the ranking should be that isn't simply the most relevant result is a worse ranking.
I don't care if the top search result is the fastest, leanest, shortest, straightest, most adless, most equitable answer to my query if it's not the best answer to my query. I'll take the slowest loading, most verbose, popup ridden, mobile-unfriendly site if it's the one that has what I asked for.
Trying to add weights for things other than relevance is probably exactly where Google started going wrong. And then when it turned out badly, people propose yet more weights beyond relevance to fix the problem of irrelevance?
Even better for Google the worse the organic results are the more you need to rely on ads or some sort of ai snippet.