G/O media will make more AI-generated stories despite critics
vox.com
vox.com
They write page after page of content that barely get read. It all just exists to make the google bot happy. And that works. They bring in more customers & revenue apparently. (honestly, I wouldn't be surprised if it's all wasted and the added revenue is due to the better website and backoffice, but what do i know.)
So it seems only natural to have AI generate that content for the Google bot and stop wasting all this effort. Obviously I can't prove nobody reads this content, but seriously.. A landing page about hot-tubs and how cool it is to have one in your holiday home? Doubt that anyone actually reads that..
EDIT: I guess my point is also: it's already flooded with crap, just look at al the blog spam for things like 'how to start nginx'
1. Good old days, humans talk to humans.
2. Someone invents a bot, humans talk to bots.
3. Someone invents another bot, bots talk to bots.
This is the story of SEO, banner ads, stock trading, consumer pricing.There are probably many other cases.
The tech help spam: initially an attempt for a bot to talk to a human, but with LLMs crawling search results and aggregating answers, your answer to "how to start nginx" becomes the output of bot-bot interaction: an average of several generated answers (some initially and occasionally by humans).
If someone could provide a way to programmatically detect accurate, high quality content, it would solve the SEO spam problem. It would also as a side-effect solve the AI-content problem. Or maybe it's the other way around and it would solve the AI-content problem and as a side effect the SEO spam problem. Well, they'll soon be the exact same thing, and either way I'm not sure it's solvable.
Their lies are simple: they rule us.
(-;
But real barbarism begins when no one can any longer judge or know that what one does is barbaric.” — Ryszard Kapuściński
Human society is not fractal; we die off, new social value stores come about; optimize for self could be defined as “not build environment destroying technology.”
A synthetic machine kept running but never allowed new inputs will just exist in long term belief in its old bullshit
If you are aware then it's a "Lie". If you are not aware then it's a "Mistake".
LLM's are not aware, so saying "lie" is not correct.
I just asked (another anthropomorphic verb) ChatGPT for a better word and it came up with "fabricate".
> The term "fabricate" could be used to describe AI's tendency to generate fictitious information. This term has less anthropomorphic connotations than "hallucinate" and directly refers to the creation or construction of something, which could be a suitable representation for what an AI system does when it generates false data. Other alternatives could be "synthesize" or "generate," depending on the specific context.
I think the best way to refer to these things is the more accurate "error". The LLM isn't lying, it's in error.
There's one side of the current zeitgeist that over-anthropomorphizes these models. But there's another side that seems to be terrified that LLMs could be anything resembling intelligent
Both sides are pretty emotion over facts though. A lot of their "human-like" behavior is emergent from things that don't work the same way in a humans... but we also don't understand cognition enough to then say there's no overlap between what LLMs are doing and what some part of our own thought process is like. If anything we have more studies that imply the opposite going back decades: https://www.sciencedirect.com/science/article/abs/pii/S09266...
Correct. Which, as near as I can tell, isn't the correct characterization of what's happening with LLMs.
> But there's another side that seems to be terrified that LLMs could be anything resembling intelligent
I certainly have no such fear. I say that only to indicate that my biases are not of that sort.
Is it an unfounded or mistaken cluster of tokens which resemble an impression or notion then?
At the end of the day it's producing an output which systemically was intended to be accurate, but ended up not being accurate through a mistake in generation: That's a hallucination.
It is completely reasonable to say that "This machine fabricates widgets." or "This algorithm fabricates pseudo-random song lyrics." without being worried that the machine or the algorithm is exercising "cognition".
> Confabulation is distinguished from lying as there is no intent to deceive and the person is unaware the information is false. Although individuals can present blatantly false information, confabulation can also seem to be coherent, internally consistent, and relatively normal.
> Can include autobiographical and non-personal information, such as historical facts, fairy-tales, or other aspects of semantic memory.
> The account can be fantastic or coherent.
> Both the premise and the details of the account can be false.
> The account is usually drawn from the patient's memory of actual experiences, including past and current thoughts.
> The patient is unaware of the accounts' distortions or inappropriateness, and is not concerned when errors are pointed out.
> There is no hidden motivation behind the account.
It's a distinction worth making as people are much better at spotting lies than they are at spotting bullshit. The bullshitter may even be accidentally correct. AI models in particular will make up things in random places where nobody is sceptical or has their guard up because they see no reason for a lie, but bullshitters don't need one.
AIs do not have intentions, the only thing in an AI is a statistical model of how likely words appear in certain contexts.
"Hallucinating" is a much better term, even though it still pretends a staristical model is akin to a human.
I do agree with you that "serverless" is a misnomer and not a useful technical term. I'm guessing the marketing team wasn't fond of "dynamically provisioned and auto-scaled pool of virtual machine instance-based services".
In a technical forum, most people have the domain specific knowledge and the “well actually” is not actually needed.
[1] This isn't sarcasm. I think that's the world we're headed to. Hundreds of thousands of little creators serving their own unique niches. A world without Disney.
What they want is to optimise the ranking Google gives them. If Google's ranking filters out falsehoods then lies can be optimised away automatically.
I keep seeing these "AI will destroy X" articles, I've never seen one that wasn't ultimately referring to a minor incremental extension of something already prevalent.
Keywords gave way to link farms gave way to shitty outsourced "content" and now LLMs will be part of SEO. It's not really different from getting a content farm to write you a bunch of crap, just cheaper and faster. Search companies will adapt, and seo scammers will find some new thing. It's all such a minor part of the internet anyway.
Not sure they have adapted. Most of the links on the front page of a search are crap.
That’s by design. They load the results with crap so you click on the ads instead of the result. Ads generate revenue for Alphabet, search _results_ less so.
Simple as.
Using a gillnet to catch 1,000 fish in an hour is not really different from using a rod and reel to catch a few a day. It's just cheaper and faster.
But the gillnetting can easily lead to extinction and death of an ecosystem whereas recreational fishing rarely does.
Scale matters.
So the scale already exists, and we're really just talking about quality.
Being able to generate 100x the spam will certainly hurt things even more.
Very well said.
As for why it might not make Google worse, just think about YouTube algorithms:
1. One good video is substantially better than hundreds of bad videos, or thousands even.
2. Clickbait works in terms of getting people to click on your video, but if you don't deliver what the clickbait promised people would leave and it would hurt your stats.
3. Generating large amounts of videos may be profitable because it's like collecting peanuts with automation, but that would never reach the mainstream audience like Mr. Beast.
This is because the algorithm promotes videos that are frequently clicked and fully watched. Generated or not, the content must be good enough to be watched through, in order to please the algorithm.
You see, there are a lot of bootleg versions of Stackoverflow, but many of them are terrible. Some of them get ranked because they're still helpful (to the 80% majority of programmers, probably not for the HN elites). We can imagine there would certainly LLM generated versions of those, but very likely with better quality. Generated or not, it's still an arms race to generate more helpful and higher quality content to get traffic.
Also, HN people have been complaining about how the quality of Google degrade over time. It might be Google's fault, but I believe there's another important factor that hurts search engines as a whole - There are too many walled gardens nowadays: Instagram, Discord, and now Twitter, etc. Value information is hiding behind those, which would certainly make search engines less helpful over time.
For example, when I play not very popular video games I wish there are wikis or guides for that. In the old days, I know I can find most of the information in several forums and check the top posts. Now I need to go through Discord and scroll through casual chats and wonder if I missed something.
With LLMs, people would likely be draining that information from walled gardens and compiling them into helpful insights.
Again, there are a lot of dynamics, and it's hard to tell how things would turn out. Maybe it would kill Google and possibly the old Web, maybe it would make the Web great again.
I was searching a municipal website that had a bunch of pdfs. Instead of downloading them I'll I figured I'd sure the "site:__" to look through them. It didn't work. It seemed to index only 3 of those files. sigh.
One of the appeals of using "chat bot" is you get an answer without all the ad crap the web delivers. I think thats the appeal of stackoverflow and Redit.. Using AI to make the web worse is something...
This is an interesting take. Google can only "reward" videos that people watch for longer because they own the site. That is, they can tell when people leave. For a good chunk of sites, they still can because of Google's ad market share, but for a really sizeable chunk of sites that isn't the case. For like 80% of users, they also still can because they own Chrome.
I'm not sure I like the idea of allowing Google to reward sites that use their ads or browser monopoly in their search algorithm.
Plus, I've never once seen enshittification reward better quality. I highly doubt this will be the first time.
however, with all of the debates around attribution and ownership for human creators in the age of AI art, combined with the apparently legally and ethically dubious means by which these megacorps obtain their training data, i have began to think about what are some avenues that the general public could try to protest the actions of these megacorps by discretely poisoning the well of their training data.
with the gigabytes and gigabytes of data i have generated of mostly incoherent audio, i have considered releasing this music for the first time by innocuously labeling it as a music audio dataset with the intention of trying to make it appear extremely attractive to megacorps scouring the internet for free data. my individual contribution probably couldn't amount to much, but if a concentrated mass of people did this in their respective fields, perhaps this could be a way of at least obstructing these corporations from freely capitalizing on the hard work of real artists.
But I agree with the sentiment—if the blatant disregard for IP is not curbed, this sort of thing would have to be done…
I'm definitely pro "personal data poisoning," e.g, absent legislation and/or incentives with teeth - we should all be working on ways to confound and confuse and generally screw up the companies who are sucking up all the personal info.
This kind of runs parallel to that. Google search has been pretty bad for some time now, might be worth "burning this thing down to save it."
Google makes money off of SEO spam. They get ad revenue from sellers and then give some pocket change to the scammers from the sellers. All of the money goes through their hands first.
They're evil, they've sold whatever soul they might have had to the god of "money at any cost", and the current shitshow of their terrible search results is what you get when a soulless wreck is left to shamble the earth until its heart stops beating.
Quality, human curated, perhaps niche, search is what we should be encouraging and paying for with actual money, and not by renting our brains out.
It's hard to argue against the basic idea of human curation as a necessary component. I'm envisioning something like a search engine on top of a community-curated, categorized list of sites. I'd prefer something that works like Wikipedia, rather than a service controlled by a private company.
Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html
So AI generated content will now take up the single non-sponsored link space to replace the ad-stuffed SEO-gamed clickbait article?
Huh.
This strategy of avoiding backlinks and technical BS & focus 100% on content quality has worked really well for me.
If you think about it, Google has all of the data to do this:
- Android
- Chrome
- Google Analytics
All of the big platforms use UX metrics to influence reach - LinkedIn, TikTok, Instagram, Twitter, even YouTube uses UX metrics to influence reach, why not Google search?
It seems like people using AI see some limited amount of success, for some limited amount of time before losing their rankings.
If I was Google, I would let people think their AI content is working for a little bit, then de-index their website to demotivate + scare webmasters away from using AI.
First things first, using AI to generate CSS advice is probably the most false positive laden use-case I can imagine. Especially because mistakes in CSS aren’t always readily apparent.
But more importantly, there’s really nothing stopping folks from omitting the fact that the content on the page was AI generated. It’s the new hotness right now, so I can understand why they say so, but I can’t help but think they’ll start hiding this fact.
It's always worth verifying the answers, but that's also true for google results, and with GPT I don't have to dig through mountains of trash advertising sites to get to something useful.
Really, I'd just like to see Google as useful as it once was.
So what if Google has a few less users? Since Google Ads is auction-based, it means a few less impressions will be available, and the cost of impressions will go up. Advertisers will pay a bit more for less but Google's sort of the only game in town so it won't make a huge difference.
Now if Bing achieves 30% market share on the back of its great GPT integration or something, maybe that gets Google off its ass. I think LLMs might be the beginning of the end for Google but they're so entrenched it'll take 20 years.
For example, consider the example of Demand Media (e.g. https://variety.com/2013/biz/news/epic-fail-the-rise-and-fal...).
They invested quite a lot in cheap content generation and for a while it served them quite well. I struggle to imagine a scenario where lowering the cost of polluting would have made this strategy more viable.
Most really good material is already better found in books (even if ebooks) from real publishers than by trawling the Web, anyway. Even before the LLM apocalypse, the Web has always been rather disappointing at the whole "all the world's knowledge" thing. Great for trivia, great for a few aspects or presentations of a few topics, kinda shit for the rest. AI garbage wrecking the Web doesn't seem like a complete change of the current state of things, but a shift in where exactly the "Web is good for this / Web is not good for this" line falls.
The big problem is false positives—human-written text that gets incorrectly flagged as being written by AI. There's going to be so many false positives in any scheme like this, it'll barely be worth it.
Expect lots of false positives and honest businesses being affected. How much should the filter threshold be tuned to smash more bots instead of real people?
You also have to remember the future of AI is not just LLMs getting larger and larger. There will be something after them, and something after that. I'd guess we're maybe 2-4 years away from an AI that hooks up to an LLM but has "actual" knowledge of things so it doesn't confabulate new facts, which would remove one of the major signals I'm currently looking for in GPT content.
Getting the actual facts to put in the story normally took far longer when I was working in a newsroom.
Longer, harder to write stories, like deeply researched news, or long features are not something that AI can do (yet).
Dangerous to their business model, on multiple fronts.
I just don't have any interest in reading a story written by a LLM.
Due to the incentives of the players involved, there is no solution that would lead to a neater, less blogspam-infested internet. The concept of a 'website' as Tim Berners Lee and the early netizens understood it, is commpletely outdated.
The second we were able to add an Adsense script or Paypal button or Amazon affiliate link, the fundamental motivation for building websites changed. The artisanal personal websites of yore aren't coming back. People who used to build those have long since moved into less crowded, less commercialized territories where their work might actually be seen.
If you don't want to get angry when looking up a recipe online, buy a recipe book. Everything good on the internet is already paywalled anyway.
If you don't know how to find what you're looking for, just ask. Paywalls aren't a real thing that have to be obeyed, nor is there ever only one spigot for any digital fountain.
> "In early July, managers at G/O media, the digital publisher that owns sites like Gizmodo, the Onion, and Jezebel, published four stories that had been almost entirely generated by AI engines."
More about them on Wikipedia: https://en.wikipedia.org/wiki/G/O_Media
I won't, cause I have largely stopped using Google.
And first they'd have to figure out what that even means, and the impact on non-professional writers. Blogspam will get worse before it gets better.