CNET Has Been Quietly Publishing AI-Written Articles for Months
gizmodo.com
gizmodo.com
Oh piffle. My wife and I roll our eyes and mock the sports reporters when they run down to the field and ask the question "How does it feel to have won the $CHAMPIONSHIP?" "Oh man, it's... it's just so incredible, you know, you work all your life to reach this height and... here I am... it's just... sniffle it's really amazing. I have a really great team and I'd like to say HI MOM and... just.. oh man..." "We'll let you go join the celebration. There you have it, $OTHER_ANNOUNCER."
(Our favorite is: "Hey, $COACH, what do you plan to do to win this game?" "Well, you know, that's a good question, and we've got a plan we're going to execute to, you know, go out there and play strong, we're gonna try to, you know, score points and, uh, prevent the other team from scoring points. We're pretty confident that if we can execute that plan we're going to win." In fairness, sometimes the answer is better than that. But a lot of the times, if you think about what they just said, that's what it amounts to without any particular loss. After all, even if they have a good answer, they're not terribly interested in giving it to you before the game....)
I mean, I get it, I'm not so cold hearted as to not understand why they ask the question. It just doesn't happen to speak to us. But I would rate it more on the easy-mode side of sports writing, way easier than, say, being accurate about statistics or being correct about what strategy a team needs to pursue against another team for victory or what strategy they're going to pursue.
Extremely, extremely stereotyped responses. Can't hardly be wrong, it's not like saying "excitedly" instead of "energetically" is going to be the difference between accurate and inaccurate. Easy mode for a transformer architecture.
A broadcast from the Super Bowl:
Reporter: "How does it feel to have won the Super Bowl?"
Star quarterback:
Completion:
"It feels amazing. We worked so hard all season and to finally have won the championship is truly a dream come true. I'm so proud of my teammates and coaches for their dedication and hard work. It's been a long journey and I'm so grateful for this victory."
(I was in a Zoom chat yesterday with the transcription on and I noticed it was doing the same thing. It was aware of the "ums" and "uhs" and sometimes those phonemes would be in the original transcription, resulting in the wrong words/phrases, but once the transcription stabilized they are edited out for the most part.)
You could replace the entire Daily Mail
Interviewer: “The most exciting moment during the race weekend?”
Kimi Raikkonen: “I think so it’s the race start, always.”
Interviewer: “The most boring?”
Kimi Raikkonen: “Now.”
Interviewer: "Kimi, you missed the presentation by Pele at the front of the grid."
Kimi Raikkonen: “I was having a shit.”
He wonders if, in order to be a top tier athlete, they almost have to think this way. Simple, repeatable mantras drilled into their conscious and subconscious. Too much thinking would throw them off.
[0]: https://bpb-us-e1.wpmucdn.com/sites.psu.edu/dist/7/59784/fil...
"So I was sat there waiting for the cross, and it comes over me head and I just hits it and bang it's in the back of the net"
I think ChatGPT could easily replace commentators too. Of a missed shot - "He'll want to have got that on target!" - Oh he will? Such insight!
- I’m just here so I won’t get fined.
- It’s only game, why you heff to be mad?
- They had us in the first half, not gonna lie.
"Well, after analyzing hundreds of hours of video of the opposing team, we've pretty much decoded their play-calling signals. We're also going to monitor them remotely and communicate with our players on the field to give ourselves the edge we need. We're also planning on deflating the ball a bit to make it easier when we have possession. Also we've stolen our opponents' play sheets, so we hope to make use of them as best we can. And we're hoping that our opponents' headsets might 'mysteriously' stop working.
And of course as always our players are going to bring their A game to the field to really win this!"
Crash: "You're gonna have to learn your clichés. You're gonna have to study them, you're gonna have to know them. They're your friends. Write this down: We gotta play it one day at a time. ...."
Nuke a.k.a. Meat: "That's pretty boring."
Crash: "Of course it's boring, that's the point. Write it down."
The thing is that webpages don’t ‘exist’ in the way most people mentally kind of think of them - as ‘documents on the file system of some computer’. And google is not an index over all those ‘documents’.
How search actually works is Google bots go out to a bunch of servers and ask them ‘hey, what kind of documents can you make?’, and the servers respond with dynamically generated lists of links to those ‘documents’, and google it requests them, and the server makes them and sends them to google, which indexes them so that later when someone searches the internet, google can send them to the same link, and the server can make the document again for them.
In an internet that works like this, using an LLM to generate documents that google indexes is a waste of everyone’s time.
This turns google into an index of the output of various LLM prompts that someone has run that they think someone might later search for.
This wastes everybody’s time.
An LLM outputted summary of every earnings report of every company doesn’t need to be pregenerated - we can make that later, at the time the customer asks the query ‘what was in the earnings report of Foocorp in Q4 2013?’
This breaks the ‘document-centric’ conceit of search; and that might not be a bad thing.
Just take a look at a random stock in apple stocks, the news section is just filled with this garbage.
AI contamination is already starting. There's no way that content written by an LLM hasn't already been used to train an LLM.
This issue is more about the value of search finding ‘documents’ when ‘documents’ are thin layers of low-value-add LLM cruft over raw factual data.
The facts are what you’re searching for; search indexing has always involved trying to figure out the underlying informational content of a document; LLM prettification is just adding obfuscation over that data that search indexers then have to try to remove.
There has to be a more efficient way for a website to say ‘we have all the earnings data for every US public company’ without them having to dress that up in LLM prettification that google’s language model can index.
in the early oughts, I said to anyone that would listen that the powerful popularity searches shown by google will be sort of self-referential given a simple definition of what can be indexed and what cannot..
Frankly that might be better then everything being wrapped up in prose.
It's usually pretty obvious because you'll read something and some form of uncanny valley starts to creep in where you're like "no sane person writes like this."
We'll always need someone to be on the scene to capture the events / investigate the stories, and I don't think the world is dystopian enough to start manufacturing AI news drones.
Efforts will be put in place to filter out the garbage AI content with human made, real world reporting.
When the content on the other end is written by a sapient human (or, eventually, an AGI), as much vigilance is not needed. Vigilance is always necessary, but the level required for parsing the output of language models is much higher.
This requirement is why publishers like CNET quietly mislead their audience and do not clearly mark each submission as AI generated. If it were not an abomination and an abuse of the reader to the gain of the publisher, then they would proudly claim to be doing it.
There is no interaction to be had in the consumption of most material online, and any perceived interaction is surface-level as best.
Were you reading that stuff before "AI"? Why?
The interesting thing about this moment in AI, it seems, it’s not how good the AIs are, but how bad the lower end of human work and ability is relative to the AI.
In this case, how much SEO-farm rubbish content has been put out by salary-slave-humans and is it really better than being tricked into reading AI content?
Two evils I know. But maybe there are deeper issues than the met existence of decent non-general-AIs …?
Which raises the interesting question of how self-referential and useless search results go, if they are primarily great big generative, predictive models feeding huge indexing models for the purpose of predictively generating the right additional content intended to snag an actual human neuron (that is, to select advertisements)? Once the whole thing is computers talking to computers hoping some gullible living person is watching, it becomes indistinguishable from bitcoin mining - a kind of viral, self-referential internet onanism.
In a way its ad-money mining via SEO/engineered content to get high page views/CPM. The issue with this comparison to crypto mining is that the surface area of low-quality content prevailing means their ad campaigns are likely going to be less effective because humans reading the low-quality content just aren't going to stay long enough and/or will avoid the content. I can imagine the ad companies shifting to other types of content like video, then the AI generated content doesn't bring ad-revenue.
Except now they're not on 3000 sites, they're on 3000 sites plus CNET.
https://www.minitool.com/backup-tips/crucial-bx500-vs-mx500....
Is this ChatGPT written or minimum wage copywriter written? Does it matter? It doesn't help me in the least and it's the first result for "crucial bx500 vs mx500".
Edit: seriously, doesn't it sound exactly like ChatGPT articles?
They are writing reviews with no experience? lol.
For those who aren't aware, feature checklists for web hosting services are 90% bullshit. Not that they aren't true, but the items that get checked off are often trivial nonsense which are implicitly available from any commodity hosting provider like "control panel", "access logs", or "password protected directories".
The real differentiators are usually things which don't show up in marketing materials, like "is the support team halfway competent" or "how badly does this provider try to nickel and dime you". Which is why you need to actually work with the provider to give them a meaningful review.
It's atrocious. And Google just pumps this garbage. I need to finish writing about all bullshit from big brands doing these reviews. CNET is just one easy example.
https://www.theatlantic.com/magazine/archive/2021/11/alden-g...
What better way to cut costs than to replace the journalists with AI?
Many pieces are now just regurgitating twitter posts anyway.
What is funny is that this is exactly that and to make things worse, this article from gizmodo is just regurgitating an article from https://futurism.com/the-byte/cnet-publishing-articles-by-ai that is regurgitating a twitter thread which is more informative because it includes information on how the person found it in the first place https://twitter.com/GaelBreton/status/1613110185995771905
Now we are on, hacker news, talking about a gizmodo article, based on a futurism article, based on a twitter thread.
In years past, I think it was much easier to discern such low-value content. But lately I find that it's not immediately obvious, and I waste time in drawing that conclusion and moving on.
So my search strategy, at least for the kinds of things most susceptible, is to shift from "open-ended with exceptions" to "only trust sites whose names I recognize".
Sad to see that I should consider moving CNet from the good column to the bad.
> In years past, I think it was much easier to discern such low-value content. But lately I find that it's not immediately obvious, and I waste time in drawing that conclusion and moving on.
I know what you mean, but I wouldn't describe that as an increase in "quality."
> So my search strategy, at least for the kinds of things most susceptible, is to shift from "open-ended with exceptions" to "only trust sites whose names I recognize".
IMHO, one interesting thing "AI" might do, is finally kill the open, free-to-access Web, returning things to something like the 90s, where if you wanted to know something, you had to buy a newspaper, magazine, or book published by an institution. It would do it by filling the web with low value stuff that's too hard to detect.
DALL-E/Stable Diffusion are already killing the appeal of certain styles of fantastic art, due to overexposure and mediocrity.
I look forward to the day when Wikipedia is overrun with "AI"-generated vandalism that's difficult to detect, but designed to corrupt it according to various agendas. They could even include generated citations to documents that sound like they could support the vandalistic claims (because few actually go through the trouble of verifying those).
"write a glowing review of product X, assuming you are an Y using it in your line of work as Z"
the output will look better than printf("Template blah %s with %s and %s is really good, check out our website %s, use coupon code %s");
but it's not necessarily better.
My goal is good information, but these technologies (especially currently) mainly enable bad/mediocre information at scale.
This is the thing that actually scares me, because I don't believe them when they say they have editors scrutinizing everything the AI writes. They may give it a glance, but I would bet any amount of money that they don't rigorously check everything. And as time goes on, they're going to spend fewer and fewer resources on editorial, I guarantee it. My basis for saying this is that editorial has already been deeply cut — forget rewrites, forget fact checking, it feels like half these articles have outright typos that never get caught. I'm not sure anyone but the author closely reads an article before it goes into publication, and if the author is an AI...
They aren't going to staff up on editorial if AI can generate content, they will just spend $0 on writers, and as little as they can get away with on editors, approaching $0.
See, for example, https://automatedinsights.com/customer-stories/associated-pr...
The fact that we just noticed appears to be driving the angst, not that it's suddenly here. Because it was here all along.
source: shared an office with them
This will be interesting to see over time how Google responds which it looks like they somewhat have regarding Bankrate: https://www.seroundtable.com/google-ai-content-guidelines-ba...
To date there are well over 100 of his videos that appear to be "inspired" by my content. For example:
Me, 2008: https://www.damninteresting.com/the-third-reichs-diabolical-... Simon, 2021: https://www.youtube.com/watch?v=8mP4RQ783lk
Me, 2005: https://www.damninteresting.com/the-man-who-was-a-dwarf-and-... Simon, 2019: https://www.youtube.com/watch?v=TpdMolMlDzw
The list goes on and on. I've been writing online long enough that I've been poached and plagiarized six ways from Sunday, but no one else has come close to leeching so long and systematically as this unoriginal, abject freeloader. Simon Whistler is downright seaward.
People are raising quite specific concerns about this technology and its potential impacts, and are doing so in a world that has been coping with the unintended consequences of new technologies since the industrial revolution. You may perfectly validly disagree with their concerns, but to ascribe them to an emotional need for humans to be "special" doesn't contribute to the debate. It's just incredibly condescending.
You could do worse than to avoid making arguments of the form "your points are so wrong you can only possibly be making them out of spite/jealousy/insecurity/other" in all debates on principle.
What currently exists is far from a human-level AGI, and it's not likely that just making a few tweaks would even get closer to that goal. The model doesn't in any sense understand what it is writing, it generates text that is optimized for appearing as if there is intelligence behind it, as long as you don't read too closely.
And while we don't have to fear a godlike superintelligence turning us into paperclips yet, there are a lot of potential negative consequences to this kind of limited AI, and very little of value to society. We don't need more efficient ways to generate bullshit or propaganda.
It's kind of like the deal with the More Plates More Dates YouTube Channel, and the Matt Does Fitness guy. Matt has an amazing physique, but because of that people thought he was on steroids. He paid MPMD to randomly drug test him over the course of six months like WADA (world anti-doping agency) to prove he was natural. He had to do this because steroids exist in the world.
Same thing with writing. Using perfect grammar and a more formal sentence structure will probably cause people to think you are a robot. But writing about your life, your hobby projects and what you are doing with them will show people that you are a real person
Also, I feel the rise of Chat GPT will result in more value for verified human-written or human-supervised content. I will much prefer hearing an organic, person’s personal experience in another country, for example, as a verification for things that actually happened, compared to a fictitious AI generated story. Now combine a verified human + AI assistance, you have a reliable narrator and good content.
The hardest hit will probably be to the anonymous writers and commenters, who cannot provide verification of being actual sources, and will be washed in the storm of AI content.
Many times people are bad at writing, I agree. I see what you mean, but if you have a concept and cannot express it, do you really "have" a concept. It is like knowing mathematics and being unable to solve any exercises.