AI-generated content, other unfavorable practices get CNET on Wikipedia banlist
tomshardware.com
tomshardware.com
Some more discussion here: https://news.ycombinator.com/item?id=39556041
Wikipedia downgrades CNET's reliability rating after AI-generated articles - https://news.ycombinator.com/item?id=39556041 - Feb 2024 (10 comments)
Btw you can check this kind of thing (more or less) using https://hnrankings.info/39556041/ - it didn't catch the 8 minutes but the approximation is good enough.
I suppose I could've merged the current comments thither and re-upped that one, but I didn't think of it!
Welcome to the enshittening. There’s only so much a single dedicated operator can take before they pack it in. We need legislation to catch up fast and some big symbolic restitution cases decided in the courts.
>It's important to remember that while Wikipedia is "The Free Encyclopedia that Anyone Can Edit,"
The worst thing about AI is how it can easily betray you, manipulate you and have swarms execute long term sleeper plans at scale!
> It's important to remember that while Wikipedia is "The Free Encyclopedia that Anyone Can Edit," it's hardly The Wild West.
My reading is that the article is average human prose (not great, not unreadable), not LLM prose.
But as most Chess Grandmasters say, you only really need to use it in one difficult spot to change the result of a game.
Which is, of course, exactly what ChatGPT is trained to produce. A lot of people's mental AI detector is actually a mediocrity detector.
There's no average. Pre-training incentives being able to predict the smartest string of text in the corpus as readily as the dumbest. It doesn't converge on "average" and it doesn't really make sense that it would either.
Base models don't talk like GPT. This is strictly an artifact of post training fine-tuning/RLHF.
It's not average human prose either(for one thing expecting GPT to converge on some average doesn't really make sense in the first place).
Base models with no rlhf or fine-tuning don't talk like that at all. This is specifically an artifact of the post-training fine-tuning/RLHF process
"It's important to remember" is a phrase that plenty of people used before ChatGPT. I've included three examples below from a quick time-boxed search. Just because ChatGPT says a phrase doesn't mean that every time you see it it came from ChatGPT—someone wrote the training data that ChatGPT was trained on, and someone else wrote the data (or selected the responses) that it was fine tuned on. ChatGPT isn't inventing new phrases out of whole cloth, it has a stereotyped style that is pieced together out of many existing phrases.
A collection of such stereotyped phrases in a single piece would be stronger evidence of GPT authorship, but I see no evidence of that here.
https://old.reddit.com/r/gravityfalls/comments/bgtnez/while_...
https://twitter.com/AsteadWH/status/1050813462673264640
https://stackoverflow.blog/2019/12/19/what-senior-developers...
"It's important to remember to consult a mechanic" for example.
It's just some priming.
not THE reason, not THE ONLY reason, so it is not COMPLETELY untrue. agitated?
Someone who's studied the subject quite a bit might visualize an actual Einstein chalkboard from memory. Their output could be verified and referenced, but looking for new Einstein-level math on the chalkboard would be madness.
Someone who Einstein himself would consider a peer might use this visualization method as a way to do actual work.
If we're assigning value to our chalkboards we'd be able to explain why we chose the numbers -1, 0, and 1. This would bias any math we'd do towards one of the chalkboards depending on our intent. At this point my chalkboard is useful as a filter.
Putting this all together, we'd see that the overall look of the cartoon responses would be a blend of 0 and 1 styles and would depend on our requests e.g. a reference request would look mostly like 0's art style. My own personal art style will be intentionally absent because it's only ever framing nonsense, by my own admission.
I'm guessing that a human did edit this AI output, just not very well.
I'm still surprised to see that, as far as I can tell, no news outlets have made a public commitment to never use AI in their writing. Seems like it would be an easy way to promote the brand on commitment to quality.
I've griped about this before, but here we still are. We now know that MSN news, for instance, has no credibility due to their publishing AI-generated misinformation. https://news.ycombinator.com/item?id=39043135
Given its nature as an LMM and a complex next word predictor, the phrase "it is important to remember" could be a way that was inadvertently trained so that it keeps itself on track.
If it has "some points, it is important to remember {something}, some more things" it may be able to better generate text compared to "some points, weird tangent".
Since it doesn't have a hidden memory, everything that it "thinks" is out there in the text including its own cues for what it should do. It also can't go back an edit its previous text to remove the self hints or clarify earlier points without calling them out. That style of writing differs from natural human writing since we are able to keep on topic (or not) without needing to write messages to ourselves that others can read.
When we do, it's pointed out rather than trying to slip it in casually. "Note to self or reader" for things that are to be pointed out and break the flow of the text or "as an aside" for the tangents.
Wikipedia is supposed to use primary sources, AI generated articles, by nature, can't be primary sources. In particular, AIs love to use Wikipedia in their training dataset: it is a free, high quality source of information, but it is not flawless either. So there is a good chance that if Wikipedia cites an AI generated article, it has Wikipedia as its source, starting the "citogenesis" process.
https://en.wikipedia.org/wiki/Wikipedia:No_original_research....
To wit: AI articles are to be found along a dramatic spectrum of quality. If such an article is high-quality, relevant to the subject matter, asserts the fact in question, and makes proper and veracious use of a primary source in support of such an assertion, why isn't it a reasonable source for an encyclopedia?
I can imagine a future with rich educational materials with these layers:
* raw experimental data -=>
* publication ("primary" source) -=>
* AI-generated review of many publications ("secondary" source) -=>
* Human-authored encyclopedia article (with one or two people following the fact pattern all the way back down to the data, and many more people helping to synthesize the higher layers into a rich, readable, considerate, diverse synopsis)
An article that says Abraham Lincoln wrote the emancipation proclamation would be preferred to Abraham Lincoln himself being asked and then the Wikipedia article citing “my interview with Abe”.
(Your general point is valid, I’m just confused about the terminology)
The final hurdle at that point will be the de-democratization of writing and the dilution of creativity and novel writing. There will probably always be a market for that, but for things like reporting events, it seems like AI could easily overtake the industry.
Here's a spoiler-filled article from a major site (IGN) whose last paragraph (and likely more) has clearly been authored by ChatGPT in order to meet tight publishing deadlines and coincide with a movie's release: https://www.ign.com/articles/dune-part-2-post-credits-scene-...
The indicators are all there: "In short, (vague sentence)... Is (hypothetical from article) or (other hypothetical from article)? Will (vacuous statement) or (inverse vacuous statement)? We'll have to wait for (eventuality) to find out."
There is no attribution to AI for the generated content, though, and this lack of attribution is going to become the norm once LLMs become just another authoring tool like spell-check. Coupled with the race for clicks, the "excellent blog posts" are going to be drowned out.