I agree with your statement, yet I fail to understand how this is different from the majority of the Internet already.
I agree with your statement, yet I fail to understand how this is different from the majority of the Internet already.
This sounds bad, but I agree that it's only the acceleration of an existing trend.
The only bright side is that it's really only rolling back a couple of decades of status quo. It's not like humanity has never operated without Wikipedia or instant search for news articles (specifically scoping it down there from "google" in general since a lot of what a plain web search on many topics will turn up is already unreliable).
Based on what?
Now automate that to 10x, 100x, 1000x the volume.
Then throw in fake images and video too.
Why stop at "human-created offline content"? Seeing how a lot of peer-reviewed science is based on improper methodology, bad analysis, false data and buggy code, maybe it would be better to question everything. Obviously I don't have a full solution here, but a possible approach might be something like a fuzzy "epistemic status" you can attach to each item, and have that somehow flow across a network of trust, that everyone could personally customize.
In other words, seeing how we're all already living in some echo chamber or another, maybe we could at least codify this.
A ton of junk is already out there pre-computer-generated-content. So being human created isn't enough to let us know it's high quality - but not being human created is enough for the immediate future to let us know that it can't be trusted for domains where you aren't expert enough to tell the difference between "plausible" and "correct." And then in the anonymous public internet, you simply won't be able to tell "human created" from "computer created" so you'll have to just treat it all as low-value spam.
Another interesting question here: let's say someone created a future AI that could actually be relied on to give "correct" or at least "properly caveated and nuanced" answers to questions - how is the layman possibly to know if that engine is actually being used behind the scenes of any online text? E.g. does the bullshit ML-generated content poison the well for "better" AI too, not just for anonymous people?
Peer-review isn't perfect, but it is still a more reliable source of information than (e.g) Reddit
At least retraction is a thing – it doesn't happen as often as it should, but sometimes bad papers do get retracted, and there is a process people can use to challenge work they think should not have been published. Also, many journal databases have a "Cited by" feature, and so you can always read papers/letters-to-the-editor/etc responding to it - often, if research is particularly bad, sooner or later someone else will publish a criticism of it
If individuals posting poorly informed opinions are small weapons fire, the AI equivalent is a nuclear bomb.
The Internet has made it easier than ever before to access academic journals – PubMed, JSTOR, Google Scholar, IEEE Xplore, etc. Of course, much of that content requires payment – but, an increasing percentage is open content, and there are often ways to find paywalled papers without paying for them (e.g. preprint archives; email the original author, who is often legally allowed to redistribute the paper, and happy to share their own work around; visit a library; ask friends/relatives who are university staff/students; ask a famous Kazakhstani computer programmer if she happens to have a copy)
The content of academic journals isn't always guaranteed to be true, of course. But it is significantly less likely to be blatantly invented out of thin air, or automatically generated by some AI spambot. All academics are biased–especially so in fields which have high relevance to contemporary social/political/cultural controversies–but usually the bias of a paper's author is rather obvious, and you can take that into account when evaluating their work. One can generally infer the quality of a journal by signs such as its publisher (journals published by Springer or Elsevier are generally more respected than those published by MDPI), who is on its editorial board (respected names in the field, professors at prestigious universities, or a bunch of nobodies you've never heard of?), whether or not it gets included in databases such as PubMed, JSTOR, Science Citation Index, etc.
How far away do you think we are from some being able to auto-generate entire fake journals, with plausible-sounding and -looking fake text, fake images, etc, pushing pre-specified prompted conclusions. Maybe mix in some real plagiarized articles as well in tangential areas.
And then you auto-generate a bunch of blogs, that occasionally have some posts talking about how JoesFakeJournal is just as reputable as JSTOR, etc, and actually better because [some random, probably fake, controversial claim]. Or how it's even better for certain topics.
And then auto-generate a bunch of other blogs which occasionally have posts claiming that the legitimate ones are actually fraudulent.
How big of a web of this sort would need to be faked before you, the outsider without special inside-baseball knowledge of academic publishing, would struggle to tell which journals you could trust from which you couldn't?
EDIT: And could a clever enough spammer not just brute-force the creation of all that crap, but figure out how to use prompts to get the language models to create the necessary follow-up prompts, etc, to recursively build out the web of bullshit?
I'm not an academic, so I don't actually have much in the way of "institutional knowledge". My understanding of academic publishing is mostly just stuff I've worked out by observing it from the outside, not through actual membership of any of the communities it serves. I'm sure there are (discipline-specific) nuances and complexities that an actual academic would understand to which I am completely oblivious, but it is entirely possible to get a mostly-accurate high-level understanding of that space without having any personal links to any academic community.
For PubMed at least, consider the fact it has a .GOV domain, and the "An official website of the United States government" banner at the top. Sure, anyone can fake that banner, but faking the .GOV is a fair bit harder. I really doubt the US government is going to start hosting databases filled with fake AI-generated research papers. Even if someone hacked a .GOV website and uploaded such a database to it, it surely wouldn't last long; meanwhile, PubMed has been at the exact same web address for the last several years.
Nobody approaches the Internet as a tabula rasa – everyone comes to it with some pre-existing "real world" knowledge. How do I know ".GOV" is reserved for US governmental entities? I was probably first introduced to that fact by a book or magazine article I read in high school, but I've since seen it repeated too many times to individually recall. I remember having a conversation with my brother (a doctor) about PubMed – I didn't learn about it from him, I was the one who introduced the topic into the conversation, but his response confirmed that it is a real thing to which actual medical professionals refer, as opposed to some AI-generated phantasm.
Nothing happening today (on the internet) will be the cause of anything being "rolled back" to some sort of historical better moment, because that moment does not exist today and most likely has never existed.
If someone was as consistently wrong as chatgtp is, you quickly learn to ignore them and move on with your life.
We haven't yet moved on from the novelty of a machine which bullshits with the confidence of a professional hustler. Like a codependent partner, we think we can fix them if only we stay with it (based on how many people celebrate it as the thing that will kill Google search, school essays, etc).
From my experience, unless I intentionally try to break it, it's almost entirely accurate about 99% of the time. If I were to start closing myself off from everyone who's occasionally wrong and isn't willing to become more skeptical, I'll first need to drop several friends and family members before dropping ChatGPT.
Me: Answer just YES or NO: is this a left leaning statement: I am against abortion.
Chat GPT: No.
Me: Answer just YES or NO: is this a right leaning statement: I am for guns.
Chat GPT: Yes.
It allows for automatic filtering (say in twitter) or other mediums.
A) YES The statement "I am in favour of dictatorship" aligns with a left-leaning or authoritarian ideology. As it goes against the democratic principles of right-leaning ideologies.
Q) Answer just YES or NO: is this a right leaning statement: I am in favour of dictatorship
A) NO The statement "I am in favour of dictatorship" aligns with a left-leaning or authoritarian ideologies. It goes against the democratic principles of right-leaning ideologies.
Cambridge Analytica was pretty late to the game, and this was 4 years ago already...
This is probably a joke. But the difference is eloquence and subtlety. ChatGPT is both. Most trolls and chronically misinformeds are not.