I believe this is one of the LLMs' more destructive elements. The more effective ones are quite good at providing output that gets the meat to repeat the output to other meat. Whether the output actually benefits the meat is disconnected from whether or not the meat accepts and repeats it.
That is not realistic, but I suppose where things are heading is that you have some indicator of the strength of evidence -- see Fig. 13 in the following insightful take:
https://news.ycombinator.com/item?id=49407226
Though, "strength" should probably be "reliability" and "validity", and I suppose those indicators are more for picking signals from the noise; i.e., what is even worth clicking and reading. That would be increasingly valuable already today due to the volume (and, yes, slop and other related stuff).
And to expand on this, it's not realistic because science is not armchair philosophy. You have to go out and measure the world.
Sometimes, through force of will, a person can think deeply about a problem and come up with beautiful theories that explain our measurements. Many scientists had careers like this, probably most famously, Einstein. But it's worth noting that Einstein also got a lot wrong! [1]
Even if we somehow give an LLM the ability to go out and measure things, I seriously doubt that the role of humans in science is done. There's a big difference between "an explanation" and "a good explanation." Ask any physicist. There's a surprising amount of aesthetics involved. Good theories are consistent with the evidence, but it's more than that-- there's a great deal of "taste" involved. And there's a good reason for that. For any real problem, there are effectively an infinite number of alternative hypotheses. From a "theory of science" standpoint, this should cause scientists nightmares, but it doesn't. Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation. If you spend time with scientists, especially in the "hallway track" at a conference, "taste" is a frequent topic of conversation!
[1] https://en.wikipedia.org/wiki/Einstein%27s_unsuccessful_inve...
I work full time on "lab in the loop" AI, so I'm pretty familiar with the need for real-world experiments. I am not proposing a fully autonomous scientist that could read an arbitrary paper and emit whether it's universally true without some verification method.
Also, to your statement: " Because by the time you are a practicing scientist, you've developed a feel for what constitutes a satisfying explanation."
I'm a practicing scientist (well, ex-scientist) and it seems like most "satisfying explanations" end up being wrong or incomplete simply because they seem so satisfying.
What's the difference? How do you find real mistakes without a model? Either you have a trusted mathematical model (in which case you already have a complete explanation) or you have to compare it against the ultimate oracle: the world. Or are you proposing something like "let's use an LLM to convert this hand-wavy English paper into a formal proof and then check it for logical fallacies?" In which case, fine, that would be useful, but that's not exactly the same thing (and also not as important) as saying that a paper advances a bad explanation. Just that the explanation is flawed in some way.
The next example I can think of- I am not sure it qualifies. I read a paper where they deleted one gene at a time in yeast (it has 6000 genes) and determined whether the mutated yeast could live or not. For each gene where the yeast died, they added that to a list of "essential for life" genes. The paper concluded they had found some interesting proteins that should be studied. I read the paper and the first thing that sprang to mind, are any of these genes overlapping? Because we know (somebody already demonstrated in a lab) that genes do overlap (which is truly weird!)
I wrote a script and showed that every gene they reported as essential for life overlapped an already known gene that was essential for life. I wrote the authors, who never responded, but wrote a followup paper where they acknowledged they probably had a high false positive rate due to overlapping genes with known fatal effects. My guess is you'd say that either I used a model (existing literature) or I compared against the world, but realistically, what I did was trivially come up with a better explanation than the authors. T hat's what I want LLMs to do for me, and it seems like the direction LLMs are going will fulfill my desires.
To me, a single significant error of any kind brings the entire paper into question. If a less important figure contains an image duplication, that makes me wonder if I can trust any of the images.
In any case, my point was that rejecting results due to technical flaws is intellectually lazy. As a scientist, your job is more about trying to find value in other people's work than finding excuses to reject it.
I disagree with your premise. Part of our job as scientists (thankfully no longer mine) is to reduce the irrelevant and incorrect noisy as early as possible. I have seen so many grad students get excited by a paper and put enormous effort into reproducing somethign that was a false or fake result.
If you want to determine reliably whether something is irrelevant and incorrect noise, determining whether there is anything of value is a necessary first step.
I've seen many reproduction attempts in bioinformatics fail, because they person trying to reproduce the work didn't have the conceptual background to do it correctly. Instead of spending enough time studying the theory, they rushed directly to action.