We have rare but not unheard of issues with academic fraud. LLMs fake data and lie at the drop of a hat
We can do both known and novel reproductions. Like with both LLM training process and human learning, it's valuable to take it in two broad steps:
1) Internalize fully-worked examples, then learn to reproduce them from memory;
2) Train on solving problems for which you know the results but have to work out intermediate steps yourself (looking at the solution before solving the task)
And eventually:
3) Train on solving problems you don't know the answer to, have your solution evaluated by a teacher/judge (that knows the actual answers).
Even parroting existing papers is very valuable, especially early on, when the model is learning how papers and research looks like.
A bit like how you might write a paper yourself - starting with the data.
As it turned out I thought the figures looked like data that might be from a paper referenced in a different lecturers set of lectures ( just on the conclusion, he hadn't shown the figures ) - so I went down the library ( this is in the days of non-digitized content - you had to physically walk the stacks ) and looked it up - found the original paper and then a follow up paper by the same authors....
I like to think I was just doing my background research properly.
I told a friend about the paper and before you know it the whole class knew - and I had to admit to the lecturer that I'd found the original paper when he wondered why the whole class had done so well.
Obviously this would be trivial today with an electronic search.
Or maybe give it a paper full of statistics about some experimental observations, and have it reproduce the raw data?
Practically speaking, I think there are roles for current LLMs in research. One is in the peer review process. LLMs can assist in evaluating the data-processing code used by scientists. Another is for brainstorming and the first pass at lit reviews.
After ChatGPT big cooperations stopped sharing their main research but it still happens at academia.
It would be the biggest boon to science since sci-hub though.
And since a large set of studies won't be reproducible, you need human supervision as well, at least at first.
The main reason people don't do it is because incentives are everything, and university/government management set bad incentives. The article points this out too. They judge academics entirely by some function of paper citations, so academics are incentivized to do the least possible work to maximize that metric. There's no positive incentive to publish more than necessary, and doing so can be risky because people might find flaws in your work by checking it. So a lot of researchers hide their raw data or code for as long as possible. They know this is wrong and will typically claim they'll publish it but there's a lot of foot dragging, and whatever gets released might not be what they used to make the paper.
In the commercial world the incentives are obviously different, but the outcomes are the same. Sometimes companies want the ideas to be used as they compliment the core business, other times the ideas need to be protected to be turned into a core business. People like to think academics and industrial research are very different but everyone is optimizing for some metric, whether they like it or not.
Producing novel ideas is the most famous trait of current LLMs, the thing people are spending all their time trying to prevent.
Could you please explain what you mean or give a simple example?