(1) A lot of researchers are bad at writing code, but they can audit it. This is true for sociologists, psychologists, etc. so I'm hoping something like this can help.
(2) Philosophically, I disagree with the debates that LLMs can't produce new knowledge. I think there's merit to this if we're talking about whether the LLM neural network itself synthesizes new knowledge via its weights... However, why can't we have an LLM try and merge multiple data sets, analyze them, and report back to a human?
To your point + concerns, I think a human still needs to be very careful and actually revisit the analysis for any promising findings, but at least some of the grunt work can be taken care of!