AI tools are spotting errors in research papers: inside a growing movement
nature.com
nature.com
Wonder which is more common overall? Can AI spot more errors than it creates, or is it in equillibrium and a net zero?
Other than these quips, I am actually a fan of this movement. Anything that helps in the scientific process of peer review, as long as it does not actively annoy and delay the authors is a welcome addition to the process. Papers will be written, many of them incorrectly, with statistical or procedural errors. To have a checker that can find these quickly, ideally pre-publishing, is a great thing.
But concerned about tool dependency; if you can't do it without the AI, you can't support the code.
"LLMs cannot find reasoning errors, but can correct them" (2023) https://news.ycombinator.com/item?id=38353285
"New GitHub Copilot research finds 'downward pressure on code quality'" (2024) https://news.ycombinator.com/item?id=39168105
"AI generated code compounds technical debt" (2025) https://news.ycombinator.com/item?id=43185735 :
> “I don't think I have ever seen so much technical debt being created in such a short period of time"
Even if trained on only formally verified code, we should not expect LLMs to produce code that passes formal verification.
"Ask HN: Are there any objective measurements for AI model coding performance?" (2025) https://news.ycombinator.com/item?id=43206779
https://news.ycombinator.com/item?id=43061977#43069287 re: memory vulns, SAST DAST, Formal Verification, awesome-safety-critical