LLM's still make stuff up routinely about things like this so no there's no way this is a reliable method.
The only reason why this is helpful is because humans have natural biases and/or inverse of AI biases which allow them to find patterns that might just be the same graph being scaled up 5 to 10 times.
Having seen from close-up how these reviews go, I get why people use tools like this unfortunately. it doesn't make me very hopeful for the near future of reviewing.
It is a tool, and there always needs to be a user that can validate the output.