I understand why it wouldn’t be feasible for a human to do this, but I’m quite sceptical about an AI assessing how accurate predictions turned out to be/how contrarian they were at the time. It seems like that would depend a lot on what sources it chooses, be liable to hallucination or getting poisoned by bad sources, etc. They don’t mention whether they used independent queries for each prediction either, or whether it was doing multiple sequentially.
Given that LLMs can’t really distinguish prompt from instructions etc, I’m sceptical that they can reason particularly well about things like how contrarian a view was at a particular point in time.