Nice idea and all, but I think their methodology was totally unsuited to the task. What they actually assessed was whether an LLM can recognize an accurate summary of a given article, not whether the LLM can produce an accurate summary itself. Here's their description of it:
> We used a 3-way verified hand-labeled set of 373 news report statements and presented one correct and one incorrect summary of each. Each LLM had to decide which statement was the factually correct summary.
The problem with this approach is its assumption that if an LLM can recognize an accurate summary, it'll be able to reliably produce accurate summaries. We know very little about the inner workings of LLMs right now, and what we do know suggests that they work highly counterintuitively, so I think there's no basis to make this assumption.