AI summaries turn real news into nonsense, BBC finds
theregister.com
theregister.com
"In December 2024, the BBC carried out research into the accuracy of four prominent AI assistants that can search the internet – OpenAI’s ChatGPT; Microsoft’s Copilot; Google’s Gemini; and Perplexity. We did this by reviewing responses from the AI assistants to 100 questions about the news, asking AI assistants to use BBC News sources where possible. Ordinarily the BBC ‘blocks’ these AI assistants from accessing the BBC’s websites. These blocks were lifted for the duration of the research and have since been reinstated. AI answers were reviewed by BBC journalists, all experts in the question topics. Journalists rated each AI answer against seven criteria – (i) accuracy; (ii) attribution of sources; (iii) impartiality; (iv) distinguishing opinions from facts; (v) editorialisation (inserting comments and descriptions not backed by the facts presented in the source); (vi) context; (vii) the representation of BBC content in the response. For each of these criteria, journalists could rate each response as having no issues; some issues; significant issues or don’t know."
They didn’t provide the answers models provided, so we can’t even say if the reviews of the answers were correct. And since the experiment relied on BBC giving timed access to their archive to the models, we couldn’t even type in the same prompts into the models to see the model outputs.
It would be interesting to see the experiment repeated, but with the BBC feeding the AI PDFs/text of the initial reports. I suspect it would be much more accurate.
There you will find at least some examples of questions and answers. But it is true, the whole raw dataset has not been published, yet.
Pretty good though, for consumer-grade, off-the-shelf AI.