Here is a direct link to the arxiv article:
https://arxiv.org/abs/2305.14292
WikiChat: A Few-Shot LLM-Based Chatbot Grounded with Wikipedia
WikiChat: A Few-Shot LLM-Based Chatbot Grounded with Wikipedia
Like we invented this new thing and this new measurement for evaluating it. It does great on the metric we just made up while we were making it.
[1]: We can discuss if ChatGPT passes the Turing Test or not, but I think we can now all agree that being able to have a convincing conversation is not a good test for intelligence.
Of course, one should always be critical of benchmarks, and there is an obvious opportunity for bias here that should be reviewed with care. But your phrasing suggests that this is unusual or actively suspicious, which it is not.