I'm genuinely curious to find out whether anybody is surprised by this given that it's a language model that has been trained on almost all of the information and the data on the internet will beat human accuracy on tests - computers can retain memories and data a lot better than humans can and models like these simply have a lot more non-unique information than humans have consumed in their lifetimes.