I think the point of the study was precisely to see how well the original model without adaptations (my understanding is that they only use prompting, which does not affect weights) can perform. I think that question is arguably even more interesting than 'How well can we train a model to perform on med school questions?'. I'm not saying this is generalization, because the training set surely included a lot of medical literature, but if the base model without fine-tuning can perform well in one important domain, that's a very interesting data point (especially if we find the same to hold true for multiple domains).