The evaluation was done on 246 X-rays, which is good; it would be better if they were not all chest X-rays (to see generality), and if there were more than only four radiologists from the same country.
The achieved 0.25 clinically-significant errors is impressive, although it is only for the best of 3 models, which can incur bias; averaged across models, it is 0.27, a bit worse than human error. Additionally, I am wondering where they get the human baseline; they state:
> These results are on par with human baselines from prior work [14]
but the citation[1] doesn’t give data in the same format (and its format is honestly better: it indicates that humans make no urgent errors or worse in 64% of reports).
Surprisingly, there is no improvement with model size; the largest model performs the worst.