Accuracy and other quantitative performance metrics are imperfect for sure, but how else do you propose testing before real-world deployment? How do you propose scalable and feasible testing of human students?
> The practice of medicine does not produce cute little cue-card prompts with 4 options.
Ah, but the New England Journal of Medicine's Image Challenges (designed to test the knowledge and diagnostic capabilities of medical professionals) does.
> It's of less-than-zero value to have an 80% machine accuracy and 75% clinician accuracy if the impact of the machine's mistakes are high risk, unconcerned with the impact of the intervention on the patient, or provide little therauptic upside
But this paper does not study the practice of medicine. It intentionally focuses on performance on one specific, well-known medical imaging diagnostic challenge.