The practice of medicine does not produce cute little cue-card prompts with 4 options. Diagnosis means trading the risk of various interventions against their diagnostic value, including a wait-and-see strategy. It means consulting with the patient on impact, as well as conferring with other specialists.
It's of less-than-zero value to have an 80% machine accuracy and 75% clinician accuracy if the impact of the machine's mistakes are high risk, unconcerned with the impact of the intervention on the patient, or provide little therauptic upside. Likewise, on these decision benchmarks *Machine decisions are NEVER iterated* -- this makes any claim to "performance" not even borderline pseudoscience, but imv, malpractice.
In the realworld practitioner decisions are iterated: doctors do not make high-risk mistake on every sequential decision, given more diagnostic information. It is highly highly unlikely that an inaccurate high-risk judgement call is compounded on each intervention. Whereas machine decisions are never tested in this sequential manner -- if they were, their benchmark performance would drop off a cliff.
The production of non-iterated decision benchmarks which measure accuracy and not risk-adjusted real-world impact constitutes, imv, basic malpractice in the ML community. This should be urgently called out before non-experts think that LLMs can give credible medical advice.