> Results: Full concordance with the reference diagnosis was 8% (27B) and 5% (1.5 4B; McNemar p=0.68), while partial matches were 29% vs 20% respectively (McNemar p=0.053). When correct diagnoses anywhere in the differential were counted, 51% (27B) vs 30% (1.5 4B), with 27B significantly superior (McNemar χ²=12.1, p=0.0005). Site-level performance varied widely (30–100%). Both models reported HIGH confidence in ~99% of cases irrespective of correctness.
i.e. highly confident, wrong 95% of time. in 49% of cases the real diagnosis wasn't even on models' differential. Doctor can hardly improve using something they can safely assume to be just noise.
https://ecp2026.abstractserver.com/programme/#/scientific/de...