Before you can even do the statistical analysis you suggest, you need large amounts of high quality data—which we don't have. One place where the US (and the world?) gets data privacy right is in healthcare, but unfortunately that also means it's nearly impossible to create the data sets we need to do the statistical analysis you want.
Institutions face severe penalties for wrongfully sharing patient data, so most opt to just not share any data. Any research that is performed is done internally on local populations with de-identified data sets. A few brave institutions go well out of their way to create and share de-identified data sets publically, but these data sets still undersample the general population. This is a critical problem because certain diseases are highly prevalent in certain regions (e.g., Lyme disease in New England) but unheard of in other regions (e.g., Lyme disease in Colorado). If your ML model is trained on data largely from New England, it's going to diagnose a patient with the classic "target-shaped" rash with Lyme disease even if the patient is from Colorado (high false positive rate). If the model is trained on data from Colorado, it will underdiagnose Lyme disease in patients from New England (high false negative rate). The only way I know to overcome this problem is to create even larger data sets, but this just isn't possible with data privacy laws.