The whole craze about my application i.e. dermatology successes in ML spurred from the 2017 Nature paper, where skin cancer was detected as good or better than dermatologists. But technically, such experiments have design problems: We had apriori knowledge of the dataset (White N American Melanoma data) & hence we could ascertain the model performance. Real world data is much more variable. Further, later it was revealed that ML model latched on to the little marking physicians made rather than generalizing on lesions. 3 years later my experiments could model reliably only on 10 very common diseases & of a very uniform skin type and ethnicity. Those results were nowhere close to perfect.
The proposal to keep Human-in-the-loop is a much fairer alternative in ML aided medicine, than end to end machine learning. Most direction of research is headed that way. Physician assistance is much more reliable than potential replacement.
The trouble with synthetic data is that it doesn't address the extended variability in real population & generalization will always suffer. Also, doctors take multi-path decision, choice by elimination, past cases - based on several diagnostic inputs & even gut intuition. At that scale of input multimodality, Type I & II errors are at a scale higher than correct identifications. And we don't know how to teach intuition or imagination to machine models well enough. Those fall back to rule based methods & edge cases.