My experience in dealing with data (we had sufficient, and somewhat well labeled) & methods made me realize that a lot of the prediction human doctors make are multimodal - and that is something deep learning will struggle for the time being. For example, say in detection of a disease X, physicians factor in blood work, family history, imaging, racial genealogy, general symptoms (like hoarseness, gait, sweating etc), even texture & palpitations of affected regions sometimes before narrowing down on a set of assessments & making diagnostic decisions.
If we just add in more dimensions of data to model, it just makes the search space sparser, not easier. Throwing in more data will likely just fit more common patterns & classes well, whereas a large number of symptoms may be treated as outliers and mispredicted.
We humans are incredibly good at elimination of factors & differential diagnosis. The findings don't surprise me. There is much more work needing to be covered. For straightforward, and conditions with limited, clear cut symptoms they are showing promising advancements, but it cannot be trusted to wide arrays of diagnosis - especially when models don't know what 'they do not know'.