You and the article are both correct. The disease does present itself differently as a function of these other characteristics, so since the training dataset doesn't contain enough samples of these different presentations, it is unable to effectively diagnose.