As noted elsewhere, we give confidence scores between 1-99%. We also use many different models for each modality for a more robust and complete answer with each scan, and each model has its own confidence score.
That doesn't fix the fundamental potential for abuse, moral hazard, and accountability sink[1].