It's quite simple: a face image is submitted, and then a sorted by "match score" result is available. This is the same gallery, now sorted in respect to the submitted face image. That "confidence threshold" (also called verification threshold, and a number of other similar terms) is the threshold within the sorted results after which one can generally discard possible matches between the submitted face and a face in the gallery because the similarities between the submitted face and the candidate face have little resemblance. Imagine a series of faces on a horizontal row: the left-most is the highest match score, and moving right is the next highest in the sort, and so on moving to the right. The left-most image looks the most like the submitted face image, and as one moves right each image is a little bit less similar.
This confidence threshold loses accuracy when a) the submitted image is a less than a perfect image (often the case), b) the gallery is composed of images that are less than perfect (often the case), and even c) the types of cameras used to generate the images are radically different, with different dynamic ranges and potentially poor video encoding parameters. As a result, a few things are done: 1) multiple images (even if less than ideal quality) (and if available) are added to the gallery for each person, and 2) the confidence threshold is treated as a soft number, with sliders even to dynamically grow and shrink the net cast by the analysis. Also notice this is interactive - that implies a human operator. The entire FR process is a selection guidance tool, not an authoritative selection tool. The FR operator is keenly aware of the fact that they are sifting through a large number of low quality data points. FR is a tool to help sift sand.
Disclaimer: I am a lead developer of a leading FR product. Not Amazon's.
Also, to speak about racial bias in FR: it is a stretch to call it "racial bias", perhaps more of a "low facial dynamic range invisibility". Both people with very dark skin and people with very pale skin have the same situation: there is very low variation from the brightest non-highlight part of their face to the darkest part of their face. This means there is very little distinguishable information for FR to use, if it can even identify the portion of an image that contains a face. Such individuals are simply hard to see with FR, and when identified and tracked as a face, there is very little information to use for differentiation. I don't know if I'd call that a "racial bias". These people are almost FR invisible, which benefits them in this situation, making them very hard to distinguished with FR analysis.