The "person one" vs "person two" bias seems trivially solvable by running each pair evaluation twice with each possible labelling and the averaging the scores.
Although of course that behavior may be a signal that the model is sort of guessing randomly rather than actually producing a signal.