E.g. -- a "differential discrimination" test of sorts: altering any one variable of a "protected class" such as race, gender, religion, etc. does not change the answer. You would maybe want to pick a set of canonical test profiles among real people who differ only on one axis (as closely as possible), rather than just take a test point and alter one axis directly, because you'd want all the relevant correlations (ZIP code vs wealth, etc) to remain authentic. The end result would be a set of "equivalence classes": sets of human profiles who must be considered equivalent on all relevant life-altering judgments.
Or perhaps a "unit test"-like approach: similar to how one creates a unit test for each bug one fixes, create a "criminal justice ML algorithm test suite" with canonical profiles and their results: you must judge this person to likely not re-offend, you must judge that person as a high risk, you must judge this person worthy of a home loan for $X, etc. Sort of like a body of case law. I guess the risk is overfitting -- so maybe this data set is held in trust by some regulatory agency and not revealed.
People have probably thought about this and I haven't read your links -- is building a test data set and building regulations around it something that's considered?