So to test for unknown bias and sensor anomalies, they put raw modelled data through the statistical machinery looking for homogenized outputs.
This strikes me as wrong. If you want to know if your filter is bad, you plug in ground truth and add known noise models. If you want to know if your noise models are wrong, you cant do the same thing, you need to point your sensor at a known object (e.g. calibrate).
They seemed to have done a cursory analysis of the former type, which is not the same as saying "our pea-pod hyoothesis is correct".