They post train on big math and improve calibration across five different non math benchmarks so the new claim "can only give as accurate predictions and probabilities as the underlying data they are trained on represents" is also off base. RLCR generates its own calibration examples from ordinary questions and answer keys. RL usually isn't trying to represent a data set, it's closer to search.