Deterministic feedback is precisely how frontier models are trained. It’s called RLVR. You let the agent run on a problem and then calculate a deterministic score of how well it did. Repeat 1000x times and you can “brute force” a good solution. (Which includes all thinking traces and you add it to your training data.) And then a Chinese model can copy your advance for 1000x less compute. Which is why US labs call this not learning, but a distillation “attack”. It’s an attack on the business model.