I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
They result of the benchmark does not feed back into the training, it simply serves to provide a measurement of progression over time.
It's not the direct feedback loop of RL but its not far.
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
It's actually close to the second. "Here are 200 software questions, will drill you on *other stuff* until you can pass exactly these 200. If the other stuff isn't improving your scores we will change ratios of it till it does."
The reason it benchmaxes is that *other stuff* ends up looking more and more like SWE Bench without you realizing it.