The whole point of his model is to optimize for a very specific benchmark.
BUT, he does not use labels when training, so the model does not know the answers.
BUT, he does not use labels when training, so the model does not know the answers.
But benchmaxxing is what we generally try to avoid for training, as there is no point really for it. We used to call it "overfitting", now you're saying this person does it intentionally? Why?
I would not call this overfitting, it's finetuning for specific task where you have a benchmark.
Also, the complexity of the task he is using occupies an interesting middle ground of ultra high dimensionality (for a “simple” problem) while being limited in width to a narrow set of solves- a space where one would be tempted to imagine you would need a much more capable system.