> weight update during eval: this is a form of test time training and not really cheating.
Possibly "not really cheating", but it does make benchmark comparisons unfair - especially as the other models are unlikely to have their weights updated during the eval.