On prompts only, with answers presumably from the teacher model (Gemini).
It was not trained or RLHFd on Arena replies or user preferences.
It was not trained or RLHFd on Arena replies or user preferences.
At the end of the day, by optimizing for leaderboard scoring, it makes the leaderboard ranking less useful as a benchmark (Goodhart's law strikes again). The Gemma team obviously isn't the only one doing it, but it's important to be clear-eyed about the consequences.