> I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling).
Given that GDM pioneered RL, that's a reasonable assumption
Given that GDM pioneered RL, that's a reasonable assumption
RL was established, at the latest, with Q-learning in 1989: https://en.wikipedia.org/wiki/Q-learning
i still think my original statement is fair
people who knew from context that your statement was broadly not actually right would know what you mean and agree on vibes. people who didn't could reasonably be misled, i think.