Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.
https://dylancastillo.co/posts/pelicanmaxxing.html
Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world.
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.