I wouldn't be surprised if models were optimizing for pelican-related comment chains at this point
It's a silly fun little benchmark, and because Simon's been doing it for so long, you have a lot of examples over the years to compare. But you can always come up with and run your own test with other drawings.