I don't really see how replacing the vehicle and the animal is a good test.
It'd be better to just have it draw a completely, linguistically, unrelated scene.
Like, a single tree in a meadow bending in the wind.
Then it’s not a benchmark. I think his thesis is solid: if neither pelicans nor bicycles stick out, it follows that there isn’t special attention being given to them by the labs
I do think that the bicycles stick out. They all look remarkably similar, aside from the DeepSeek test.
Feels like it's been pretty obvious that they have been for a while now