There's also the "horse riding astronaut" challenge in image generation: https://garymarcus.substack.com/p/horse-rides-astronaut-redu...
He also claimed that LLMs were a failure because of prompts that GPT 3.5 couldn't parse, after the launch of GPT-4,which handled them with aplomb.