what I would like to see is a parameterized class of prompts which can never be solved by the LLMs even when a finite number of them are manually added to the dataset.
IE, you're getting into areas that are analogous to Turing's theories. I don't think he came up with those theories overnight.