And is RL not? All of these models are constrained by finite weights and then tuned. Are you suggesting we grow a neural network until certain criterion are met with regard to out of distribution test criteria? Hmm
I'm saying that we shouldn't expect the models to come up with things we didn't train them to come up with.