1) I'm leaning more into a broader sense, which is that given the priors a system already possesses, how efficient can it acquire competence on a novel task? If I'm reading the point you make, you're focussing on continual learning right? If so, I'm not necessarily restricting my statement above to that.
Here's another reframing: How much of the benchmark improvements come from overwhelmingly large training distributions vs improving the models for adapting to things genuinely outside of it?
2) Excellent points about crystalized and fluid intelligence. Wouldn't the LLM scaling gains be a representation of crystallized capabilities? In regard to Gf, that is exactly what I am asking about. That is what seems to be lacking, Gf like adaptation under genuine novelty.
My concern is that it's increasingly difficult to tell of what looks like Gf like behavior is really coming from better adaption vs. having broad priors from the model's large learned distributions.