This is a great point for what I was trying to explain about where we currently stand with these models - we seem to be hardware constrained in relation to the trade off between rote learning, full memorisation of information and the model's ability to use what it knows and generalise to achieve an outcome.
Your experience aligns with what I've seen: 4o and o1 are better at filling gaps from each other where there should be an upward trend of intelligence instead of having to jump around as we do currently.
So the real question is: are these models forgetting their skills, do we need to make them larger? Does distillation and generalisation work?
So many questions that might end up as LLMs being unable to escape the problem of catastrophic forgetfulness.