I suspect it wouldn't help too much. This model is meant for physics-based world modeling, while nearly all the problems in ARC are symbolic reasoning.
As usual comparisons with humans provide little practical insight for what's achievable with ML. Humans don't have to learn everything from scratch like ML models do, you aren't expecting ML models to learn language out of a few thousands of tokens just because humans can, so similarly you shouldn't expect neural networks to learn reasoning from world interaction alone.