This reminds me of this article for which symbol fine tuning helps overcome: https://arxiv.org/abs/2303.03846
Basically they show that larger models have an easier time giving up their semantic priors, so that, in the context of OP paper, learning the mapping for len becomes print and print becomes len.