This is far from unsolvable. It just means that the "apply RL like AlphaGo" attitude is laughably naive. We need at least one more trick.
As you said brute forcing the search space as the starting procedure would take way too long for the AI to build intuition.
But if we could give it a million or so lemmas of human math, that would be a great starting point.
[1] https://www.vice.com/en/article/a-human-amateur-beat-a-top-g...
That said, reachability and novel strategies are somewhat overlapping areas of consideration, and I don't see many ways in which RL in general, as mainly practiced, improves upon models' reachability. And even when it isn't clipping weights it's just too much of a black box approach.
But none of this takes away from the question of raw model capability on novel strategies, only such with respect to RL.