One funny thing is that in order to tune the models to make what they're doing explainable to humans, you need to have humans involved in the RL pipeline to indicate which explanations are good.
You can understand this as learning a mapping between the model's internal "world" (i.e., 'meaning,' which is hopefully coherent and consistent -- but definitely not always! see, e.g., https://arxiv.org/html/2505.11581v1) and language (i.e. 'form') that reflects that world.
For this to work, you need both coherent / consistent internal model worlds, and also good mappings onto human language. Supervision by mathematicians has provided the signal for both internal coherence (though this can also come from interacting with a proof oracle) and for good explanations. If models exceed human capacities, you could imagine that aligning their explanations potentially becomes harder (though not necessarily). Also, humans naturally have to do the same thing: as researchers we must find analogies to make our work legible to collaborators or laypeople. Often in doing this, we further clarify our own understanding!
More deeply I think the "end of the world" vibe arises not only from the practical need to have models that explain, but also Litt's (and many other fields' researchers) grappling with being relegating to not mattering.