Actually, for what you are mentioning, it is getting worse. There was a sweet spot somewhere around the release of GPT-o3, and ever since, the LLMs have been getting more accurate at solving problems, but worse at explaining how, and to hone in on what is interesting. This isn't surprising, as RL strategies shifted from RLHF to RLVR, so priorities during learning changed. I don't expect AI labs to reverse course on this. We can expect AI proofs to become increasingly incomprehensible over time.