We still don't beat humans in e.g. LeanDojo, to my understanding.
But we also haven't yet thrown the state of the art in graph attention technology at doing the core task of in-context learning (consider the entire explored (partial-)proof "tree" for deciding what tactic with which premises (if the tactic takes an AST/expression argument, like "use that definition/equivalence there as a rewrite rule on the current goal at this node in the proof tree") to apply next at which node in the (partial-)proof tree) in the interactive theorem proving situation, let alone with reinforcement learning that uses the attention structures to funnel back information on what things in dead branches had responsibility for the decisions that made the non-dead branches.
Because the RL agent should happily use the theorem prover as a "calculator", and only get "charged" for how much the ML inference and the actual cheap calculator execution really cost.
Exploration in-context is what humans also do, and immensely powerful. Similar to, when programming, testing some expressions to figure out edge-cases before incorporating them into a large function/algorithm that has poor observability.
So far, machines still mostly do the task of "computing" in theorem proving for us, rather than show some proper intelligence. It's damming that LeanDojo got so high marks without reinforcement learning. They pretty much just dumped prover context state during a run through the entire matlib "standard" library and handed it to two transformers, one for embedding premises and one encoder-decoder fed with the current state and as many premises as the encoder context fits after the state, generating/sampling a few tactics to apply.
No freedom of where in the proof tree to apply the next tactic, no awareness of any such side-branches.