No person or AI needs to actually understand how an AGI actually works, they just need observable behaviours indicative of intelligence. That's what we already do with image recognition, language translation tasks, and more.
Once we move to more general forms of intelligence, like math and other forms of problem solving, then you apply more general tests.
For instance, if we had an Einstein AI v1.0 that could infer General Relativity using X0 bits of information in time T0, and Einstein AI v2.0 managed to infer GR using X1 << X0 bits of information and/or could infer it in time T1 << T0, then v2.0 is clearly superior to v1.0.
I agree that sometimes the parameters for such tests aren't always clear at first, but this has always been the case in science, which is we we iterate and progressively refine our tests as we learn. Any true AGI must be capable of such abductive and inductive reasoning to qualify as a general intelligence.