Exactly the way we've been doing it: compare what it produces against known intelligence to find cases where it fails.
What we’ve been doing has produced a lot of impressive stuff, but not intelligence, has it? Again, hard to tell without a good definition.
And we've learned a lot about what is and is not intelligence, and if we keep it up we'll steadily carve out the space that defines it.
It's not hard to tell that it has not produced human-level intelligence, so the process outlined by naasking has not yet run into an insurmountable problem.