Our method of evaluating quality is not super systematic right now. For this competitive landscape task, we have a "test suite" of ~10 companies and for each we have a sort of "must-include", "should-include", "could-include" set of competitors that should be surfaced. We run these through our tool and others and look at precision and recall on the competitor sets.
In terms of errors, right now our results are a little noisy, since we're biased towards being exhaustive vs selective. There are obviously irrelevant companies in the results that no human would have ever included. Our users can fairly easily filter these out by reading the one sentence overviews of the companies but it's still not a great UX. Actively working on this.