I’m an IO psych at a large tech company that evaluated pymetrics against job performance and there was no relationship, just near 0 correlations so recommended not using them and we don’t. These AI tools are not transparent enough in explaining their outcomes and lack what we call face validity or job relevance. A coding test is at least a good filter at the top for entry level roles with large pools because it has relevance and false negatives aren’t as much of a concern (but we test still for disparate impact) otherwise for most roles a structured interview is the best option.