It's more likely to be different specialisations. Most of the people doing the evaluations are more data sciencey ML type people, rather than software engineers. This isn't helped by their culture which is very much driven towards alignment as the only possible solution to super-intelligence (which may be true, but I have my doubts that this will happen in any reasonable time frame).