"An LLM might be able to explain something to you, but it can never understand it for you."
And then something like this [0] will happen, creating a weird wasteful meta-game about the model(s) used by each company.
It makes intuitive sense: If you outsourced "generate awesome assertions" to a contractor, then someone else hired the same contractor to "judge the awesomeness of these assertions", they are more likely to get lots of "awesome" results—whether they're warranted or not. The difference might even come down to quirks of word-choice and formatting which a careful human inspector would judge irrelevant.
[0] https://www.theregister.com/2025/09/03/ai_hiring_biased/
As anyone in the US healthcare insurance claims-adjacent spaces can tell you: yes, in about 2 years.
the boss is nowhere