For example, it's long been known in the physics education research community that students come away from introductory courses with very little physical understanding, even if they can do the plug and chug problems on typical tests just fine. Students can all recite Newton's third law, but immediately afterward claim that when a truck hits a car, the truck exerts a bigger force. They know the law for the gravitational force, but can't explain what kept astronauts from falling off the moon, since "there's no gravity in space". Another common claim is that a table exerts no force on something sitting on it -- instead of "exerting a force" it's just "getting in the way".
For research purposes, we measure physical understanding using a battery of tests, such as the Force Concept Inventory, containing only simple conceptual questions with unambiguous answers. So then everybody asks: if ordinary tests are so hackable, why not just switch to these conceptual ones? But that wouldn't work. There are less than ~100 distinct FCI-style questions. If these conceptual tests were the norm, students would just memorize the answers and parrot them back, with a flimsy understanding that crumples the second any follow-up question is asked. It would be just the same problem as before, except they would be worse computationally, too. The FCI only works as long as it doesn't count for a grade.
The problem isn't tests, it's scale. If the people aren't motivated, any standardized measure will miss the mark -- even the entrepreneurship Paul Graham advocates for. God knows I've seen a lot of bullshit in that direction.