I think this often happens in order to be "objective" about the evaluation. I can see how it feels like cheating to coax the model to produce the answer you want. But... it's not! An off-handed prompt isn't more objective than a crafted prompt. You just haven't investigated its biases and flaws.
This lazy assessment is common everywhere, of course. It's one of the reasons bias gets into testing so easily: you setup a test and you assume that it is objective because you give everyone the same test with the same rubric. But if the subjects don't understand your terminology, or the proctor doesn't understand the subjects' terminology, it's easy to mistake misunderstanding for something else (intelligence, opinion, whatever you are testing for).
Systems based on communication need feedback loops, and that's just to get to the _starting point_. Prompt engineering is one of those feedback loops.
[1] https://www.vellum.ai/blog/best-at-text-classification-gemin...