Might be a result of using LLMs to evaluate the output of other LLMs.
LLMs probably get higher scores if they explicitly state that they are following instructions...
LLMs probably get higher scores if they explicitly state that they are following instructions...