When I see tests like this, I have no idea what I am even supposed to expect. Should the model do what the pretraining examples show in aggregate? Is it supposed to follow some post-training RLHF? Is it supposed to do exactly what the prompt asked it to do?
What is it even supposed to "align" to when the above are in conflict? No matter what it does, someone can construct a case where it fails.