Interesting, but couldn't a model "cheat" in this task by being very good at telling model outputs apart? How far do you get with a classifier simply trained to distinguish models by their output?
It seems to me many models - maybe by design - have a recognizable style which would be much easier to detect than evaluating the factual quality of answers.