How was ground truth obtained if not via human annotation?
For each of these datasets, we specify task guidelines/prompts for the LLM and human annotators, and compare each of their performance against ground truth labels.
Fixed it for you.
Is there some noise in these labels? Sure! But the relative performance with respect to these is still a valid evaluation