I can't find how many people labeled the DWA task descriptions, where did you got that number?
The article seems to describing the labeling here:
> Human Ratings: We obtained human annotations by applying the rubric to each ONET Detailed Worker Activity (DWA) and a subset of all ONET tasks and then aggregated those DWA and task scores at the task and occupation levels. To ensure the quality of these annotations, the authors personally labeled a large sample of tasks and DWAs and enlisted experienced human annotators who have extensively reviewed GPT outputs as part of OpenAI’s alignment work (Ouyang et al., 2022).
I understand the authors, four, did the initial labeling and then asked an undefined set of people to the rest of the labeling.