I work at Meta. Scale has given us atrocious data so many times, and gotten caught sending us ChatGPT generated data multiple times. Even the WSJ knew: https://www.wsj.com/tech/ai/alexandr-wang-scale-ai-d7c6efd7 https://archive.is/5mbdH
We intentionally didn’t use them at all for Llama 2 and mostly avoided using them for Llama 3, but execs kept pushing Scale on us. Total mystery why until now, guess this explains it.