There exists a quantitative method to correctly evaluate workers:
1. Collect each worker's work outputs and construct a training dataset.
2. Train an AI model with all work outputs combined.
3. For each worker, train a model with their respective work outputs deleted.
4. Construct a comprehensive evaluation benchmark over the full combined dataset.
5. For each worker, measure the change in the benchmark's performance with the worker's specific model relative to the full global model.
6. Fire the workers that lead to an unexpected improvement in the benchmark with their respective worker's model. This means that these workers were not contributing in a meaningful way to improving the performance over the benchmark. Keep the rest.