144 karma · joined November 11, 2021
We actually just published a blog today that includes our perspective on building “AI red teams” and best practices for AI alignment/safety: https://www.surgehq.ai/blog/ai-red-teams-for-adversarial-tra...
https://www.lesswrong.com/posts/r99tazGiLgzqFX7ka/playing-wi...
Ha, fair point. I must not realize how old I am, because I was attempting to reference the music of the 1960s and 70s, not 1982, which I agree is not many people's idea of the golden year for music ("Come On Eileen" notwithstanding).
> Sophisticated tools are a bit of a trap. People tend to create in ways that their tools make easier.
No doubt. Ableton, logic, and protools have drastically altered the norms of what modern music is "supposed" to sound like (ie tuned vocals, quantized drums etc). I do wonder what the next generation of music tech will bring.
The modern process of producing music would basically be unrecognizable to anyone 40 years ago — it's completely intertwined with technology, and far more automated. Yet music is as important as ever, and amazing music is being made (will politely side-step the pitfall of debating whether music was better 40 years ago!)
So I'm excited to see how visual artists incorporate tools like Dall-E into their artistic process.
For example:
1. Most data labeling systems don't allow you to communicate meaningfully with your workforce. In contrast, we prize two-way communication; your data labelers are the ones going through tens of thousands examples, so they often have amazing feedback for how to improve your data and design your tasks better. And of course, you often have questions for them as well.
2. Context matters. I can't label Spanish hate speech; someone from Mexico City often can't label Madrid slang either.
3. The majority vote isn't always the best one. Real-world data is often personalized and subjective; your opinion on a funny or angry story may not match mine, and that's okay. Our training sets and AI models should reflect that. We just wrote a blog post on the subtle nuances when considering majority votes and inter-rater reliability metrics: https://www.surgehq.ai/blog/the-pitfalls-of-inter-rater-reli...
4. Curating annotator pools. Our product is designed around helping you build custom labeling teams that you trust, who learn the nuances of your domain and stay with you over time.
Funny to suggest that Spotify may NOT intend to dominate the industry? Surely that's their goal.
For example, imagine you're running a Search Relevance task, where search raters label query/result pairs on a 5-point scale: Very Relevant (+2), Slightly Relevant (+1), Okay (0), Slightly Irrelevant (-1), Very Irrelevant (-2).
Marking "Very Relevant" vs. "Slightly Relevant" isn't a big difference, but "Very Relevant" vs. "Very Irrelevant" is. However, most IRR calculations don't take this kind of ordering into account, so it gets ignored!
Cohen's kappa is a rather simplistic and flawed metric, but a good starting point to understanding interrater reliability metrics.