GPT trainer says he's traumatized from the RLHF work
bigtechnology.com
bigtechnology.com
He was traumatised by millions of unfiltered thoughts shared by millions of folks.
He was traumatised by the garbage, racism, homophobia, child abuse, hatred and biases that some harbour and release when unsupervised.
We focus too much on advancing tech but too little on advancing humanity.
The people meant to advance our way of thinking are marginalised, underpaid, and even threatened by replacement with the very tools that ingested our darkest thoughts.
Are we sure that is the future we want?
Even law enforcement don't visually classify child porn anymore; they just look at file hashes and obfuscated thumbnails. You could randomly redact half the words in any sentence and ask "does this look like a message you would send to your mother?", then compare results across 3 reviewers. Or do a sentiment-analysis first-pass and make sure at least 75% of what these people read is positive.
Humanity has advanced if poor Africans only have to read about this stuff now. In the last century, his Liberian neighbors were having the acts described actually happen to them.
That's effectively what we currently have. No one is using threats of violence to coerce people to read it, they're forced by economic circumstances as are MTurk workers.
There's also an underlying issue that the people most likely to be interested in volunteering for this are people who enjoy the content. It becomes a fine line between moderating CSAM and snuff scripts, and being a distributor of that content.
> The job should be crowdsourced, not concentrated. We can all survive the occasional slip-up and report it as needed.
That's going to be fundamentally incompatible with a lot of businesses, and especially advertising. I'd imagine Microsoft would be pretty uncomfortable with using ChatGPT if it came with a "this thing will probably regurgitate CSAM at some point, just have your users report it".
> You could randomly redact half the words in any sentence and ask "does this look like a message you would send to your mother?", then compare results across 3 reviewers.
I don't think you can reduce the fidelity of ideas in the same way you can reduce the fidelity of an image. E.g. that sentence, skipping every other word, becomes:
You randomly half words any and "does look a you send your?", then results 3.
It's just gibberish. Cutting the number of pixels in an image in half is still the same image, but harder to make out. Cutting out half the words of an idea makes it not an idea any more.
You can water down a sentence and the underlying sentiment, but doing so would require understanding the underlying sentiment. At that point, we wouldn't need humans to read it at all.
Crowdsourcing this stuff could work, but I think it would require substantial shifts in current perceptions. I don't think most people would tolerate having CSAM and snuff scripts show up in their feeds with any degree of frequency.
A hard answer is to maintain a community properly to ensure self-policing with good backup moderation, like they do here at HN. There are only two (?) moderators but the culture here is very well maintained to self police threads.
> OpenAI told me it believed it was paying its Sama contractors $12.50 per hour, but Mathenge says he and his colleagues earned approximately $1 per hour, and sometimes less.
> OpenAI knew these workers were supposed to get routine counseling, but Okinyi and Mathenge found it insufficient. “At some point, the counselor reported,” Mathenge said, “but you could tell he was not professional. He was not qualified, I’m sorry to say. Asking basic questions like ‘What is your name?’ and ‘How do you find your work?’”
The final outcome when OpenAI pressed for more details about working conditions was that this outsourcing company simply exited content moderation entirely. Hopefully the folks running this program on the OpenAI side are revisiting how they choose and manage their vendors.
Will this be easier when mature or when its young?
I strongly suspect the iterative training of successful AI's will resemble how we raise children.
We got GPT 3.5 in part by training models on RLHF data, not just to answer questions, but to rate questions like the humans were.
And in turn we've seen those models used to fine tune models like LLAMA without requiring anyone to sift through that.
Still this is a horrible injustice, OpenAI is going to take on unimaginable wealth and even if that wasn't guaranteed when these people were being paid $1 an hour, it's not too late to try and fix some of what they did.
North American companies don't like to do that legwork. They like to outsource every little thing, because that's how you minimize costs. You don't send your employees to Africa, that's risky. You don't pay North Americans $12/hr, either, because you don't want the hassle of directly employing people. Workers have rights here and if they get paid $1/hr they get uppity.
This isn't an Africa thing, it's typical third-world grift. What they're saying holds true anywhere, even with immigrant communities in America.
You offer X Solutions $100 to do a job (what a steal!). X pockets $60 and solicits Y Recruitment for $40. Y pockets $30 and hires 10x Z Workers for $1 each. It's middlemen all the way down.
>To make our models safer, more helpful, and more aligned, we use an existing technique called reinforcement learning from human feedback (RLHF). On prompts submitted by our customers to the API, our labelers provide demonstrations of the desired model behavior, and rank several outputs from our models. We then use this data to fine-tune GPT-3.
Were there a bunch of user-submitted horrific prompts in the dataset? Was raw GPT-3 producing graphic text in response to benign prompts? Was OpenAI adding a bunch of graphic prompts into the dataset?
^ they were reading and labeling explicit prompts provided by users. And explicit responses by the model to those explicit prompts.
I presume people tried that with chatgpt as there is demand for such content.
Then there’s also the training content. All forums and online platforms have some shape or form of abusive content. Filtering that can be very traumatising because it reveals what some among us think and do. Such content is not rare - it sometimes surfaces on this forum as well but to a lesser extent.
The frequency of such content is a good indication of what prompts people might be coming with.