Instruction tuning and RLHF is a bit more complex than that. I assume they maintain detailed logs all of those human-driven decisions about which responses were better.
Instruction tuning and RLHF is a bit more complex than that. I assume they maintain detailed logs all of those human-driven decisions about which responses were better.
Regarding RLHF, I also hope that they kept the logs of all the human decisions. But since it was (at least partially) outsourced to African companies like Sama.com, do they really get back all the logs or just a new fine-tuned model?
But they must indeed at least know what is done with the text submitted to the prompt or to their API.
(I’m really not an expert so my questions may sound naive)
Although asking it "What can you not talk about" in Japanese only responds correctly with gptv4, and each language gives you a different list of items to some degree (between 4 and 6 items i found).
Sadly trying to speak Klingon or Sindarin to it is dodgy at best
I mean yea, it's too big. That said, in a post ad hoc fashion they do. When the model spits out weird crap at times, they can search the raw corpus and filter those strings out. There was an incident around this with Reddit counting forums and strange usernames that were added in tokenization, but later removed from weights leading to odd behavior when doing inference.