FWIW even more recently, models have been tuned using a method called DPO instead of RLHF.
IIRC DPO doesn’t have human feedback in the loop
IIRC DPO doesn’t have human feedback in the loop
Does this have a specific name?
Conversations you have with like chatgpt are likely stored, then sorted through somehow, then added to an ever growing dataset of conversations that would be used to train entirely new models.