Fine-Tuning GPT-2 from Human Preferences
openai.com
openai.com
(Someone didn't read the whole article..)
Though I hesitate to draw strong conclusions from fiction.
I'm not certain how random users feel when knowing their selection knowing it's being used to improve an algorithm. In this case, it's relatively easy for me to log the prompt and the option that was selected -- just doing it felt a little ... bad, and I'm a little too scared of GDPR for a fun project.
The other thing is humans sometimes select the funnier option even though it may not necessarily be the best one. In a Show Reddit post, the most upvoted response is a Game of Throne's character Tyrion and a brothel story.
https://www.reddit.com/r/FanFiction/comments/d5s9yh/i_made_a...
Open-sourced Code: https://github.com/jeffshek/writeup-frontend
One hack is for the medium-level models, you can actually run them on Cascade Lake (which are sort-of more optimized for ML) than traditional processors. There's a 30% performance there. Mathematically, GPU vs CPU in performance at inference time, you're paying an annoying premium!
Right now, the default writing style "medium" is running in Cascade Lake (no gpu). I also (over)optimized the microservices running the endpoints too.
> Note that we provide pre-trained models, so you can skip directly to RL fine-tuning or even to sampling from a trained policy, if desired.