If you see Gwerns experiments with GPT-2 you notice that his websites are actually just extremely large samples of text / image data. Essentially the design is that the whole website is practically statically generated.
Refining on AWS was costing me $100 a day.
I also met Shawn (https://github.com/shawwn) who decided to attempt to refine / train large models using TPUs. That sounded interesting but since I figured out that I would waste time trying to understand TPUs because I have a day job I instead just bought a Titan RTX.
My first experiment was that refinement of GPT-2 with your own data would somewhat work. At first I used Donald Trump because he is a very active person on twitter that I believed that people would have some kind of ability of detecting whether the tweets are fake or not.
The above was a bad choice for the following unique reasons. 1. People would generally believe that Trump would say nearly anything. 2. I noticed a weird situation where someone basically duplicated everything that I did had commenters which were apparently much better at figuring out that the person performing the tweets were fake.
Since that experiment sort of failed I created an expanded refinement of GPT-2 on general twitter data. It was 200 MB of tweets from https://www.kaggle.com/kazanova/sentiment140.
I then repeated the experiment and eventually figured out that some people were really bad at figuring out fake tweets and I found a person who used twitter a great deal would actually perform better. I'm not sure if its that you tweet a bunch or just read twitter enough. I had a low sample (n=5) where I carried out the test. So my test could be completely biased.
The test procedure that I followed was to make a set of 10-20 questions and then have the user pick which one was fake and which one wasn't.
I was also doing these experiments on a Titan RTX (I'm thinking about getting a V100 or maybe just another Titan RTX so I can train two models at the same time). I accidentally upgraded the memory to 32GB which worked for a bit but instead you should probably get at least 1 or two multiples of your VRAM so that you can keep your operating system running and having the dataset loaded into memory.
Also, during refinement I don't think my loss ratio was improving. I think that either I wasn't using the system long enough.
But, as a conclusion I figured out that there is a huge shortcut that I never considered. It would be MUCH easier to just take any refined GPT-2 model or even use GPT-3 as an API above and then make it LOOK like a tweet. Just adding a hashtag and a t.co link would work. (the funny part is that GPT-2 actually seems to have some notion of a t.co link and will happly generate t.co links that don't work. Removing t.co links before refinement would be one way to get this out.
I did some of these experiments using Google CoLab initially. But as I used it I got out of memory options. I asked someone at pyOhio who worked for google if there was a way to connect Google Colab to a paid instance. They responded to refer me to Jake Vanderplas https://twitter.com/jakevdp . They said no at the time. Then a few months after Google came out with Google Colab Pro. But then I was able to run out of system memory in the high memory instances. I upgraded my Deep Learning Rig with a AMD 3600 and I am now waiting for my 64 GB memory that is in the mail.
The latest thing that I have done is use local voice synthesis and recognition so that you can talk to GPT-2 locally. My tutorial is at https://www.youtube.com/watch?v=d6Lset0RFAw&t=2s