It uses their small model and a tiny dataset in comparison (and a small amount of training). It is more showcasing how much it learns (and doesn't learn) with those limitations in place. As well as allowing you to recreate it with perhaps a few minutes of work and less than an hour of waiting.
Also, I wouldn't say the results are nonsensical - I think it has learned a lot more than a markov chain or a simple rnn but I agree that especially on the surface they dont even sound like they surpass Eliza by much. Moreover, it is significantly more apparent how much it learns about the different people you've talked to AFTER you run it on your own data.
For a somewhat more novel/interesting result with fine-tuning GPT, I can recommend checking out gwern's post[1] on training it on a big poetry corpus.
1. https://www.gwern.net/RNN-metadata#finetuning-the-gpt-2-smal...