This looks very interesting but I don't really understand what he has done here. Can someone explain the process he has gone through in this analysis?
Feeding an empty prompt to a model can be quite revealing on what data it was trained on
>> i sample tokens based on average frequency and prompt with 1 token