12 karma · joined October 18, 2021
This is correct. We did almost all this work on a Macbook Pro. Although for the pile-10k dataset analysis we used an A100 GPU because it would take many hours to run the whole thing through GPT-2 on a laptop.
1. It's kinda interesting because this is a clear case where the model must be thinking beyond the next token, whereas in most contexts it's hard to say whether the model thinks ahead at all (although I would guess that it does most of the time).
2. More importantly, the key question here is how it works. We're not surprised that it has this behavior, but we want to understand which exact weights and biases in the network are responsible.
Note also that this is just the introductory sentence and the rest of the article would read exactly the same without it.
> it doesn't seem that surprising that there might be cases where there is a dominant bigram [...] that would trigger an "an" prediction, without actually predicting the following word first
btw I don't really understand what you mean by this. Bigrams can explain the second prediction in a two word pair but not the first.