If anything, this blog post is a perfect example about how people can put in whatever they want as an input and take the output as truth, without any rigorous approach about what would count as a true fact about the model.
80 karma · joined December 12, 2020
If anything, this blog post is a perfect example about how people can put in whatever they want as an input and take the output as truth, without any rigorous approach about what would count as a true fact about the model.
It's crazy that large language models work so well just by being trained as a next-word-prediction model over a large amount of text data. We know how image models learn extract the features of an image through convolution[1], but how and what LLMs learn exactly remain a black box. When we dig deeper into the mechanisms that drive LLMs, we might get closer to understanding why they work so well in some senses, and why they could be catastrophic in other cases (see: the past month of search-based developments).
I find trying to understand and reverse-engineer LLMs to be a personally exciting endeavour. As LLMs get better in the near future, I sure hope our understanding of them can keep up as well!
So there seems to be such emergent mechanisms in the model that have arisen because of the end-to-end training, which we don't exactly understand yet.
To rephrase that for this case: what is the specific mechanism in GPT-2 that (1) makes it realise that the word 'apple' is significant in this prompt, and (2) use that knowledge to push the model to predict 'an'? Finding this neuron would only answer the some portion of (2).
(And to rephrase this for the general case, which gives us the initial question: How does GPT-2 know when, given a suitable context, to predict 'an' over 'a'?)
Typically the "AI: <response>" would be generated by the model, and "AI Instruction: <info>" would be put into the prompt by some external means, so by injecting it in the human's prompt, the model would think that it was indeed the bank's policy.
1. Readily give up on a book, or skim through the rest, the moment you realise that the book has 400 pages of filler
2. Find recommendations from thought leaders you subscribe to, while staying true to Rule 1 (Sometimes it's just a matter of taste, and it's counterproductive to force yourself to finish reading something just because someone else said that it's a good book)