But the naive approach of having a table of how much each individual training item influenced every weight in the model seems impossibly big. For DALL-E 2's 6.5B parameters and 650m training items, that's 4.2 quadrillion associations. And then you have to figure out which weights contributed the most to an output.
I would love to see any research or even just thinking that anyone's done on this topic. It seems like it will be important in the future, but it also seems like a crazy difficult scale problem as models get bigger.
> How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction.
If I include the tag "floor", do I get some (tiny) percentage of every image that uses "floor" in the prompt, even if the bits from my image did not end up affecting model weights much at all in training?
Worse, for tags like "dramatic lighting", it's likely that the important source images will depend on the other words in the prompt; "sunset, dramatic lighting" will probably not use the rely on the same weights or source images as "theater interior, dramatic lighting".
And then you get the perverse incentives to tag every image with every possible tag :)
I'd love to be convinced otherwise, but I'm not seeing prompt-to-tag association working.
Why do you have to, though? What do you hope to trace back to exactly?
(Also, it's entirely possible that eg a model could generate images resembling your work without "seeing" any of your work and only reading a museum website describing some of it. Resemblance is in the eye of the beholder.)