Ok, so let's assume it's using a Visual Transformer architecture, since AFAICT, we don't know.
What the model does is turn the picture into numbers in a clever way (embedding).
Then, the model compares the relationships between all the embedded tokens (self-attention).
Then the model does the fancy linear algebra stuff we've come to know and love (feed forward neural network).
What you're basically getting at with your temperature comment is "rather than having it semi-randomly pick one of the top n options that the feed-forward network recommends for which pixel goes where, we can always have the model simply choose the top answer each time."
That is not "we can make it reproduce training images"
The point of embedding is to see multiple examples of similar images and generalize from them. The process in and of itself breaks the ability to re-create a photo, unless the training set is very skewed.
If that's the argument of legal cases, they are likely to fail or win on ignorance.