The model knows the original context too, though. So it doesn't make a lot of practical difference if it found it through a secondary source.
But the discussion has mostly centered around whether Anthropic have properly scrubbed benchmark data out of their training corpus before training. So it seems to me that even if they did it properly (maybe especially so), the presence of the canary string in Claude's output is likely if it saw it in a secondary source (e.g. a blog post), which should then be fair game, no?
It makes literally all the difference that matters?