That's a good observation. Theoretically, 0 strength should give you the content image. What I think is happening is that the algorithm is not trained to generate style vectors for photorealistic images, and so the mapping it learns from pixels to style vector doesn't work well. Maybe a term should be added to the loss function to generate itself when content and style image are equal.