Identifying Stable Diffusion XL 1.0 images from VAE artifacts (2023)
hforsten.com
hforsten.com
It should be possible to identify images that have not been generate by the VAE since they are not part of the set images that the VAE can generate. The other way round is a bit more difficult as there may be images that can be mapped to the latent space and back without loss but have been generated in another way
-> there may be false positives.
Of course, you will need to remove exifs or other metadata, but this sounds like the kind of domain that NNs are good at.
For everyone else though, the people making common tools everyone uses explicitly want their images to be easy to automatically ID as fake. And, the users largely prefer it that way too.
The toolmakers might have a mild incentive to tag images generated with their tool to prevent too easily creating these 'fakes' or just for analytics purposes so they can see how their output is spreading.
But regardless, it's a fools errand like the other poster said. Anyone that is serious about tricking the mass media will strip out the watermark or use a different tool.
It will either match a known VAE well, or be an image that's altered enough (jpeg noise, photoshopping in post, etc) to not match the output of a VAE.
And meanwhile the "loss system" for training AI idea is still hobbled by how the image generating AI systems are all using VAEs. At best, you've swapped one VAE for another. Just another new VAE to add to the list of VAEs to check.