This beach does not exist
thisbeachdoesnotexist.com
thisbeachdoesnotexist.com
The network did bad where it had to take into account global illumination, or distance to the object and its level of detail: 1) tree trunks are mostly flat and do not depend on the possible sun location; 2) shadows don't match the trees and the rocks; 3) nearby rocks often lack texture and are not properly illuminated (they look flat).
It's not surprising given that the method is fundamentally local (CNN-based). I think better results will be achieved when a generator creates a scene graph rather than a raster. Are there works which actually create a 3D scene (or use it as an internal representation)?
[1] https://thisbeachdoesnotexist.com/data/seeds-075/4375.jpg
[2] https://thisbeachdoesnotexist.com/data/seeds-075/3300.jpg
[3] https://thisbeachdoesnotexist.com/data/seeds-075/7638.jpg
Actually the original “this person does not exist” wasn’t so bad, but things like this, DALL-E, or the sheep one that was posted the other day I find extremely unsettling and can’t explain why.
Anyways, if you were raised in a machine with these images and then shown the real world, I wonder if it would be a relief or also uncanny?
I know that when it comes to sound, it's really hard to create organic sounds. And that part of our ability to hear sounds happens (theoretically) before the brain, in the form of grouping harmonic series together into isolated sound sources. Perhaps eyes do something similar? Some sort of pre-processing that functionally works for all natural phenomena, kind of like the harmonic series does for sound.
I thought the latest thinking is that most of what we "see" is really a projection from our own mind (similar to how these images are generated) and our attention is drawn to those areas where our projection and sensory input don't match. So yes, we are evolved to be alerted when "something feels wrong" because it's unexpected and could be a danger.
The mismatch can happen at any level of the hierarchy. One person here said some waves appeared to be moving the wrong way for the rest of a scene. Things like that are at a higher level of abstraction than rock textures or tree trunks with gaps. Some things might seem wrong but take more time to identify or explain what the problem is.
It looks nice, but I don't feel safe using untrusted pickle weights of pretrained models, as they can allow for arbitrary code execution.
I fully understand your concerns, but I don't know how to guarantee that the pkl is ok. Don't run it on your computer, but in some isolated environment, like Colab.
I'd argue that running untrusted code in a Colab is even worse, as you'd risk your account instead of just your computer.
I don't think pickle files can be loaded safely, it's better to use a numpy archive npz to store the weights which can be loaded without a security risk by using allow_pickle=false when loading (the default since numpy 1.16.3)
1. Express each pixel as a vector of numbers, probably with RGB.
2. For each of the color vectors, for each of the colors, subtract the color 1 from color 2.
3. Square those differences, and add them.
In this way you can get mean squared error for each pixel or for the entire image.
recursion paradox joke :)
Does this artifacting have a name?
1. https://arxiv.org/abs/2106.12423 (project page: https://nvlabs.github.io/alias-free-gan/)
In theory you could train a generator on stock footage. In theory the stock footage copyright holder could sue you for creating a derivative work without paying them. In theory that claim could be bunkus if the creator used a different training set - how do you prove a generated image's origins?
We probably need ThisMonumentDoesntExist.com for that.
Anyway, to grab all these beach images, I used this command:
wget -w 3 --random-wait https://thisbeachdoesnotexist.com/data/seeds-075/{1..9999}.jpg
Average file size is 130KB, total 1.3GB.Plenty of photographs have crisp object definitions in the real world. These do not. Trees blend into humans standing on the beach.
Only a small feature request: When you click the button to get another set of random images or another cluster in the knn, the page waits until the new image is loaded to update it, so it "does nothing" for half a second. I'd like that it instantly shows a spinning thing or other signal to show that it's waiting for info and will be updated soon.
Maybe there's a brief moment of "I'm impressed with how much work you put into this" but for that to happen, you kind of have to see the work being done.
I think that's why impressionism has such staying power, impressionists realized that your brain doesn't need photorealism. They were panned by their contemporaries who were focused on photorealism, but time has proven they were onto something.
Even now the impressionist-style neural nets don't get it right.
Or be knowledgeable enough to estimate the effort and skill required. Ironically, things I took for granted when I was young have impressed me more and more over the years. Probably because I've either tried to do them myself or learned about their backgrounds.
So I'm not surprised artists and at enthusiasts would value realism more than most people. Just like how speed metal is probably more popular among guitarists. The technique becomes a source of value rather than just an implementation detail.