Learning to write programs that generate images
deepmind.com
deepmind.com
But man does this scare me. I remember a quote: “You furnish the pictures and I’ll furnish the war.” – William Randolph Hearst, January 25, 1898. This was in the lead up to the spanish american war.
Can you imagine tech like this/tech like "deepfakes" being used today? Fake news that was text alone has done and is doing damage in elections around the world now. Imagine that armed with pictures?!?!
In a dueling NN architecture, many say the discriminator will be able to detect the fake images. I wonder is there a threshold that a produced image is just too damn close to a real picture that even an equally good NN that is discriminating can no longer differentiate? In the end, both real and fake images are just pixel values... what would we do then?
Cool tech, scary possibilities.
Will be interesting if it ever reaches the point where it can be automated and scaled. I predict a modernized repeat of the war of the worlds tipping point followed by ratcheted skepticism for all kinds of temporal simulacra.
Then 3d imaging will enter widespread consumer usage and prove to be very difficult to convincingly reproduce by neural networks, until it is. Trust will be restored in some kind of media until it's broken. Rinse and repeat.
[0]https://www.youtube.com/watch?v=ohmajJTcpNk
[1]https://research.googleblog.com/2018/03/expressive-speech-sy...
The point of OP is that you're learning to generate little images in a much more difficult setup: you have to control some complex blackbox system (like a paintbrush robot) to try to generate an image with only crude success/failure feedback at the end of the sequence of actions. The hope is that by going through this intermediate environment, instead of generating an entire image in a single shot via convolutions, it'll be learning more abstract structure about what makes up a face etc, so hypothetically it could do things like rotate faces in 3D (whereas something like ProGAN only sees faces as 2D blobs so it can do things like add/subtract sunglasses or change hair color, but 3D transformations are beyond it). And, with this more abstract deeper understanding, it should be able to speed up learning in settings like robotics etc (instead of paintings, imagine they are videos of humans controlling pick-and-place robot arms); you can see this as one way of approaching unsupervised learning and providing primitives which a higher-level agent can learn faster from (somewhat like the GAIL architecture uses GANs for semi-supervised learning).
Is it possible to (a) filter out these duplicate strokes, (b) convert them to heavier-weight single strokes, or (c) change the training regime to not produce duplicate strokes?
I can see that being useful for e.g. a real robot with a limited amount of ink or lead (or time to draw each character).
I'm starting to understand how Juergen Schmidhuber feels.
That's interesting. Do we know how artists draw? Is it as "algorithmic" as the article lays it out? I don't draw so I always assumed it was more intuitive and personal rather than a "step by step" process.
http://www.dailymotion.com/video/x65w5fu
It is eye-opening, even among fellow manga artists, to see how different sometimes their processes are.
Some may start with a definite sketch, others may go straight to ink with only the barest suggestion of a layout. Sometimes they struggle with expressions and may whiteout and re-ink (up to seven times in one of the videos.)
Some artists start inking with the eyes, some may start with an outline of the face. And so on.
What any specific artist uses will vary greatly. But it usually falls into one of those three camps.
But depicting those traits is another matter. You can render a chin meeting the hair in all sorts of ways; but your choices are limited to your aesthetic preferences, and your ability to draw that form.
Drawing is a highly mechanical process; choosing what/how to draw is a curated one.
A good striking example, do a video search for "two point perspective drawing", and look at some of the tutorials / demonstrations that come up.
See my results here: https://forwardscattering.org/post/42
Fast painting is a benefit I guess. My search/painting program is very computationally intensive.
Edit: I think I see the point of the paper now. Unguided search is going to be difficult in high-dimensional search spaces like this. So the NNs become a hopefully-effective heuristic guiding the search.
The (semi-) obvious next step is to do object/digit recognition with a Bayesian probability calculation, with probabilities bases on this image reconstruction process. In other words, we choose e.g. digits based on how likely they are to have been drawn to give the target image.
I have experimented a little with this idea, but with no successful results so far (plain old NNs still beat it).