Image-to-Image Translation with Conditional Adversarial Nets
phillipi.github.io
phillipi.github.io
Even though the sketches are fairly crude, with no shading and a low level of detail, many of the generated images look like they could, in fact, be real handbags. They still have the mark of a generated image (e.g. weird mottling) but they're totally recognizable as the thing they're meant to be.
The "sketches to shoes" example, on the other hand, reveals some of the limitations. Most of the sketches use poor perspective, so they wouldn't match up well with edges detected from an actual image of a shoe. Our brains can "get the gist" of the sketches and perform some perspective translation, but the algorithm doesn't appear to perform any translation of the input (e.g. "here's a sketch that appears to represent a shoe, here's what a shoe is actually shaped like, let's fit to that shape before going any further"), so you end up with images where a shoe-like texture is applied to something that doesn't look convincingly like a real shoe.
https://phillipi.github.io/pix2pix/images/index_facades2_los...
Notice white triangles (image crop artifacts) present on the original image, yet completely absent on the net input image. They make re-appearance on the output of 3 (4 even?) out of 5 nets despite the lack of corresponding cue in the input image. Looks like network cheated a bit here, i.e. took advantage of small set size and memorized the input image as a whole. Then recognized and recalled this very image (already seen during training) rather than actually reconstructing it purely from the input.
Same (but less prominent) for other images where "ground truth" image was cropped.
What I like about the "Day to Night" example is that is clearly demonstrates that these sort of networks lack common sense. It expects light to be where they are clearly (to humans with common sense at least) no things that can produce light. E.g. in the middle of a roof or in a tree. Of course, there can be, but it's fairly uncommon.
And the opposite as well, no lights where a human would totally expect a light, eg. in the front of buildings or on the top of, well, lighting poles.
I suspect a neural network better specialized for this task (i.e. that has the data interlaced for both day and nighttime during training) would have no problem feature detecting trees and leaving them unlit.
Makes me wonder how this can apply to image and video compression. You could send over the semantic segmentation version of an image or video, and system on the other end would use these technique to reconstruct the original.
There are even more traditional tricks that don't make it in things like H.265 because it is too costly.
From that, the AI could generate books, movies, and do a lot of things.
Could this change, if, for example
- inputs are augmented with the network state (or derived version thereof)
- previous outputs of the network / external memory are fed back?
This seems to be the kind of self reference self awareness requires.
Also, do asynchronous networks have fundamental advantages over synchronous networks? What about static vs dynamic networks?
Does anyone have any experience in this area?
You can pipe these product sketches directly into focus groups who tell you which product is most likely to sell. You don't need massive staff to come up with product variants any more.
Perhaps what these networks are generating can be labeled better as "Guided/constrained imitation" rather than real creativity.
What is real creativity? Creativity is just random noise converted into patterns. Is the computer variety of creativity not real enough?
This is not a consensus definition. Creativity doesn't actually seem to be very random at all according to the people who study it.
Humans are not magically creative as much as they'd like to be.
If i ask you to think of a random number, you don't just pull it out of thin air, It can be based on tens to hundres of things: -Should i do a relaly low or high number? -People always use round numbers that end in 0 or 5, maybe i shouldn't do that, or should i to make it seem truer -what other large "random" numbers have a heard? -i remember seeing a number recently, maybe try a modification of that -you used {x} as a random number last time, go similar to that?
All this adds up in that under a second thought you have when i asked you to think of a random number. the literal same thing goes into all creative works, the output is a function of the input.
You're acting like the brain is some kind of simple algorithm, more goes into a painting or a composition than just a bit of simple logic. A composer is not sitting there at 2am on her Piano going "Hm, I like round numbers, so I might make this note a F because it's the fourth note in the C major scale".
I might be wrong but from my understanding, we don't even really understand how neural nets are able to make certain decisions or generate certain pictures yet, correct?
It's not binary but it is a computer. The alternative is to believe in magic.
> we don't even really understand how a computer is able to make certain decisions or generate certain pictures yet, correct?
I don't think that's correct, no. We understand how the process works. We may not understand the weights a specific neural net ends up with, but that's an issue with just having so much data to deal with. Similarly we don't "understand" how a web page ends up with a specific PageRank. We understand the process, but we can't manually reproduce the result because it's just too much data.
Randomness is injected into all brain processes on account that biological neurons are stochastic. So there is an amount of randomness mixed into everything the brain does.
Some neural nets can map real images into a Gaussian, and back. That means they disentangle the factors of the image into a mix of independent factors that map into the standard deviation. Any set of random numbers could be converted back into an image, by the reverse process.
On what basis do you make this claim? Humans are empirically terrible random number generators. If you ask someone to pick a random number, the result is very not random. Our biases are large and obvious, so it seems faulty to claim that our "seed" number is in any way truly random.
> Randomness is injected into all brain processes on account that biological neurons are stochastic. So there is an amount of randomness mixed into everything the brain does.
There's also some amount of randomness in what happens if you drop a rock but the net result is largely the same: it falls down. The fact that there is some randomness to a process does not mean that the randomness is actually driving the process.
> Some neural nets can map real images into a Gaussian, and back. That means they disentangle the factors of the image into a mix of independent factors that map into the standard deviation. Any set of random numbers could be converted back into an image, by the reverse process.
I don't see how this is relevant.
While this is a highly philosophical topic, I can say, on the difficulty of defining creativity from a scientific point of view, is that more and more what I think is considered "true creativity" is fundamentally a social concept, and as such, even though computers are good at reproducing patters and variations on patterns, they will never be "truly creative" until they become independent members of society, whatever that means. So as cool as machine learning is getting these days, this would require a leap in AI that is still rather far away (imho).
That's why I think it's a little optimistic to label something as complex as creativity as being "just" anything. Creativity is social, insofar as it is defined and recognized by human society, and machines are.. not. Yet.
It is the interaction between the relations and components that drives the kind of analogical reasoning humans exhibit: traversing networks of relations and thinking up new ways of combining various components.
Without a large base of knowledge to draw from, a creator simply does not have enough data to do anything more than copying. They would not have learned the rich relations nor a diverse enough set of components and operations on them. They would also be unable to know which combinations are truly novel nor have enough experience to predict what might be meaningful or impactful. This is why good creators must continually sample a large set of works.
The limit to this is that humans will generally struggle to wander far from experience. Computers on the other hand, can cover more of a space to make connections humans would find difficult. And the less random the decisions, the more interesting the founds structures are likely to be.
Where computers struggle and humans shine is the selection process. Ashby defines selection as: "problem solving" is largely, perhaps entirely, a matter of appropriate selection. Take, for instance, any popular book of problems and puzzles. Almost every one can be reduced to the form: out of a certain set, indicate one element".
In this work for example, the excellent results are a better grasp of structure induced by having the GAN's networks condition on an input image. This architecture does have the downside of essentially limiting the generative capacity. Humans can generate many variations conditioned on an input but these networks: "Despite the dropout noise, we observe very minor stochasticity in the output of our nets. ...GANs that produce stochastic output, and... capture the full entropy of the conditional distributions they model, is an important question left open by the present work." AGI more and more looks like it will be about striking the right balance between generation and selection.
For an artist, selection ability also plays a large part in taste. Having not just a good enough understanding to create novel combinations, but also a good enough model of people to predict what they are likely to find somehow compelling.
Computers are good at generating, humans at selection. It is for this reason I disagree with those who believe tools like these herald the end of creativity nor can I agree with those who believe this heralds the end of creators. The quality floor of average art will rise but there will be less evidence of selection talent as the ceiling of what counts as great art must also rise.
Joule for Joule, it is unlikely that anything for the forseeable future will beat teams of machine and man.
> AGI more and more looks like it will be about striking the right balance between generation and selection.
I was thinking AGI is essentially reinforcement learning on top of rich, predictive models of the world. But you could see it as the balance between generation and selection.
I see lots of papers that go in this direction, of creating a rich, semantic, predictive representation of images, video and text and then using it as the basis for reinforcement learning. Learning to understand the world and to act based on that understanding.
...
I get a feeling this could be used in game design to do some really cool stuff with map and texture generation.
We've got the pieces of visual processing and imagination here and the pieces of language input/output as part of Google's work. It feels like we just need to make some progress on an "AI executive" before we can get a real, interactive, human-like machine.