Interested to know where my assumptions here don't line up with reality.
Interested to know where my assumptions here don't line up with reality.
One possible difficulty I see is still how do we collect the face data / body data. Image generation has a strong bias toward domain -- meaning if we use one kind of faces during training, we will need that same kind of face during inference (with the same angle, lighting, etc..). It's possible but need more thoughts on how to ask users for good image.
I think that's one reason we avoid user uploaded images so far, bc it's hard for users to understand exactly what kind of image we need, and why their images doesn't work well sometimes. There's a lot to explore on this front before we can get a market ready product.
Regardless, cool idea, appreciated:)