Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing
phorhum.github.io
phorhum.github.io
I've been primarily tackling the synthesis from the GPT-3 + 3D animation side but the 2d prerender idea is genius. I'll need to try to hook it up one of these weekends. The only hard part, it seems, is finding the navmesh.
To me, that means a version DALL-E which can do art direction without me in the loop at all is not too far off.
If you're not recreating something from actual data elsewhere, it's just being made up on intuition level guessing.
My concern is what the purpose of this "tech" is being used for, and more specifically, how this "tech" will be mis-used for other purposes.
I guess, by hedging this way, it managed to improve its evaluation score without having to learn how clothes actually work.
If you can take a photo (effort required: seconds) and turn it into a textures low- to mid-poly model within seconds, that's a huge time saver right there.
Combined with other techniques [1], [2] this could indeed present a viable business model - again for certain niches. It's the same old routine: identify potential, present solution, profit (if only for a limited time). On its own this might not look too useful/impressive, but I'm certain there's people out there who could already benefit greatly from this work.
It's not always about "the next big thing" or "turning an entire industry on its head". I could see this work its way into various products without much fanfare and gazillion dollar valuations.
[1] https://github.com/Shimingyi/MotioNet
[2] https://deepai.org/publication/end-to-end-learning-for-3d-fa...
Example? A fashion billboard that doesn't use model photos but renders passers-by in fancy new clothes. Of course, the gathered data gets syphoned of into a database somewhere.
Smartphones and CCTV cameras are much more effective sources of information than a smart billboard anyway. Smart billboards, however, sound quite awesome.
Computing and related tech is an enabler of so many new possible applications. It has become really tough to develop something useful for consumers that can be commercially successful, but doesn't at least have the potential to be turned into something creepy.
Corporations have an increased ability to do things w/o oversight, but the government can literally do whatever it wants when it wants as long as those in power allow it. Government has the monopoly on power and violence. This concern is increased even more in democracies where there is great division. One side can abuse government to enforce their aims on an unwilling opposing side.
Additionally, corporations generally just want to make money so they are fairly predictable and easy to contain (i.e take your business somewhere else if you don't want to deal with them).
Government operates at the beck and call of political whims which, let's face it, are more often than not emotional and knee-jerk reactions. Lastly, only government can make it illegal to avoid surveillance and put you in jail for it. The worst a corp can do is close your account or something.
Anyway, I think this is getting too far from the actual topic here. I'll leave it at that.
Although, they are getting better years after years, so maybe we'll be out of the uncanny valley at some point ?
On one hand you have traditional photogrammetry based on classical image descriptors (from the beginning of the 2000s). A nice open source solution is Meshroom [1]. I have very little experience as I have just tinkered a little but I would say they work OK with geometry of medium complexity and detailed textures. They fail horribly with untextured objects, like for example a candle. They are semi-automatic: you upload some photos and run the pipepiline but for best results it's easy to tweak as there are a lot of knobs to try.
On the other hand you have these deep learning research papers. A lot of them have Colab notebooks you can try yourself (which is awesome, thank you). They can give you these very nice demos for some kind of objects (the ones we have a lot of training data, like human bodies), but they are not truly general and sometimes will generate artifacts. There is some opportunity here if they are integrated in a semi-supervised workflow, like this [2].
Anyway, just curious and not an expert.
[1] https://alicevision.org/ [2] https://keentools.io/products/facebuilder-for-blender
Not exactly the same thing but I have been thinking about the idea of 3D construction of cloth/jeans you want to purchase onto your own body, so that you have better idea before placing order, which could be really useful for online shopping.
I was thinking in game graphic way (bottom-up), such as building 3D model of myself, shopping website providing 3D models of their clothes, and then it would be a "fitting" problem to properly place clothes onto the body.
From this paper it seems that it is not necessary to have low-level data to achieve the objective.
Their original idea was even better, IMO. It was supposed to send the model to an on-demand automated clothing factory that would custom-sew entire pieces of clothing to your measurements.
I think the technology wasn't quite there yet at the time. Even with four first-gen Kinects, the body model was probably too inaccurate to compare to an actual tailor. They seem to have moved away from clothing in general since then.
For gamedev you often have a drawing of a Character made up which then gets scuplted, retopo, texture ...etc. This is very time consuming especially for characters that are in the background. If you could create a bunch of characters from drawings that would make large crowd scenes much easier to create.
For archviz you don't need an accurate 3D model but one that looks good and casts shadows.
I'd imagine smart mirrors in clothes stores as a much more intriguing application.
A image-to-skeleton pipeline might even be useful for 2D sprite rigging.
This is actually technology we already have today. You wouldn't want X-ray goggles, but infrared "goggles" would do the job just fine. Like X-rays, infrared light will pass right through clothes; unlike X-rays, it won't pass through skin.
Hook up an IR camera to a VR headset and you're there.
If an ai can reconstruct a view of the far side of something from data like the near side view and the shadows & other effects on the surrounding environment, then it can do the same thing to reconstruct a view of what's underneath clothing, theoretically, which may be reasonably restated as eventually.
This is actually not something that people want. The body within clothes is generally constrained by them in ways that look odd if you can't see the clothes; what people really want is to see what someone would look like if they were naked, not what they would look like if their clothes were invisible.
Would be super creepy though.
You're not thinking evil enough. Provide app for free, then use face recognition to allow people to pay to look better naked through the app.