High-Resolution 3D Human Digitization
shunsukesaito.github.io
shunsukesaito.github.io
Suppose you tried this on your own camera or maybe in a different lighting: it would totally break this.
Paper definitely seems like an improvement, but I guess no one should get excited to use this in experimentation/production.
https://colab.research.google.com/drive/1GFSsqP2BWz4gtq0e-nk...
[1] http://www.eurecom.fr/en/publication/3189?&theme=mobieurecom
[2] http://www.eurecom.fr/en/publication/3247/copyright?popup=1
That being said, its usually pretty difficult to get code from papers to build/run successfully due to dependencies (e.g. depending on a specific version of Python and OpenCV 2 and requiring CUDA support)
TensorFlow is a crapshoot as far as I understand, constantly changing making newer versions incompatible; but why do other libraries break backwards compatibility without a major version bump?
You might be relying on a function that only exists in 3.7 and not in 3.6, code written in 3.6 would work but new code using 3.7 features won’t be backwards compatible. With compiled code the errors are usually very hard for “more used to scripts“ people to decode. You get stuff like missing symbols in the linker phase.
ML projects usually have a lot of libraries so you also get in the transient dependencies breaking quite often...
Some language-specific package managers can do similar things, but you really only get reproducibility for the whole system with a general-purpose package manager. Poetry gets you pretty far within the Python ecosystem, but if you need a specific version of Python/specific native libraries/etc... it doesn't get you all the way there.
Some gifs of the test results: https://github.com/cardboard-q/pifuhd_demo_model_test
Here is PiFuHD: https://arxiv.org/abs/2004.00452
I guess my idea of "high resolution" differs from what the rest of the world describes as such.
I would have guessed this wouldn't be a problem given that there are multiple photos of the same view, such that you can resolve the depth ambiguity from a single camera. I.e use feature detection to identify feet in > 1 views, and then use the two images to resolve the depth of the epipolar lines that lie on the optical axis and reconstruct the shoe in 3D space.
https://www.reddit.com/r/MachineLearning/comments/8zm4kl/d_l...
Edit: I've been looking at the details on the rear view of some of the "Single-View Reconstruction" examples, and I'm starting to worry that this may actually not be reproducible.
Dr. Iman Sadeghi is the man behind hair rendering tech for Disney and Dreamworks. He left his job at Google to join Hao Li's company Pinscreen (which by the way is funded by big names like Softbank).
When Sadeghi saw red flags inside the company, he raised the issue from within, and finally wanted out. While he was leaving one day, Li and his colleagues legit assaulted (violently) Sadeghi to give up his company laptop. This, by the way, is recorded on CCTV cameras and can be viewed online.
The fraud case is about falsifying results in their SIGGRAPH 2017 Technical Papers submission. They claimed to generate avatar hair shapes in their paper, and when their reviewer asked them to give results on many faces, the company hired artists for as much as $100 to generate them manually for the results. Of course, they later claimed to make it fully automatic AFTER publishing the paper, but that doesn't justify the stance that they published false results in one of the biggest computer graphics conference. Hell even at that public demo, they showed pre-cached avatars and claimed them to be generated real-time.