One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing
nvidia-research-mingyuliu.com
nvidia-research-mingyuliu.com
Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.
That defeats the entire purpose of using facial and body expressions that only video provides.
We already have video filters that remove wrinkles and blemishes in videoconferencing to make you look better.
Even if we replace ourselves entirely with computer-rendered images, they're still going to be reproducing our expressions, movements and gestures, which is what matters.
We've been skirting the line for a while.
If I could, right now I absolutely would prefer to be sending a synthesized avatar then the real me - my desktop setup doesn't allow very optimal camera placement with large monitors, but for maximum impact I ideally want to send my face making direct eye contact with the camera.
When you upload, submit, store, send or receive User Content to or through the NVIDIA Research AI Playground, you give NVIDIA (and parties NVIDIA works with, including its affiliates, suppliers and customers) a worldwide license to use (including without limitation for neural network training), host, store, reproduce, modify, create derivative works (such as those resulting from translations, adaptations or other changes), communicate, publish, publicly perform, publicly display and distribute such User Content. The rights you grant in this license are for the limited purpose of operating, promoting, and improving the NVIDIA Research AI Playground and content available to all users, and to develop new NVIDIA offerings. This license continues even if you stop using the NVIDIA Research AI Playground. The NVIDIA Research AI Playground may offer you ways to access, download, and remove content that has been provided, but make sure to keep your own back-up copies of your User Content. Also, the scope of services is limited and not all content in all formats can be loaded in the NVIDIA Research AI Playground.
“Permission to use content you create and share. [...] when you share, post, or upload content that is covered by intellectual property rights on or in connection with our Products, you grant us a non-exclusive, transferable, sub-licensable, royalty-free, and worldwide license to host, use, distribute, modify, run, copy, publicly perform or display, translate, and create derivative works of your content”
https://www.cgtrader.com/free-3d-models/character/fantasy/3d...
I made it shamelessly with https://www.fantasy-faces.com/ for the GAN and myheritage.fr for animation -- it was still a lot of work to select 1000 of the most beautiful.
Ultimately, I did it for the lulz...
It's interesting to see how the model fails at extreme values. I can see why they chose the cutoffs they did!
Its pretty cool, but one thing that always bugs me - why is the demo page so bad?
For the majority of people, this is the only way they are going to see this thing working - would it really have hurt to have had an actual frontend dev nock something together for this? It takes away from the great work behind the scenes imho
I wish this was deployable to browsers so it was fully stand alone.
The full paper is on https://nvlabs.github.io/face-vid2vid/main.pdf . (It only mentions GPU once, for the training set.)
I'm quite impressed by how NVIDIA Broadcast cleans up a simple webcam image already, on a 3070 GPU; the background blur will get the gap between headphone bridge and head with sharp cuts - it's impressive enough in my books to warrant such a gaming grade GPU for work purposes, if a remote worker.
I have my cam off to the side; I'm really looking forward to being able to try the angle correction!