Nvidia replaces video codecs with neural networks for virtual meetings
youtube.com
youtube.com
https://appleinsider.com/articles/20/06/22/facetime-eye-cont...
It seems like this could /also/ be used for video by using this technique along with residual coding.
Not trivially-- in the pixel or DCT domain the residual would almost certainly be extremely non-sparse, and would take a lot of bits-- potentially similar to just sending an image (for a given target MSE level). Consider, edges (other than the keypoint controlled ones) aren't even in exactly the same place-- so the residual doesn't just need to code the edge, it needs to code it twice. This has been one of the big impediments in using 'synthesis' techniques in video coding generally.
It might be possible to code a residual in some latent NN space and get more useful results, however.
Looking forward to have these available in consumer hardware soon.
More examples here:
But I haven't seen anybody touch on the compute cost required to implement this. As I'm not in the machine learning field I don't have a good idea what the compute cost is for something like this. Can anybody chime in on that?
If this "codec" were to require a somewhat beefy gpu I don't see the benefits at all. Current H264 is usually done by hardware decode and sometimes even encode. In areas where bandwidth is constrained I would imagine a lack of computing resources, thus nullifying the entire premise. That said, in current times it would save a substantial amount of data transmitted. But I'm not sure if we should lock-in our entire videoconferencing system to nvidia just to save some bandwith.
I thought comparing it in KB per frame was a strange way to measure it, since video codec are used to measurement similar to Network in kbps or mbps.
So the Video Codec was actually 50kbps, which is indeed a very low bitrate. But this was done on H.264, which is now nearly 20 years old. Modern Codec like HEVC and VP9, or State of the Art like AV1 and VVC would have done much much better.
Next problem, would this only work on Nvidia GPU? Apple are already doing something similar to FaceTime, but only with respect to eye contact. Are we entering an era where even AI video codec are bound by devices?
I used to hope and wish Apple introduce these kind of features to iPhone. But their act and response on App Store is making me wary.
I don't think it's fair to call this video compression as much as real-time photo-realistic animation via motion capture.
Future headsets might start to implement some Leap-esque sensors pointed at the user's chin and some eye tracking cameras inside the headset to address this eventually, but that's going to be prohibitively costly for some time (not just in terms of dollars, but in added weight/heat as well).
Face time calls on a remote satellite internet setup will be revolutionary.
Need a fixed camera, one face and a fairly static background so there goes mobile or conference room use.
Reminds me of a section in Hofstadter's 'Godel, Escher, Bach' about there being knowledge in the signal vs. the receiver, or something akin to that.