Nvidia Uses AI to Slash Bandwidth on Video Calls
petapixel.com
petapixel.com
A nitpick, perhaps, but isn't that three orders of magnitude?
We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish views, but actually be this new invention in real time. There's been a lot of discussion centered around deep fakes and its problems, this is essentially deep faking yourself into whatever you want.
Video calls are a very important form of communication at the moment, if this becomes as accepted as background modification, that would open the societal door to a whole range of self presentation that up till now was restricted to in game virtual characters.
I wonder what kind of implications that could have. Would people come to identify themselves strongly with a virtual avatar, perhaps stronger than their real life "avatar"? It is an awesome freedom to have, to remake yourself.
I dunno that I'd call it three orders - it's close at about 830x - but it's definitely not even close to being one order either.
"An order-of-magnitude difference between two values is a factor of 10. For example, the mass of the planet Saturn is 95 times that of Earth, so Saturn is two orders of magnitude more massive than Earth."
$ bc -l
l(835)/l(10)
2.92168647548360208478
That would make it 2 orders of magnitude by that method. Happy to accept that it's 3 orders of magnitude by the N=a*10^b method though. Either way, it's definitely not one."a new car is an order of magnitude difference in price compared to a used car" is appropriate even if a new car is 40k and a used car is 5k
Electric cars have two orders of magnitude less energy storage than gasoline cars, but newer ones are only one order of magnitude.
Perhaps the example was a best-case, and the usual improvement is about 10x. (That or 'order of magnitude' has gone the way of 'exponential' in popular use. I don't think I've noticed that elsewhere, though.)
Viola! My deep fake can stand in at meetings now while I code.
Today's meeting was an hour-and-a-half spent on bike-shedding the position--and color--of the "logout" button on our product page.
15 minutes were spent debating the resident usability expert who suggested white text on a dark blue background would be more readable for people with low vision. The department manager insisted on retaining pastel blue as it is his favorite color.
45 minutes were spent arguing over whether "logout" or "log out" comprises proper semantics. Our linguistic expert was unfortunately not able to attend as she was sent on a business trip earlier in the week.
The last 30 minutes were focused on team-building exercises as Bob doodled on his tablet and Susan smiled politely at the speaker whilst screaming internally as she had 3 more meetings to attend before her department could move forward.
WFH hasn't lessened the amount of bullshit, but it's become more tolerable. I mute my mic and put on a playlist with elevator music.
Now all I need is an AI standin and this summarizer and we've basically achieved universal basic income.
45 minutes were spent by humans arguing over whether "logout" or "log out", while I created both buttons (GPT-3 can already do that) and A/B tested it.
Real-time silly hats for people I talk to and I'm sold.
They had no idea.
sure is .. if you stream your face at >30-50Mbit/s. For contrast highest bitrate available on Twitch, used for streaming high motion full screen updating twitchy 1080@60 gaming, is ~6-8Mbit/s.
I mean taking away focus on things that doesn't matter in a virtual meeting such as: Where you are sitting - via Virtual Background Your daily hair style status or if you have a nose pimple - Via NVIDIAs AI showcased here. Would be great.
Though replacing yourself with a "digital" avatar I think takes away many of the benefits an actual live meeting provides.
1. the phenomenon of VTubers https://en.m.wikipedia.org/wiki/Virtual_YouTuber
2. in the virtual animal crossing late night show, Animal Talking, the presenter's (Gary Whitta) avatar doesn't really resemble how the presenter looks in real life https://en.m.wikipedia.org/wiki/Animal_Talking_with_Gary_Whi...
3. I watch a lot of interview s with people in VR Chat and it's very interesting how people seem to find it easier(?) to open up while they are embodying a character. https://youtu.be/KZWOXgc7PA4
Being able to experiment with identity in this way is really interesting to me, and I hope it becomes more mainstream with the proliferation of this technology
Humans becoming more and more dependent on virtual face-to-face meetings and also relying on embellishment of their supposed appearance through the screen. It reminds me of how SciFi authors predicted technology, but with a complimentary commentary on human psychology.
Sorry if it isn't directly related to the post, but it is so striking to me.
In his universe, both the interstellar net and combat links between ships are low bandwidth. Hence, video is interpolated between sync frames or recreated from old footage. Vinge calls the resulting video "evocations".
The plot point being that when the bandwidth gets too low, the interpolation AI has to make lots of stuff up, you are not quite sure exactly what was said.
I seem to remember the bandwidth in the book was very tiny, small number of bits per second (?) so the AI was taking the speech and compressing it into something more compressed than text then decompressing it at the other end into something that was more or less the same.
I believe this is huge and would create higher engagement if everybody was acutally looking into the camera instead of to the side or up all the time. Creating a more human an emotional relation with the people you are talking to.
https://techxplore.com/news/2019-06-intel-eye-contact-video-...
It has a very very cool twist to explain the Fermi Paradox and is a really good example of a universe with one modified rule.
It's effectively a motion-mapped keypoints of the person projected onto a simulated model. I'm assuming the cartoonish avatar was used as an example to partly avoid drawing direct lines to the full implications.
- There's no reason this couldn't extend to voice modelling as well. (much clearer speaking at much lower bandwidth)
- There's no reason this couldn't extend to replacing your sent projection with another image (or person)
- Professional looking suit wearing presentation when you're nude/hungover/unshaven. Hell, why even stop at using your real gender or visage? Imagine a job interview where every candidate, by definition, visually looked the same :)
- There's no reason you couldn't replace other people's avatar with one's of your own choosing as well.
- Why couldn't we model the rest of the environment?
Not there today, but this future is closer than many realise.
This would be interesting as an upgrade to the “name on resume” test.
Could also see a future company policy that runs peoples data through a “sameness” filter before letting them into the company to scrub bias.
Compression was expensive, because finding good morph points is hard. But now hardware has caught up to doing it in real time on cheap hardware.
As a compression method, it's great for talking heads with a fixed camera. You're just sending morph point moves, and rarely need a new keyframe.
You can be too early. Kerner Optical went bust a decade ago.
Perhaps the nets they’re using are compressing out facial microexpressions and when we see it, it seems just a little unnatural. Compression artifacts might be preferable because the information they’re missing is more obvious and less artificial. In other words i’d rather be presented with something obviously flawed than something i can’t quite tell what is wrong.
It is even less real after it went through lossy compression. The whole point of a lossy compression is to remove details that you perceive as unimportant. For example, leaves on a tree may look like a greeny mess, but that's fine, from afar, you don't make a difference.
Using neural networks for compression is far from being a new concept, and the result is not more or less real than any other technique. It is just that Nvidia implementation is really good at keeping the most important details in a small size.
If you want a more "real" image, you can just use the AI as a predictor and use the remaining bits to encode the difference, like in a traditional MPEG-style codec.
So I'd have a hard time saying this is not dramatically less "real" than some lossy compression technique. Is there some way to formalize "realness?" Maybe it would be inversely proportional to the hardness of manipulating the medium as a user..?
Also know that you can "deepfake" yourself using a traditional video encoder, just change the keyframe to someone else's face. Of course, it will look broken and totally unconvincing but because of motion compensation, you can sort of map the movement of your face on someone else's face.
The technique in the paper simply has way better motion compensation, so good that it still works if you change the keyframe. Traditional video compression algorithms don't work like that because they are not just for talking faces and can't use such advanced techniques for performance and ease of implementation reasons.
Traditional compression has no high level knowledge of what a video call looks like, the algorithms are about patterns in how pixels change in time and space, so the artifacts from the algorithm are pixel-based effects (blockiness, blurriness, etc).
A neural net that has been trained on a million faces and is setup to draw a face may draw a perfectly clear image of a face on low bandwidth...but when it doesn't have bandwidth to be accurate, it doesn't blur things, it fills in details learned from looking at strangers who aren't on the call.
I'm afraid I suspect this isn't that far fetched.
I disagree. If real / not real is binary, sure, it's not real. But if we allow "real" to be a range of values, it is less real.
Adding a secondary, uncontrolled layer of perception is definitely not the same thing, but in general, what you "see" is barely "real".
In addition to the colour grading and interpolation, lens defects might be 'fixed' before you ever get data on disk.
ML is just really good interpolation.
Is that how this works? Because it should be how it works...
Not sure if this is a viable way of compressing the video stream or if a transmission of the “diff” would be too costly. If this method would give 1/2 the original bandwidth I’d think it’s more impressive than cutting 3 orders of magnitude via “avatars”.
So either your idea will revolutionize video compression or the "diff" would bring you back to the ballpark of lossless codecs.
Now if I can use it to add a Klingon skull ridge and hollow eyes to my boss or scribble notes on my scrum master’s generous forehead we might be on to something.
In a work environment, I would expect the person I'm talking to to be presentable, ie their avatar would be presentable, so no goofy backgrounds or annoying accessories.
But the key for me is, I'd actually have something to see. So often in my work in in meetings and three people have cameras on and the rest don't. I don't really care what they look like, I care if they're engaged, nodding their heads, their facial reactions.
I don't always have my video on either, I don't have great upload speeds so I usually appear as a big blob anyway. I'd happily have whatever representation of me be in my place if it meant people could see my reactions
It feels like people really haven't put any thought into how to handle video conferences.
There are numerous articles on why this impromptu remote environment isn't the same as traditional remote environments. Are people turning on their camera during meetings? Are people actually responding when talked to during meetings? Is there some type of plan outlined for team members who have kids in virtual school or are unexpected caretakers? Have teams been given the proper collaborative tools to work together remotely? Are working hours being respected?
There are lots of things people didn't think about when places went remote. The problem is that they never went back to address them either.
https://www.youtube.com/watch?v=t4DT3tQqgRM
But after that, I was reminded of the paranoia (or not?) around Zoom and that, for an extreme example, the CCP was mining and generating facial fingerprints and social networks using video calls. It seems like this technology is the same concept except put to a useful purpose.
And they still don't often enough (see recent ExamSoft issues discussed e.g. here: https://news.ycombinator.com/item?id=24641063 )
We don't really understand just how little information is actually in a photo (we add huge amounts of info in our perception).
My guess is that predictive systems are using contrast as a guide to essentially 3D structures which, simply, just cannot be reconstructed from 2D. And therefore, probably struggle more on dark faces which have different contrast properties.
Now, I don’t know much about neural networks (AI), but my understanding is that if you provide a training set representative of the population makeup, (in America at least) it’ll be biased towards white people as it hasn’t “seen” enough black person images. My limited understanding would then make me think one would need equal white person photos as well as black person photos.
Black faces and white faces are "statistically equivalent" in 3D.
The issue is more, in my view, the hubris of calling this system "facial recognition". It isnt: it's pixel pattern color recognition which sometimes coincides with certain facial patterns.
Using their own NPU ( Neural processing unit ), you can now make FaceTime call with ridiculously low bandwidth. From the Nvidia example, 0.1165 KB/frame even at buttery smooth 60fps ( I could literally hear Apple market the crap out of this ), that is 7KBps or 56Kbps! Remember when the industry were trying to compress CD Audio quality ( aka 128Kbps MP3 ) down to 64Kbps? This FaceTime Video Call is using even less!
And since the NPU and FaceTime are all part of Apple's platform and not available anywhere else. They now have an even better excuse not to open it up and further lock customers into their ecosystem. ( Not such a good thing with how Apple are acting right now )
Not so sure where Nvidia is heading for this since Not everyone will have a CUDA GPU.
What's this claim based on? Last I looked into FaceTime tech, they didn't do anything special - their quality comes from use of H.265 and the fact that iOS devices have good quality HW encoding blocks which provide good compression at low bandwidths.
FaceTime stream is usually also low motion / change so it's possible to achieve very good compression even with basic quality. Although they still don't quite match the quality of AV1 powered Google Duo on very poor connections.
[a]: I know the password is stored on it, but idk about the face mesh
That doesn’t mean you can’t read a point cloud for other purposes.
They currently expose ARPointCloud API but I don’t know what sensors it uses to produce it;
https://twitter.com/nobbis/status/1292262455490629633?lang=e...
128kb mp3 is good enough for most people most of the time, but it isn't CD quality. Having said that, 64kb Opus is almost or about as good as 128kb mp3.
I wonder how well these techniques can be applied to audio.
And we'll notice it because somehow our and other people faces in facetime will start to subtly translate strangely aggravating emotions, as current memojis do.
Maybe it will be exclusive to Android devices.
Or maybe it will work on any device (consuming CPU or GPU depending on the hardware) but only on Nvidia's communication app.
https://blog.emojipedia.org/apples-new-animoji/
And how well does that work when you switch to screen sharing?
[0] https://github.com/tensorflow/tfjs-models/tree/master/faceme...
This is a brilliant idea, even if the hardware is pricey today.
Disclaimer: no affiliation, but I use Zoom/Teams/Slack/FaceTime/YouTube.
Big question though - is this just substituting the problem of not having good internet with not having a really fast nVidia graphics card?
The bandwidth and requirement to participate correlate.
Knowledge workers are all about the knowledge that they build up over time when working with a particular environment (be it an industry, a system, or a person/group of people). That knowledge is non-deterministically synthesised in the brain based on the experiences of that person, and being non-deterministic, no AI will come to the same conclusions about every item as this particular human would.
In that case, an emulated personality that is meant to make themselves available as your replacement will be an impostor. One that is less of an expert than yourself at best, and at worst one that is misinformed or misled on various issues (which in turn causes other people to be misled or misinformed).
If the goal is to make it seem like your present when you aren’t, I can believe we’re only a few years away from Reynolds‘ Betas — a Markov chain can’t mimic a human well, but it can mimic a human; GPT3 can do better, and while it still isn’t great, the main reason it feels like it might not be enough for public figures is how easy it is to get it to answer as if it were someone else rather than as The Right Honourable Sir Obvious Madeupname, MP for Oxbridge-upon-Wells who is paying for the chatbot to mimic him in particular.
It's not uncommon to see video calls at 100kbs-150kbps, which is ~10KB/s, and this is for 7fps or so, including audio. So "per frame" that would be 1KB or so (more for key frames, less for I frames).
So they say it can be 0.1KB, so better than that... Exciting, if realistic.
Also, add on top audio, and packet overhead :-) there is at least 0.1KB overhead for sending the packet (bundle it with audio if possible!)
Then use first-order-model to extrapolate 2 seconds of video from the keyframe.
Rinse, repeat.
Very doable. AMAZING!
The original first-order-model could not do 30 frames per second, but maybe this Nvidia model has some improvements.
1 - https://aliaksandrsiarohin.github.io/first-order-model-websi...
It's like with those news "the new battery type has been discovered", with very little actual data, just guesswork.
Agree, the woman's mouth in the video looks _very_ off at 1:03 in the video.
Petapixel is a blog spam site btw. Why not go to the source that is linked in the post?
Even JPEG takes advantage of human perception, throwing away first what we can’t perceive.
spotty mobile connections with a powerful device (iPhone Android etc...) it might make sense.
There's a real network effect with things like codecs - unless some significant proportion of calls can use it, it'll remain a cool but obscure experiment.
I hope Nvidia have the foresight to release something that'll run on any hardware, and under a permissive license, but I suspect not.
The idea is out there already (it's basically deep fake tech, right?), and I'm sure it won't be so long before some open source version of it gets released. Nvidia would be wise to get out in front of that and at least have their brand associated with a widely used variant on the theme.
The issues are social. I would hope that the receiver is the one able to choose between original or AI stream, as I can understand some people being uncomfortable with the artifacts, gaze, expressions, and other edge cases. But when the quality is higher I could see a lot of people preferring this option as a default.
Recall the promise of 5G is an order of magnitude increase in speed (among other things like low latency).
If we can get there by reducing bandwidth requirements by an order, that will be great. Wonder if it applies to Netflix...
Not sure if I understood your post correctly, but you're slightly misled here.
Watch the video. This is not a new general purpose video codec, it is basically Deep Fake - taking a still image (key frame) and superimposing detected facial expression/movement on this key frame (leaving out some technical details).
This is an improvement over (not-anymore-)state-of-the-art h264 since transmission of only a few coordinates mapping your facial expression is significantly less data to transmit then a delta of arbitrary video and periodic keyframes (again generously leaving out important technical details of modern video codecs). Trying to reconstruct a moving car/background/etc. from these facial expression key points will lead nowhere.
> Next step would be to just predict both sides of the conversation and sever the real-life link entirely.
Gmail already does a little bit of this. Google books appointments over the phone on your behalf.
We're on the road to this...
In the final form, use "You=" as reference and just press it one at 1 seconds to simulate keyframe.
AMAZING!
> Called “Free View,” this would allow someone who has a separate camera off-screen to seemingly keep eye contact with those on a video call.
Am I the only one who thinks eye contact on video calls feels creepy? I think I would prefer this feature to remove eye contact on video calls rather than add it.
So a system that maintained eye contact continuously would indeed risk looking creepy!
[1] https://www.forbes.com/sites/carolkinseygoman/2014/08/21/fac... [2] https://www.businessinsider.com/heres-how-long-you-should-ho...
I could even get an accomplice to do it while I'm talking to you. They would have your today clothes on and you'd be tied up talking to me.
I'm dubious on the tech being as good as they say now. But it's getting exciting.
So they're trading bandwidth for CPU load at either end. I wonder what the tradeoff is in terms of energy? Would this result in higher CO2 emissions?
Given it's Nvidia I would imagine it's more likely going to be GPU load rather than CPU load. Don't underestimate the current computational overhead of existing lossy video compression.
The server is in the same city as me to avoid excessive latency.
The benefits of being able to develop anywhere (including from a phone in a pinch), and being able to add extra CPU's and RAM at the click of a button outweigh the need for a network connection for my usecases.
(Yes, I know this is realtime webcam footage, not recorded footage, I'm just curious).
After all, it doesn't need to be said that when you're on a regular videoconferencing call and bandwidth starts to suffer, the resulting images don't really look anything like a photorealistic person either. I think this is actually a really good use of NN.
Edit: [1] I meant to say doing a mixture of these (with the NN image as the "base", with H.264 to improve accuracy) seems really hard. On the other hand, just a hard switch from H.264 to NN when quality degrades is probably quite practical?
I don't want bandwidth and things spent/wasted without it providing a significant benefit (I'm probably one of those 1080p/720p is good enough for most things type guys).
I definitely don't want work making large bandwidth or resource claims on my connection when I'm working at home. And if any of this remote working has taught me anything, it's that most of my colleagues don't have steady/reliable tech or connections, so i'd almost want it used pre-emptively as a default so we can spend the rest of those resources on robustness or other qualities. (I realise of course that at the moment none of them have high-grade Nvidia graphics cards, but I'm talking hypothetically in the far off future).
In short, I want a world where the cost to benefit ratio of things is orders of magnitude larger, because things like this let us spend network/resources on things which matter.
Yes, when I'm calling my parent/grandparent one on one I might want to upgrade the signal, but I don't need to see random colleague's face in all their HD glory, or remote people whom I have no idea who they look or sound like anyway (i believe that's also been one of the findings with deepfakes, that you don't notice the eerieness/falseness as much if it's a reference of a face that you don't have pre-determined knowledge of).
I tried to pay attention to what you and others flagged (parts not moving, facial expressions, shoulders, etc.) But I cannot spot anything out of place.
That said, looking at other comments here, you're definitely not the only one
That way the doctoring isn't apparent, you still benefit from the massive bandwidth savings (which I consider most important anyway), and it appears more believable in the real-world context of variable bitrates.
I made no claims about the cost. :)
I guess I'm also talking about the theoretical maximums in ideal conditions as well, which is sort of cheating on my part. There would be times when you can't get the optimal speed if, like you say, the sun is the way.
In fact, I struggle to see any difference from the original. It totally feels genuine, I'm not sure why you perceive them to be npcs
This is not fundamentally different from other lousy compression algorithms.
The bits that are missing are reconstructed using a bias built in the compression algorithm. Deep learning based algorithms simply have a more realistic bias, so the artifacts that introduced to replace the missing bits are less noticeable.
Sidenote: I always had this idea for "video compression" I'm by no means an expert in compression.
1) Take like the top 10 most "VISUAL DIVERSE" movies (imagine Forrest Gump(slow drama) VS ToyStory(anime) Vs Rambo(fast paced visuals)
2) "Encode/compress"(I know these terms are not interchangeable) the movie as a "diff/ref" to the "most similar" movie from step 1
2b) This the "diff/ref" can be many forms "sliding window" over x section of y amount of frames.
3) The "end-user" or "destination" have these "10-master-movies" locally on hdd and together with the local-data can construct the original "frame" or "movie" from the compression and local-movie on disk.
Tl;DR Try to compress a new movie by saying "the top corner of frame 1-120" is very similar to MasterMovie-2-FrameXYZ-Frame-ABC
4)