How to make video calls almost as good as face-to-face (2020)
benkuhn.net
benkuhn.net
We run https://cleanfeed.net/ to focus on full-duplex low latency audio (no video). Of course we are dealing mainly with studio environments and good quality equipment, making it a different challenge to a regular meeting tool.
But since the first prototype, we were using it ourselves as an "always on" line between developers. The ability to have 3 or 4-way conference in full duplex with low latency appeared to foster much more of a 'human' relationship than the meeting apps. The lack of picture may even have a positive effect, because it was part of privacy that made people more relaxed.
I have tried to build an "audio office" product based on low-latency spatial audio over the web but it's way too far outside my expertise. Glad to see that someone else is on it!
I had someone send me a recruiting message on Stack overflow jobs once about a product like this, and it's one of the few times I declined on the basis I fundamentally disagreed with its existence. How do you focus?
* Move the audience’s video as close as possible to the camera (make it smaller), that will have you look like you’re “looking” at the camera versus off to another screen
* have the camera as close to eye level as possible
If you just open up a camera view there's huge latency without it even sending data anywhere. The internet latency is probably only a small portion of over all latency and WiFi latency isn't much either. There's already lots of latency with audio (Perhaps having a microphone could make audio travel faster then through air?)
I wish people would take latency more seriously. Using a rolling shutter, encoding a single line at a time and sending it via UDP could probably save almost all the latency. Sub 30ms latency IMO would make it much better.
Wi-Fi is also a huge contributor to latency. In congested areas or with problematic APs or end stations it's not uncommon to see packets delayed more than 500ms by retransmits. The average might be lower, but with frequent excursions to much higher than average latency.
Home WiFi generally has really bad worst case latency and variability (jitter), removing it from the path is one of the best things you can do to improve realtime communication performance.
When it's a fight between Bluetooth and wifi, Bluetooth almost always loses -- big time.
once you have to scroll in the ap-list there is a very good chance that you are contested on air-time because some neighbour is active on any given channel. [0]
add to that interference from non-wifi or other unclean rf-equipment like babymonitors, wireless doorbells or bluetooth ... 2.4ghz is quite crowded, 5ghz is crippled by a wall...
Due to bad engineering, a WiFi chip can either be scanning or transmitting, and background location tasks can request a WiFi scan at any time. Depending on your OS and what else you run, this can cause stutters in the 250ms-1second range.
I consider this irresponsible in the face of the Zoom CEO saying to moneteize the knowledge they have about you. Very different from face-to-face in my opinion.
Face-to-face is not only about interaction of people present, it's about eavesdropping as well.
Praise for the rest and for caring about the other participants.
Impossible in Google Meet or Teams, which really sucks.
You both are looking at each other’s faces, which are not the camera.
Seems like it would be possible to transform the person’s eyes so they’re looking straight forward, not down.
This matters more in applications like therapy over video calls.
NVIDIA also have an implementation in their Maxine SDK (https://developer.nvidia.com/maxine#ar-sdk).
So if you break eye contact with the image of the person on the screen, it will show you breaking eye contact in your video.
This could be terrifying when it goes wrong.
It can even run the same noise cancellation algo on incoming audio! So you can filter colleagues' noise for yourself if they don't want to bother
Tips for immersive video calls - https://news.ycombinator.com/item?id=24610166 - Sept 2020 (138 comments)
1. We still don't share a common digital space when having meetings, because everyone sees tiles in a different order. "Let's go around the table and give our updates!" Uh, in what order? There is a lot of power in the subtle social cues available when everyone knows who you are talking (or listening) to by shifting towards them.
2. There's still no effective way to non-verbally signal that someone wants to speak. If everyone was push-to-talk, then holding down the unmute button would let the UI emphasize your tile and signal to others. This would cut down on verbally stepping on each other. Nobody uses those "raise hand" buttons, because they are too single-purpose and relatively high effort.
3. Often, there's no way to tell at a glance if you're muted. I can't believe Zoom still does this: you have to move your mouse to get it to show you the UI. Why are you hiding the UI.
4. What is the point of reactions appearing on your own tile? You're not talking: nobody is looking at you. Reactions should appear on the tile you are reacting to!
It's small things like this that make video calls a mess, and nothing's changed since the beginning of the pandemic except for video filters.
The most baffling thing to me is how common it is for video conferencing systems to either lack a push to talk system or just have a really really bad one. I agree this would CVS fantastic but they’ve gotta actually put it in first.
Did they change this recently? On Windows I always see an icon on the lower left corner showing that I'm muted without having to move my mouse or show the controls. I do prefer to go into the options and enable the Always show controls setting under General just so I have a bigger indicator. I think pressing Alt also toggles the controls.
Some tools have a raise hand feature. Though not sure if it brings such users to the top of the participant list or otherwise highlights them beyond an icon.
Really? Twitter spaces seems to do this pretty well.
#1 I’d love to be fixed.
It's orthogonal to having decent gear, however. Using a flattering aperture and a sensor with adequate sensitivity, not to mention a real microphone, will create a positive impression in a mostly invisible way, and building software which isn't painful to use would magnify this effect rather than flatten it.
also as debate from an HN post from a couple days ago covered well, using an external camera instead of webcam is a fool's errand. my how I've tried.
the most important change for reducing fatigue in my experience: turn video off. audio only.
As for Google Meet, I still can't get a non-grainy image of myself or the person at the other end of the call.
The biggest issue is that some video conferencing services don't support standard protocols, so you can't connect to Zoom without buying a special license. A WebEx subscription with some of the newer equipment is sufficient to connect to Teams and Google Meet through WebRTC as well as any standards based systems (Chime, etc.)
Also on shaky ground is the advice to use a DSLR camera as web cam. As pointed out in a recent discussion here, the problem with adding more devices to the chain is... more devices in the chain.
Dedicated diffuse lighting, wireless headsets, separate cameras etc.. participants will need to be patient while you prepare and configure your film set.
I used a wireless headset for awhile, but swapped it for a wired headset after too many times joining meetings without the headset pairing in time.
I forget to charge it sometimes but that's on me.
I like pacing during meetings
There are also DECT headsets which seem to have very low lag. Quality isn't great, but you wouldn't expect it to be given the size of the phones. But it works great for voice and has a very long range.
In my case, I was using a Logitech bluetooth headset. When calls came in, I would need to switch the headset on, wait, hope it connects, hope Teams hadn't switched to default mic, and hope the battery was still good. I grew tired of the uncertainty and went back to a wired headset.
Imports tariffs for digital video cameras are significantly higher than tariffs for still cameras. So companies like Canon, Nikon, and Sony artificially limit the maximum record time on their DSLRs to get around the higher taxes.
For personal use it's fine. But for work meetings where someone calls you unexpectedly is where things can go wrong if you have unconventional equipment dependencies.
The author recommends an "old DSLR". But you will need utility software running, it's not just a simple driver. Manufacturers make and update these utilities at their discretion. It's a gamble whether it all runs smoothly.
The DSLR needs to be in a special mode, and won't wake up in that mode automatically for incoming calls if powered down, etc.
[*] A power adapter that feeds into a battery shaped terminals, thus allowing an "infinite" battery which. It could be better than USB charging while using the battery because USB charging might not keep up with the battery drain. The downside of course might be malfunctioning dummy battery causing your camera to fry.
If you really want that full-duplex audio experience on linux today, disable audio on the call and move everyone to https://cleanfeed.net. This is not an ideal workflow, but it gets you there.
Why other products don't just take audio seriously is a pity.
I've been using my DSLR as a webcam for the last 2 years. I set it up on a short tripod so it sits just over the top of my monitor. The only time I have to screw around with setup is when I take it down to shoot my kids for family Christmas card. Takes 5 minutes and then everything is back in place.
It never ceases to amaze me how much a bunch of so-called hackers can argue about something they've never even tried to do. Hours spent bickering back and forth about something that takes literally 5 minutes to test out.
Anyway, I use a fixed 50mm lens on my camera, which is a classic lens for portrait dimensions and gives me a very shallow depth of field to blur out the junk in the room behind me.
For portrait photography, yes. For meeting with colleagues, no.
Audio quality and network performance is light years more important than focal length and bokeh of your talking head.
Of course if you sound like crap and can't stream at a decent bitrate, fix that first / as well. But with crisp audio and a decent Net connection, there's plenty of room to improve video quality from awful to decent, and your colleagues notice both, whether they realize it or not.
Perhaps you mean "care".
To not realize a visual quality difference, is to not notice.
I was pretty dissatisfied with my work at the time. I was spending a lot of time arguing with an IT manager who thought he knew things about software development on one hand, or just writing boring-ass Crystal Reports on the other. Then this newfangled device came out called the Oculus Rift DK2. I also had a Leap Motion device that I had glued to the front of it. It was an exciting time, as decent enough VR to not make the majority of people sick or have headaches hadn't been available--and certainly not at a consumer price point--until then.
I hacked around on some things, got a small amount of notoriety amongst other burgeoning VR developers at the time (I'd basically made the first JavaScript library for making VR-capable apps), and then one day I got a call from this company called AltspaceVR. They wanted me to interview for a job.
We did it in Altspace. This was my first experience with VR chat. The person I talked to had an HTC Vive dev kit. We stood in a futuristic room and talked, face to face. And it just worked. It felt like actually talking to a person. The conversation was so easy, compared to 2D teleconferencing. I wasn't distracted by my own video feed, worried about how I looked or whether I was providing enough "camera contact". There was body language and a feeling of the bodily presence of the other person.
It was immediately obvious that this was far superior than traditional teleconferencing. I had a brainwave about the future of work, meetings with clients without having to travel, my mom no longer destroying her knees, no longer having to fly around the country. And I could be involved on working on this future! If I just abandoned everyone I knew and moved to California.
I asked the guy I was interviewing with if he really thought that was necessary. He said they really liked having everyone in the same room together. I asked him what were we doing, just now, talking to each other as if we were in the same room. He vacillated, admitted they weren't dogfooding their own product.
And that's when I decided that I had to work in VR, had to work on productivity apps and not games, and had to dogfood it and stick to living on the East Coast. Now I'm the head of R&D at a foreign language training company, where I make a VR teleconferencing app for people to meet together in culturally-appropriate environs to practice their language skills.
Congratulations!
I've always wondered about one thing -- with VR, your eye convergence distance is whatever the software wants it to be, whether that is something in your virtual hands, or something across the virtual yard. But your focus convergence distance is just millimeters away from your eyeballs.
But for the millennia that human beings have evolved, our convergence distance and our focus distance has always been the same. If it's in your hands, then the distance is about three feet, or whatever. If the object is in the yard, it might be thirty feet.
So, how do you deal with this problem in VR?
you soon realize that only certain meetings work over video and some are only worth doing IRL.
Nice clothes are good to have regardless of work, and many jobs DO offer commuting benefits. If a job requires me to drive to work, at the very least they need to provide a place to park. If my job requires me to have a better camera or microphone than what is built into the laptop, why is it on me to buy that?