2021 will demand new kinds of video conferencing
axios.com
axios.com
From the hyperlink in the article. This for me is the kicker. From the start to end of a zoom call I feel like I'm on a stage. When I am in an in-person meeting I feel like I am "hanging out". One drains me of my mental reserves, the other fills it up. Zoom just feels draining. They noted the fact you see yourself and there is real time self critique to the hissy sound and pixelation. After a year of it .. I'm close to done.
I wish it had a two-tone high-contrast filtered version of myself, that I could use merely for positioning. Might be a potential solve?
(Incidentally, I really dont like working remotely and want to get back to real meetings as soon as I can, even if it will make me more anxious in the moment)
I haven't tried it for a big call yet, but Jitsi seems better. For just a chat Discord still reigns supreme for call-quality of the services I've tried, although it may be bad data because I'm usually playing a game or helping someone solve a problem rather than solely conversing.
I do a lot of video pair programming, and I feel supremely comfortable screen/video sharing. Getting to work on my own machine with my own IDE and shell set-up, my own hardware in my comfortable home office, 10000% better than awkwardly scratching at a whiteboard speculating about what may or may not work if committed to code.
I'm a firm believer that video adds nothing to the call other than add distractions.
I agree. And I would add that voice does not add anything either. It's better to do everything by chat or by mail, that at least leaves a usable trace.
I’m the opposite. In-person meetings are extremely taxing while zoom meetings are very relaxing. I’d love to never have an in-person meeting ever again.
My experience with video conferencing is that it's mostly draining people's attention away, while being completely unnecessary. We switched from seeing each other's distorted, pixelated, and badly lit faces in a video conference back to plain old phone calls. Nobody missed the video, so problem solved.
In my opinion, the writing was on the wall when mainstream media started having makeup tutorials for video conferencing: https://www.vogue.com/article/how-to-apply-makeup-for-zoom-v...
or when people started having legal battles about video avatars: https://www.theverge.com/2020/11/24/21591488/projekt-melody-...
But also, the problem is by no means new. Here's a great quote from 1996:
"And the videophonic stress was even worse if you were at all vain. I.e. if you worried at all about how you looked. As in to other people. Which all kidding aside who doesn’t." - David Foster Wallace, Infinite Jest
Read his funny take on the rise and fall of video conferencing, and you'll be cured from the hype :)
Yet, I find most of the things people visually do in person to be fluff that can be done without. Meanwhile, video-conferencing either requires an expensive setup (not everyone here works in the US, let alone FAANG) to really start emulating "in-person", or it forces somebody to self-police themselves to a much larger degree sitting in front of a webcam with no video or audio privacy (I guess this at least emulates the open office disaster). And it still misses the forest for the trees. Additionally, trying to shoehorn fun into work rarely works.
(Also, as far as I know, there is no way to stop looking at your own face in Teams besides turning off the cam. Looking at a mirror 24/7 isn't even close to reality).
DFW passed away 12 years ago, quite a lot has changed since then, let alone people the pandemic social isolation due to Covid19.
Isn't it the core of the problem?
Sure, I check if I'm not showing overly dominant nostrils due to camera angles or have food anywhere on my face but other than that, I's just how I look.
I'm personally an advocate for more asynchronous communication, for less video (I find video very tiring and time-consuming), and for more audio at work. I feel audio is the sweet spot of informal, low effort, and still more emotionally engaging than pure text. I believe this so much that I'm actually building a chat application, called heysync[1].
It's somewhere between slack and discord in terms of features, but intended to allow for asynchronous, conversational audio to fit directly into the way teams already chat. You speak, and it can play live to any teammate who's listening right then, but it also records the audio, and transcribes it inline with regular chat, so the message can be easily read or listened to in-context later, allowing for true asynchronous audio conversations without dedicated audio rooms. Of course, you can also just chat via text like you would with slack or IRC.
There is an early info site up at https://heysync.chat/ if anyone's interested in this idea :)
Anyone who prefers text, could view a transcript or submit text.
"Inverted Slack / Zoom / IDE" ... turned inside out so they support attaching multiple simultaneous editors, and attaching apps to each other to form ad-hoc "desktops".
This has been done over the years, in different forms. And it is a lot of fun, speaking from personal experience. And the market need was not there, for a business to be built on this, or a standard, until perhaps now because we're forced to be separate for at least four more months.
Video conversations for the talkative, in a mutable / auto-transcribed, resizable / mutable widget. Polls for those who just "need a decision", where additional "poll options" can be added by others. Text chats for asynchronous discussions. App-scoped shared copy-paste-buffers. Live people Directories showing which apps your colleagues are "in" now, groups of which can be collected into ad-hoc dynamic work-desktops, to focus on a task.
I swear, the best video chat experience I had was from The Go Game. At no point were you dumped into a massive shared room. Instead, people can join others at-will into small "rooms" of up to six people, which is much more manageable. Then you can implement whatever games you want and let people decide what they want to do. No need for next-gen hardware, no fancy "VR spaces", just a different interface paradigm that actually scales.
So you do need somebody technical to make the space work how you need, and if your users are all non-technical you probably only want one type of rule everywhere or it'll confuse them.
If you do that, it works a treat.
1. You don’t see who is speaking when there are many people present— you have to scroll through participants to find the person talking. Lots of new employees, I didn’t know who was talking about half the time
2. Connection issues. I couldn’t hear the speaker but others could, for extended periods of time. Missed entire presentations due to this unfortunately, as I couldn’t hear or have any means of connecting to the speaker.
Don’t get me wrong it’s fun software, but it’s not ready to be your rock solid all-hands replacement quite yet. Had the pleasure of experiencing Welcome at one of the YC networking events, and that worked really well, curious to hear if they’re interested in filling this gap
Was the speaker at a podium? Were they spotlighted? Were you close to them?
There are several reasons why you might not have been able to connect that are part of the design. We found initially people were confused by the connection rules, but once they got the hang of it, it all "just worked".
I agree completely that it takes time to get used to, and perhaps you did have genuine connection issues, but we ran with 150 to 200 people for 48 hours without any significant[0] identified connection issues, which is why I wonder if it was mis-identification of a "symptom".
I'll check out "Welcome".
[0] There were occasions when people needed to reload the tab, but doing so always fixed the problem.
FWIW the person I could no longer hear was located across the globe, so maybe that’s related.
I apologize for not having a more thorough write up of what happened, but it was certainly a real issue.
As for the other issue, active speaker detection is on our near-term roadmap. It's become more pressing for us too, as our team has grown and no longer fits in one page :P
Thanks for the feedback!
I’ll keep the feedback method in mind for future usages!
Also, I get it, having started to work with the WebRTC stack myself. While the service is free to use I think I can stomach a few issues while it’s actively being worked on :)
Indeed!
I just had longer post about it https://news.ycombinator.com/item?id=25689370 because people try to solve problems that are already solvable by P2T
Yep, it's pretty abysmal for class settings.
I know it's partly down to configuration/instructors but being shoved into a room by an instructor with 40 other people forced to stare into a webcam while watching everyone else doing the same is not an ideal environment.
There needs to be a lot more attention on how and when to select additional participants in the general case where there's only one or two priority speakers and the rest are basically just sitting around listening.
Now of course me personally I'd just rather not use a webcam at all (Or only when specifically called on), but often that's not a choice unfortunately.
But I 100% agree with you that this is a design problem to be solved first with better interfaces and workflows, not AI. Innovation is the sum of behavior change and AI is not going to be doing our knowledge work for us anytime soon. Given we're doin the work, we deserve better tools to do it ourselves.
I'm currently building grain.co that layers a new kind of video conferencing on top of Zoom, focused on brokering the recorded information gathered during sync calls into the async tools where knowledge lives and actual work is done after the convo ends.
Our approach is to make it easy to annotate in real-time, clip out those parts you want to save/share, and then push them into async tools where anyone can view them like it were text in a document or a message in Slack.
EDIT: not a criticism, just made me think of a tangentially related thing:
This is so emblematic of how AI is a moving target. Once something becomes an everyday tool it's no longer AI?
Not so long ago, being able to transcribe speech from multiple people in natural conversation would be the hallmark of AI research.
This is one of the things that bugs me the most. It makes it really difficult to have a shared conversation, such as going around the room answering a question.
I've been using [1] for system-wide push-to-talk on OSX and it has worked well for the most part.
It's because none of the video chat apps are competing on features, only on user lockdown. Their KPIs is how many users they have, and that only needs enough features/price to convince the corporate buyer which rarely accepts feedback from larger set of employees.
As such, outside of certain older very expensive corporate videoconferencing setups that remember targeting inter-corporate meetings and the like, the actual feature set ossified. You get some things that are easy to demo in 1:1 call, you get locked down room systems that leave me weeping for Webex ffs, and you get omnipresent WebRTC in-browser setup that often murders your CPU unless you get lucky.
I find it really bad that the best video conference experience I ever had involved Cisco Webex hardware and Bluejeans, all interconnected with so-maligned SIP. And before that, the glorious days of wider ITU-T based video conferencing, like NetMeeting with its whiteboards and active screensharing.
Imagine if video conferencing solutions competed on top of common protocols, and you'd separately invest in infrastructure and clients? Instead of the rare enterprise cases where Lync/Teams/Webex/other SIP is set on dedicated lines in large corporation, make similar setups common everywhere (WebRTC/SIP bridges/transcoders on the edge, with dedicated bandwidth bough by company, etc.?) with vendors competing on something more than "you have to use our stack because we don't interoperate with anyone else".
Sincerely, someone who has had enough keeping 3-5 video conferencing apps around.
I used to multitask during meetings fairly frequently. I’m trying really hard to stop.
If there’s the chance that someone can multitask at a meeting easily and very effectively, maybe they don’t need to be at the meeting?
Speaking from personal experience, you often don’t have control over this. If I don’t attend meetings I get invited to, it’s seen as a lack of interest and being disrespectful. So I attend them to show my face, and then just do my regular work. And if I don’t, then people stop inviting me all together and assume I’m not on the project anymore. That leaves me with little choice.
The whole thing about “leave meetings where you’re not needed”, only really works when you’re in a position of power.
[1] I mean, ISDN video calling is probably low latency, but good luck getting an ISDN line installed these days, and video quality was trash.
I really like cameras on so I can see people's faces. It feels much more like seeing people in person and the fact that my new team doesn't do that is making it harder to feel part of a group.
While I really like cameras on for calls, video quality isn't important (to me), but audio quality really is. Bad audio whether it's background noise, bad internet, echo etc. makes vc very tiring. The video could be 640x480 and 1 FPS and that would probably be fine - helps me know there's a real person on the other end of the call.
- Zoom definitely has mastered the ease of logging in whether you have the application or not as well as the ability to see multiple people, but obviously they have had their privacy issues.
- Teams seems to have good screen sharing and chatting integrated as well as some other small features here and there, but if you dont have the application, it takes some time and the browser based isnt bad but could be better.
- Google Meet- All sorts of improvements are needed...
- Jitsi - Definitely has its positives but I have always had connection issues for whatever reason.
Would love to see Apple do something more with Facetime...
At the end of the day we want to be heard in the first place. Nothing is more irritating than someone's (or your own) poor connection that makes people ask each other to repeat, watch participants freeze, wait, waste time trying to reconnect etc. etc.
Can it be solved? Degrade the quality, compress more but deliver the human speech as nicely as possible, on time. I feel not all options have been tried in this area yet.
My only real qualm with video conferencing is the camera placement issue that makes it difficult to have eye contact with the remote side. I consciously try to look up at my camera so that it appears I'm looking directly at the other side, but that's a bit unnatural.
Not perfect, but it was pretty good, and people liked it.
Some of the user interface was confusing as a first-time user, but overall it was effective and people wanted to do it again.
* A virtual 2D 'office' with positional audio
* Super low latency (~10ms)
* Minimal data/power/CPU usage
This would simulate what it feels like to be in an open office (when they work well). I can hear interesting things going on around me but because the audio is in 3D space I can pretty much filter it without thinking. If I say "hey has anyone gotten this error before?" the context clues of my volume and position make it clear to others who I am talking to.I actually started building this last year but then the pandemic hit and I figured someone would beat me to it now that there's infinite VC money flowing into team chat. Still nothing!
The current solutions where it shows everybody's face is not the best solution for bigger events or conferences, where you can go from conversation to conversation freely. It would be very minimal compared to a game, like one room or something, but the most added value would be to be present, more actively participate and not just listen somebody else talking.
We could go from room to room, conversation to conversation.
I agree, I can see how some companies would dismiss it on those grounds. I'm glad I've never been involved with companies like that, because it feels like a good balance between system demands, usability, effectiveness, and ease of use.
So far I've found attempts at full 3D "proper" environments cringe-making. Perhaps they'll improve, and my fear is that when someone gets the 3D environment glossy and slick, the suits will make people use that, even if the video and audio components are vile, simply because it "looks better".
It will be interesting to see how all this plays out over the next 5 years.
Tech companies will be video-gamey and childish but fun and effective.
Old school companies will have pointless expensive VR that is boring and slows everything down (but kinda works, and is sold by Microsoft).
I've looked around a lot and couldn't find anything that isn't a complete CPU hog.
o there is no video feed
o You only have an avatar
o You need to wear VR goggles
However in practice it feels far more natural than "good" VC. The biggest reason is spacial audio. In VR VC its perfectly possible to turn to your neighbour and have a conversation without disturbing the rest of the meeting
Even better, its possible to sustain a normal social speaking flow. There are no awkward pauses, just free flow of talking like you are in the room.
For those not familiar with the "Finnish bus stop": https://twitter.com/eeroramo/status/1238524451509145603
Different people like different things. Last year I had a team of ~8 people who hadn't worked together before working closely together. Cameras on most of the time helped us build relationships. I've just joined a different team now (who have been working together for 3 months already) and it's really hard to feel like I'm part of the group.
What's your issue with video on?
This will return. We just need a "what is a client-host" / "what is a distributed server" + structured data standard for sharing the interaction part, like IDK torrent, and a way to describe dynamic groups / sets of apps which are relevant to a workflow or an experience.
The level of interactivity, the cognitive / attention demands on a user also should be modeled. Asynchronous texting, or live chatting, or even most active: producing an experience moment-to-moment.
IDK - it's possible and clearly this thread shows the interaction / synergy potential. To help keep us together.
I created:
To host these experiments. For instance, if you want to chat with people around tables organized by pubs in the golden mile (from the movie The World's End), go to:
https://drinkingclass.com/go/goldenmile.cfm
If you want to watch random questions from jeopardy with others:
https://drinkingclass.com/go/jeopardy.cfm
I am looking for creative and interesting uses for video conferencing technologies. Also interested if you find any issues.
1) the ability to direct-channel voice call someone in the room. Can mute (or dim) the rest of the conversation that is happening, but allow for focused discussion aside. This would be a huge advantage over actual meetings where we must use furtive whispering. Obviously not great for small meetings (distraction maybe) but excellent for large ones and really excellent for social gatherings.
2) when using breakout rooms, the ability to still see what is happening in the other breakout rooms. Just like when we are separated in groups in a physical room and we can still peek over at what may be happening in a different group (in case we may want to bounce to a new conversation).
We've been in lockdown for a month. We're all a month or more from our last haircut. We live in the homes we had when it all started. We don't have control over that and it shouldn't affect my perception of candidates.
Besides, it's a remote job. You don't need to look nice for it. You can work naked for all I care.
It made me realise how superfluous pictures on resumes (Germany), LinkedIn profile photos and video calls are in a neutral hiring process. Even without them, the WhatsApp or Skype profile photo betrayed their looks.
I have zoom meetings that prompt me for my name every time. I don’t want to create a zoom profile, I don’t want to login in. I just want it to remember my name, on the client.
I think stuff like this will get ironed out now that folks are using it day in and day out. Although, it’s still funny how it hasn’t hit 100% yet. My org has been 100% teleworking since April and last week someone said “this is the first time I’ve done a Teams meeting.” I was so curious and wanted to figure out how she’s avoided any video calls (my org is 100% Teams) for so long.
And you know how highly competitive e-gaming teams compete at things like Rainbow Six and other things without being in the same room? (Not for a tournament, but competitive enough to beat any casual team online).
I feel like, yes there is lots of room for improvement, but part of this is actually being overblown. By people not willing to type in a Discord or who don't own a comfortable headset and know how to use it. Or people who have not learned how to use an online whiteboarding system or Google Docs etc.
Latency is always going to be an issue unfortunately, maybe the camera could be better at recognizing facial expression and body language to indicate when someone is about to speak, or wants to speak.
I haven't tried the Quest 2, but I hear it's lighter. Also have yet to try a Valve Index.
At this point in my experience with video calls, I'm just happy when everybody's wearing headphones so they don't get their audio dropped when somebody talks over them.
there is a clear opportunity for better video experience online but I expect that 2021 will be the year that everyone runs screaming for the classroom and office.
The best managers guide that evolution, but recognise that different people work in different ways.
But video conferencing/meetings is a capability to be used, and I look forward to seeing how it improves. Rejecting it out-of-hand simply because it's currently sub-optimal and badly used seems short-sighted.
Having a real sound card like an entry level Focusrite 2i2 or Solo, a sub £100 condenser mic and decent headphones would make a lot of difference
I found one called "onespace" in my search now but there are quite a lot of them.
The creator of Nim was involved with one called "3dicc".
The are multiple recent top-down localized audio ones.
It isn't quite ready for a show HN, but I do have an early access program if your interest is peaked.
Here's a link: https://weevee.tv
I have never understood the problem people have having video active - id describe my self as fairly introverted and have no problem with this, in fact I do between 4-6 hours a week streaming DnD
Pick a popular video chat application, have you and someone else on the same wireless network both join the same chat, go into separate rooms to avoid crosstalk on the mics.
Have one person talk, then realize there is a half a second lag.
It doesn't help that many bluetooth headsets add up to 200ms of latency (!!) https://www.rtings.com/headphones/tests/connectivity/bluetoo...
We'd be better off using analog telephone lines and 900mhz wireless headsets from the 90s.
I don't know if video calling tends to multiplex audio and video, or send separate streams; multiplexing audio onto the video could help reduce audio delay though. Standard voip uses a packet rate of 50 pps for audio, which means at least 20 ms between sampling and sending. Reducing that would be nice.
Switching to Peer to Peer for 1:1 or small group chats would be nice.
Low latency requirements being part of the BT spec would also be nice[1].
OS's fixing their drivers and entire end to end audio stack, that'd also be nice. Get that end to end latency locally to less than the e2e latency of the network packets!
[1] Qualcomm has a low latency version of apt-x which is actually low latency, and Apple does fine with Airpods connected to Mac hardware.
That is... so bad. How? Did they let the same people who designed Bluetooth LE's ANCS[1] go hog wild with the desktop stack?
Edit: Also I wish that site tested latency for phone calls, since that is often a different protocol.
[1] ANCS is, or at least last time I looked at it, a protocol designed by people who had obviously NEVER worked in the embedded or low power space. No limit on message size and no headers indicating the size of the incoming payload are two giant tells.
There are two important differences with VC compared to that ping: You're not trying to exchange data with a CDN built for exactly that purpose, you're trying to exchange it with someone else's (probably crappy) internet connection. The analogy I used the other day is "your supermarket has better roads to it than your friend's house".
You're probably pinging 32 or 64 bytes. Now try that again at a more realistic video data rate (not that the CDN endpoint will allow that). Keeping latency down is a lot easier on small packets than on full pipes.
Even cabinet meetings
It's built on Jitsi.
Maybe "Gaming" suits some people, but not everyone.
<sarcasm>
Imagine if there were a publicly available system that could switch connections directly between the end users, without the need for computers, or anything wireless? It could send uLaw companded audio in a 8,000 samples per second, to cover the 3 KHz bandwidth of human speech, with almost no latency. You could even hook circuits directly up to the end user equipment to avoid any need to share the audio channel. Because the bandwidth was dedicated on a switched channel, there would be no drop-outs or packet loss!
What an amazing improvement over cellphones and Bluetooth that would be.
</sarcasm>
I believe we can transcend that threshold by adopting a shared canvas for structured information and systems.
Suppose my Fermi math is somewhat plausible when I say that the annual aggregate of all human food consumption is about 34 peta-calories and that we use 24 zettajoules of energy. What if we the people could collaborate to make that 50 pCal and 19 ZJ 5 years from now, while flattening the per capita distribution? Religion, privately-controlled social networks, and democrapitalism aren't likely to carry us there alone. Humans are good at using, creating, and distributing tools and that makes us highly adaptable. They are good at pooling resources to tackle large objectives.
Here are components I believe are necessary to build an machine for transformative virtual collaboration:
- Simple primitives that transcend literacy and align with human behavior and cognition [2]
- Tactile consumption and manipulation of data and business logic across many surfaces [3]
- Simultaneous replay of media and event data from disparate sources through a multitude of lenses [4]
- Immutability + Time Travel for all data/transactions
- Arbitrary auctions against time, resources, data access, and more
- Ubiquitous simulation, modeling, and learning systems
- Rapid, contextual, stake-driven agreement systems for contract formation and decision-making.
- Zero-trust, edge-first privacy model
- Containerized knowledge work - like each task you pull from your queue gets its own OCI container and X session with appropriate tools, files, keys, and roles. Pausable, snapshottable, sharable, replayable, work.
- Sortable task queues that align with personally defined values, priorities, and constraints.
- Personal Optimization aligned with Global Optimization
- ... better looking video :P
Such a system would work to remove friction related to misunderstanding, misinformation, and misalignment. It would empower and enliven individuals to do work that is truly meaningful and impactful where needed most. It would scale our ability as a team to sense, think, react, and observe. It puts humans and facts in a place of primacy, and allows for async work to... work.
Thank you for coming to my TED Talk.
I've been trying to unpack and realize this vision for the past 5 years. I've got mountains of notes and some meaningful progress on the data and compute infrastructure, business and growth models, some lenses, and an obscene amount of bunny trails into geometric UI paradigms (isomorphic hexagonal tiling, aperiodic tiling and armand bars, ZUIs, and inevitably ur patterns). I've also got a family, day job, and a needy old house. I want to make deeper progress, faster, and in collaboration with others, but I have no idea what the right venue is to pursue this work. Where are the people who are working on this sort of thing? I don't have much of a pedigree for academia, and I'm not in a position to take much personal financial risk. If you have any ideas please let me know.
[1] https://fs.blog/2019/01/yuval-noah-harari-dominate-earth/ [2] https://www.8ways.online/ [3] https://worrydream.com [4] I define lenses as views, filters, renderers for any device, where the audio/visual output streams of a lens are a function of the data itself. Those could range from 2D rectangle screens to 3D worlds, AR/VR, watches, voice assistants, raymarching projectors, haptic holograms, auric interfaces, smart dresses, whatevs.
https://www.bloomberg.com/news/articles/2021-01-04/south-afr...
We don't need more video conferencing apps, but the end of the pandemic is mostly irrelevant.