How to create a video call application with WebRTC
blog.phuaxueyong.com
blog.phuaxueyong.com
The last time I did any p2p networking was back in 2002 or something when you still had to do it all manually. We used all sorts of fun tricks like NAT hole punching, and using little script endpoints to capture and forward along port and public IP address information.
It was fun to see that all of this has since been formalized under the "ICE framework". I was surprised to see that the STUN spec is only 12 years old now, despite the techniques involved being used for at least 20 years, probably more like 30+.
So if anyone who's new to this whole p2p world feels that WebRTC and the ICE framework is confusing or onerous, I would point out that just a short while ago these were basically just a handful of heuristic techniques developed through trial and error over the years. It's really much easier nowadays! zonko.chat only took me 12 or so hours to build (and seems to be well-supported by chrome and ff, even mobile).
Edit: Upon reflection, I don't even remember how I learned about some of them. The concept of TURN was probably one that I, and many thousands of others, invented from scratch due to necessity (failed to punch the hole? fall back to this custom relay I wrote in perl). STUN was an easy one to figure out yourself, too. I don't remember how I learned about hole punching though. Probably a forum or a book. Or possibly just an experiment ("what if the two connections touch somewhere in the internet at the same time... hey wait, it worked?") What's interesting to me is that the core "ICE" concepts (hole punching, STUN, TURN) are still pretty simple even in their mature, formalized, scientific form. But the concept of "SIP" is much more sophisticated today than it was back then.
> It's really much easier nowadays! zonko.chat only took me 12 or so hours to build
While it may have only taken you 12 hours to build in "real time" I'd say you've been "building" it for the last 20 years. If a newbie tried to do this they could expect to spend a few weeks or more on the project would be my guess.
The bulk of that 12 hours was actually spent debugging negotiation and timing/race issues in the signaling/SIP layer. Even for me, figuring out how the WebRTC API is supposed to work was a little difficult.
Often veterans here will talk about projects they trivially did. but unless you dig into their profile and find out who they are it feels like every joe-shmo on hn is a 1/2mil+ SWE at google.
My point is that a whole 75% of the time I spent on this project was just me flailing around with stuff that wasn't working as I expected (that's relatable at all experience levels!). Perhaps it's true that veterans can get through certain things more quickly than novices can, but we're not immune to the 80/20 rule either! We just get stuck on different types of problems.
The article mentioned "This section we will just touch and go about when do you need a TURN server. It is not needed in all situations but a component needed if you have to deal with slightly less straightway use cases." a TURN server is a must in the real world...
*https://developers.google.com/talk/libjingle/important_conce...
Solutions like Jitsi Meet[3], whereby (formerly appear.in)[4] use WebRTC to great success and are fantastic for quick meetings.
WebRTC these days is pretty mature and ready for primetime -- TURN servers are a last resort, but they're a small price to pay for something that might be free (to you the service provider) most of the time, if signaling succeeds.
[0]: https://github.com/jselbie/stunserver
[1]: https://github.com/enobufs/stun
[2]: https://gist.github.com/mondain/b0ec1cf5f60ae726202e
[3]: https://meet.jit.si
[4]: https://whereby.com
This “last resort” is about 10% he says
This means products which help alleviate WebRTC infrastructure such as AWS Kinesis are not allowed (due to how they allocate turn servers with unknown IP addresses) and a company needs to manage their own infrastructure / TURN servers (which allows you to cherry pick where server locations are (HIPAA, country legal for what is streamed)) or accept Twillio's, or their competitors etc, large IP ranges (and don't have server location flexibility / increased commercial and market growth restrictions).
Whichever route you go down it is quite an undertaking!
P.s. Tsahi Levent-Levi is truly exceptional in this area. I highly recommend reading his blog and training courses: https://bloggeek.me/, https://webrtccourse.com/, AND he runs an amazing testing product https://www.testrtc.com. if you build your own infrastructure testRTC is a must.
And it would just be STUN between each participant and the SFU deployed in the internal network for example.
Someone contributed a really cool tool to Pion stun-nat-behavior[0] that I use a lot. It prints out NAT details. It uses the modern/correct details. I see a lot of docs that still use symmetric NAT etc.. RFC 4787 [1] suggests against all that.
[0] https://github.com/pion/stun/tree/master/cmd/stun-nat-behavi...
We've built a push-to-talk walkie talkie system called Squawk[0] which holds long lived webrtc connections in the background throughout the day. We use simplepeer[1] as the base to help bootstrap some of the browser shimming, but it's not perfect. So ultimately we've had to build all sorts of checks into our protocols like an audio keepalive where we send periodic frames (20ms) of silence down the media channel, and verify that we received some additional header bytes on the remote end, because otherwise webrtc would let the connections rot and you wouldn't know until you needed them which in a push-to-talk situation is too late.
Especially, how do you maintain all connections so users can listen to all channels at once.
Not allowed to browse Newly Registered Domains category
URL: https://app.squawk.to/
This is perhaps because your whois records are bare minimum! $ whois squawk.to
Tonic whoisd V1.1
squawk noah.ns.cloudflare.com
squawk kim.ns.cloudflare.com
You might want to look into that.Also, to increase adoption, perhaps give the users a link to directly use the app.squawk.to URL to use as well. I tried that on my personal device and it works as advertised.
- janus[0][1]
- mediasoup[2][3]
- Medoze[4][5]
It's never been easier to start your own video streaming platform.
[0]: https://janus.conf.meetecho.com/
[1]: https://www.youtube.com/watch?v=zxRwELmyWU0
[3]: https://www.youtube.com/watch?v=_GhdFOZTWTw
I really think it has gotten way easier in the past couple years. I am biased though, don't want to think my effort has all gone to waste :p
[0]: https://www.kurento.org/tags/transcoding [1]: https://mediasoup.org/faq/#does-mediasoup-transcode [2]: https://janus.conf.meetecho.com/recordplaytest.html
https://github.com/peer-calls/peer-calls
https://www.irif.fr/~jch/software/sfu/
Out of those, it seems Jitsi and Mediasoup to be better as SFUs than the rest because they have what appears to be decent-looking congestion control, bitrate allocation, and support for simulcast. The rest apparently do not (at least I couldn't find anywhere in the code where it happened).
Also interesting to note is that so many of the newer ones are written in Go based on Pion. If Pion ever gains the ability to do decent congestion control (perhaps based on transport-cc like Jitsi and Mediasoup do), that could improve things for all of those.
Here's a podcast interview[3] about how he did it.
[0]: https://github.com/ianramzy
[1]: https://zipcall.io
[2]: https://github.com/ianramzy/decentralized-video-chat
[3]: https://syntax.fm/show/256/webrtc-and-peer-to-peer-video-cal...
https://www.twilio.com/blog/2014/12/set-phasers-to-stunturn-...
This has meant NAT is less of an issue for native IPv6 endpoints, including P2P.
Hopefully when IPv6 is finally widespread in US/Europe we will see stuff taking more advantage of this fact.
The audio quality on zoom is just terrible no matter if you disable DSP or not.
So many yoga classes require high quality music.
It's frustrating that chaturbate provides top notch video and audio quality for free essentially, while paying $20/mo for zoom gives you what looks like 380p video quality and audio quality I have yet to find a poor comparison for...
Does anyone know how one could emulate what chaturbate does?
Any good articles outlining how they do what they do?
Ideally, the teacher would just plop their phone down in front of them, hit broadcast, and a few seconds of buffering later 1080p video and quality audio would be visible through a browser.
Why is that so tough to do??? I haven't been able to find a single article that simplifies or distills it at all.
I suppose I could dig through the code and disable anyone but the host's video feeds but I don't have a lot of time to dedicated to this project unfortunately.
I just try to help them out when I can.
Also a lot of audio codecs are tuned towards speech and filter out high frequencies. You should pick one meant for music.
Give it a try.
Start playing any form of music through your phone.... mp3... youtube... spotify app... and then open a zoom meeting.
It distorts the sound of the music so much and I can't figure out why they would do that or what purpose it serves.
Probably just poor coding.
That said, the downsides of HLS are potentially higher infrastructure costs required to transcode video to the different qualities and somewhat related is the higher latency to live. With proper tweaking you might get 2-3 seconds of latency, but it might be too much for your use case.
If there is voice interaction between the yoga instructor and their students the HLS delay will certainly be noticeable.
No, WebRTC is 1 to 1. Each connection is adapted independently. But, you can build services that have rooms with more participants, then it's up to you to shape the traffic as you want. If you use a central server (SFU), it can just send each peer the best they can receive, each independently from one another. It's a property of the service, not the technology.
Also youtube makes it difficult to restrict who can and cannot see the video.
Without a pay wall there is little incentive to make payment.
Some studios have done OK with a donation/pay what you can model but those donations and rapidly dwindling now in the third month of Covid.
* debugging. One users sound just doesn't work while it works perfectly for me with different machines. I am clueless as to how to debug it.
* ice. while it works, I had a hard time understanding, tracking and debugging what was going on.
* closing and restarting connections
* multiple clients in one room?
* echo cancellation. This was frustrating for users.
* Turn. Is there a tool or way to know which clients need a turn server? Are using a turn server?
I ended up guessing that getting it to be a product would actually be fairly time consuming
Closing and restarting connections is signaling layer stuff, ie your responsibility.
Echo cancellation is really supposed to be application layer and up to you as well, but I think this will probably shift to be the browser's/WebRTC's/getUserMedia's responsibility at some point.
Re. TURN: ICE is the process that works out whether a specific client needs to relay through a TURN server. The question is: do you need to implement a TURN server? The answer is: yes, you need a TURN server. If you built a P2P app that you want to work for all users, you will always need a TURN server. You can run coturn on the same box that you serve your app from. Most likely a side project will never hit the scale requiring more than a $5 digitalocean box for TURN.
And yes, it should not be a surprise that products are time consuming to build :) WebRTC is plumbing; you probably were expecting something more like Jitsi.
Echo cancellation typically can't be application layer. The APIs I've seen (Android, iOS, WebRTC), require low level latency and works best as close to hardware as possible.
{ echoCancellation: true } as a track constraint in getUserMedia works.
Echo cancellation is a pretty lightweight DSP/FIR task. Whether you do it close to the hardware (? I suspect this is not actually the case with getUserMedia though -- it is still an audio stream algorithm) or in the application layer, echo cancellation requires the same amount of added latency.
But in any case, I did say I suspected echo cancellation would shift to getUserMedia. It's not fully there yet, but it will be.
The team behind Kurento is working on this (I am part of it) for people who don't really care about all the intricacies of the standard(s), and just want to build a product on too of it. A single Docker container to deploy, and you're all set to write your app.
Still, this is a complex topic so there are a thousand ways this technology can be made easier to use and understand. And I agree with other comments about the issue of debugging, there is totally an empty space in the market for a comprehensive solution that can help troubleshooting when WebRTC fails.
Learn more about the details: https://brie.fi/ng#help Also see https://webrtc-security.github.io/
WebRTC is end-to-end encrypted by default. There is a signaling server that helps establishing the connections between users in a room, but after that the communication is encrypted. Also those TURN and STUN servers are only required for technical reasons to get peer-to-peer working. So no content is ever passed unencrypted.
That's the difference to other services like Zoom and Jitsi, where a server in the middle is receiving the video streams unencrypted and then redistributes. Although Jitsi is adding encryption support for that as well soon.
Finally we had to invent workarounds for Cordova: https://mobile.twitter.com/qbixapps/status/11564841564250398...
Did anyone set up a WebRTC that was super easy and worked rock solid with just a few lines?
We open sourced our implementation.
I am still wondering if we missed some simple solution.
I mean do you use kubernetes and stuff - what load balancer do you use (I assume sessions are "sticky"), etc
But the way you scale WebRTC is by using multiplexers between clients.
Let's say you have 2 clients. They can directly send data to each other.
Let's say you have 3 clients. Client 1 sends data to Clients 2, 3 Client 2 sends data to Clients 1, 3 Client 3 sends data to Clients 1, 2
As you add more clients the scaling requires n(n-1)/2 edges because it forms a Complete Graph.
The idea then is to use a server in the middle. That server could be 3rd party or maybe even one of the participants within the graph.
[0]: https://news.ycombinator.com/item?id=8952880
[1]: https://medium.com/@icecomm/how-launching-icecomm-on-hacker-...
I didn't find the article particularly good.
As others have mentioned, building a simple project is fairly simple. The difficulty comes when you want to scale to more than ~4 users without the app becoming unusable. Adjusting audio/video constraints to ensure that you get optimal media streams is quite difficult, also. Nevermind dynamically tweaking them!
I always wonder if there is a way to think outside of this 'box'.
I think you could also do it with the WebAudio API. If you throw up a repo would love to try and help :) having a backend makes it so much harder to deploy/maintain stuff.
I've built a number of WebRTC apps over the years. Recently, I built just such a thing as you described and open sourced it: https://www.calla.chat. I opted to build it on top of Jitsi Meet this time. It's actually advantageous that it's not through the WebRTC API because Jitsi doesn't give access to the raw WebRTC commands. But hijacking the audio elements it creates is completely doable.
Check out
https://github.com/meething/meething
dWebRTC Video Meetings MESH/SFU hybrid using GunDB, MediaSoup and Beyond!
This seems to be one of the most promising projects in the WebRTC space with support from Mozilla Builders.
WebRTC works without a browser too FYI