WebRTC is almost the only choice for low-latency video and audio inside a web browser. The open source libwebrtc [1] implementation that's in Chromium and Safari is now mature enough to be used in other native applications if you have a medium-sized engineering team and are comfortable with C++. (Again, WebRTC-as-a-service platforms often provide native libraries that wrap libwebrtc to give you easier to use full stack iOS, Android, (etc) SDKs.)
The three big challenges with WebRTC are that low-latency media is its own domain and the learning curve is steep, that scaling sessions to more than two or three people requires a lot of server-side packet routing code (you can't do pure peer-to-peer with lots of participants), and that there aren't yet the mature "off the shelf" cloud building blocks that exist for HTTP-ish workloads.
WebRTC data channels are the non-video/audio part of the WebRTC spec. My hot take: data channels are rarely the right solution to any problem description that doesn't start with, "well, I already have a WebRTC transport open ..."
On the off chance that you're still checking this thread, would you care to elaborate? Specifically, if one isn't using WebRTC for media, is there a better off-the-shelf solution for P2P with NAT traversal in native apps?
However for that remaining 5%, I have a lot to learn. Using an abstraction is great when it works, but I'm interested in going through OP's project to get a better sense of what's happening when things go wrong.
NewTek actually uses WebRTC for NDI's remote networking, and while the NDI software itself is prone to crashing and probably not usable for production, the connection to the remote system is never an issue.
I don't find the webrtc signaling and set up particularly noteworthy, but once you try to connect nodes on different networks you're pretty much dependant on some third party.
* FaceTime @ Apple https://support.apple.com/en-us/HT212619
* KVS and Chime @ AWS https://github.com/awslabs/amazon-kinesis-video-streams-webr.... Lots of security cameras and robots use it, not public though.
* Lightstream https://golightstream.com . Cloud compositing and other magic.
I also have something I am working on now that isn't public yet that is using WebRTC. Really excited to see what people build with it/what it inspires next.
It is kind of amazing everywhere you will find WebRTC. Stadia, Boston Dynamics, Zoom, Meet, Security Systems, Drones etc... It is probable that you use WebRTC in production everyday :)
When Google announced WebCodecs/WebTransport they said Zoom was involved, so maybe they will switch to that eventually?
There's a spec for RTP over QUIC [1]. It's really cool! But obviously very early days.
[1] https://datatracker.ietf.org/doc/draft-engelbart-rtp-over-qu...
I think it is worth measuring how Google's implementation works, but it is tuned for a very specific use case by a single company.
You're right that this is measuring the WebRTC.org codebase (as used in Chrome, Firefox, etc.), not necessarily the WebRTC protocol. Better URL and demo/talk videos are here: https://snr.stanford.edu/salsify
But the issue probably isn't with the bandwidth estimator or congestion-control algorithm -- you probably can't fix this by taking some WebRTC implementation and plugging in better ones. The core issues as we see them are about the architecture of the WebRTC.org codebase, and frankly all WebRTC source/sink implementations we're aware of, in particular:
(a) even with perfect bandwidth estimation, libvpx and libx264 (and, we think, typical hardware encoders) as configured are bad at achieving the requested bitrate over short timescales, meaning that "overlarge" coded frames are regularly being sent, and techniques like reference invalidation or golden/altref encoding are never (?) used to skip sending an overlarge coded frame or to retry encoding the same frame at a lower quality before sending. [At a technical level, the interface between the encoder VBV buffer model and the bandwidth estimator/CC algorithm is very inconvenient -- to have these two control loops running independently, trying to do similar things at similar timescales, isn't great.]
(b) loss recovery does not work well, and again features like reference invalidation to recover quickly do not seem to be well-used in practice (the second half of https://youtu.be/jaDelb4JnP4 makes this pretty clear),
(c) the WebRTC.org codebase is so complex, with so many modes that it can get settled in, that trying to reason about these behaviors or explore them systematically is quite challenging, and
(d) because there are several layers of buffering on the receiver-side, and because the sender-side code will change things like the camera's frame rate, it's hard to measure application-level metrics [e.g. lens-to-display and microphone-to-speaker latency] robustly in a deployment, especially across a diverse hardware or OS base. (And it's easy to get a false sense of security from the network-level WebRTC metrics that are available.)
It's probably possible to produce a WebRTC source/sink implementation that works better over challenged networks and has good monitoring of application-level latency, but it would be a big job afaik. Our work was partly funded by Google, we had a high-up Google sponsor, we gave multiple talks at Google, etc., but it was challenging even to find "the people in charge" of this cross-modular stuff to talk with them, because I think the codebase in some respects mirrors the org chart. E.g. you have video compression people worrying about the encoder (and wanting to be able to plug in libvpx, libx264, and a bunch of hardware encoders to the same interface), and networking people worrying about the bandwidth estimator and CC algorithm, and it's sort of way too late to say that the interfaces or architecture needs to be refactored or that the complexity has gotten out of control. To Google's credit, they have since driven the industry to produce standardized APIs for "functional" codecs, and functional decoder ASICs now exist (not sure about encoders yet), so there is progress being made on that front at least.
There are some quirks in getting all the signaling working but there’s now more of a standard to do that process the right way…
Anyway since the devices are on the same network latency is nearly 0…
I do have an ack and retry process … I should add logging to see how often that happens though
Edit: that’s Web Socket Server