Show HN: Talk – A free group video call app with screen sharing
github.com
github.com
Ended up deciding to spend the week configuring jitsi, but all it took was 1 hour to install and configure the server and I was done
The development of the frontend is a bit messy, but you'll manage
Or can you?
The NAT hole punching is done by the STUN servers listed in script.js. They appear to be public third-party STUN servers, so that could be a vector for a malicious actor. There are also third-party TURN servers listed, which will relay media in the case that NAT traversal fails. That should be ok too, but could also be another attack vector.
It's based on this open source app https://github.com/simplewebrtc/simplewebrtc-talky-sample-ap...
They also run https://www.simplewebrtc.com which is an SDK for building custom WebRTC apps. They also provide TURN and SFU servers.
And,i dont quite understand this bit:
> To get started, you will first need to edit public/index.html to set your API key.
So, for any service using talky, I can just steal the Api key rather than subscribing? I mean, it's in the index file, not even protected by login, but sent to all clients (and bots..)?
Also open source. Discussion: https://news.ycombinator.com/item?id=23523830
I created https://omen.tv/ just last weekend. Similar in that it's powered by WebRTC, but it's designed for casting your screen (e.g. Jackbox Games) to other people's TVs.
My greatest annoyances with WebRTC were:
1. WebRTC requires a STUN server, and despite the spec initially supporting default ICE servers, it has since been pulled out into an extension because browser vendors don't want to provide servers[1]. There are free STUN servers (Google etc.) but...
2. Double NAT clients. STUN is inevitably going to discover double NAT clients. The only way to connect these clients is through a proxy. Specifically a TURN server. Unlike STUN servers, these are not light-weight, and I can totally understand why browser vendors don't offer them by default. So inevitably any WebRTC use still requires a self-hosted TURN server.
3. Inconsistent access to streams across browsers. In particular the getDisplayMedia[2] API provides poor availability of the audio stream. Courtesy of Apple doing Apple things, I don't believe it's even possible to implement this API on macOS as there's no audio loopback device. I worked around this for my use-case by installing BlackHole[3] (which is great software!) However, the loopback device appears as an input e.g. like a microphone, so isn't technically the display audio.
For this all to be smooth I think all these pain points need to be addressed.
Issue 1 and 2 require NAT to be kicked to the curb, come on IPv6! Issue 3, requires Apple to expose a loopback device and browsers to implement support. These aren't insurmountable issues, but they're unfortunately out of our hands as day-to-day devs.
[1] https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConn...
[2] https://developer.mozilla.org/en-US/docs/Web/API/MediaDevice...
The identity and access management is non-existent at this point of the project’s roadmap. A client needs the hostname of the Talk Node.js server along with a unique “room” name to connect, as far as I can tell after a quick glance of server.js [2].
[1] https://github.com/vasanthv/talk/blob/master/www/script.js
So, to sum up, for a webRTC p2p app you need to:
1-3) The things you mentioned.
4) Code your own signaling server to transmit events like "call", "hung up", "room" tracking, etc.
This is not what double NAT means. If both people are behind single NAT, STUN is enough.
I think it is also enough if one person (or both) are behind more than one regular NAT, which is what double-nat usually means.
Where STUN doesn't work is symmetric NAT [1], where different destination servers will receive a packet with different source ports, even if the original source port was the same.
It also doesn't work in a few, lesser-used NAT types [2].
[1] https://en.wikipedia.org/wiki/Network_address_translation#Sy...
Just off the top of my head, the only way I could see this working is if the NATs themselves aren't transparent, and offer an API to manually map address/port pairs.
So, you bind 192.168.100.55:5555 and send a packet to a STUN server 1.2.3.4:3478. This creates a mapping between, 192.168.100.55:5555 and an outside port, we'll call it 81.82.83.84:12345. And the STUN server tells you back, in a reply, "you're 81.82.83.84:12345".
Then your friend sends a packet from his private address to 81.82.83.84:12345. It reaches your NAT, and is translated to 192.168.100.55:5555 and you receive it. Replying to this packet lets you send packets to him.
I was incorrectly assuming that all NATs dynamically assign a new external port per destination IP-port tuple. As this is what I've observed in practice.
However, it looks as though many believe this only occurs for Symmetric NATs, and never Cone NATs. It seems the common terminology just isn't nuanced enough to explain what's really going on. Wikipedia[1] actually has an interesting paragraph regarding terminology:
> Many NAT implementations combine these types, and it is, therefore, better to refer to specific individual NAT behaviors instead of using the Cone/Symmetric terminology. RFC 4787 attempts to alleviate confusion by introducing standardized terminology for observed behaviors. [...] Specifically, most NATs combine symmetric NAT for outgoing connections with static port mapping, where incoming packets addressed to the external address and port are redirected to a specific internal address and port.
The CGNATs I've encountered, which perhaps aren't representative, are mobile network CGNATs. By observation they were behaving like Restricted (either Address or Port) Cone NATs for inbound packets, but like Symmetric NATs for outbound packets. Hence, the source of my confusion in believing all Cone NATs exhibited this behaviour.
EDIT: I'm admittedly poor with the terminology as it's been years since I read the RFCs. I was only now able to express the above after doing a lot of refresher reading. I was coming at this based on experience implementing custom UDP P2P protocols, where I mostly just care about the worst case scenario; and had evidently assumed it more common than it is.
[1] https://en.wikipedia.org/wiki/Network_address_translation#Me...
A specific Nat setup would be my home connection, or pretty much anyone using the same ISP. It's CGN and then another Nat at home. Stun works fine with that.
I think the biggest issue is that often you have to set up a STUN/TURN server in an extra environment. In the XMPP world, ejabberd has come to the point, where it brings most things you need today in a single package (XMPP, HTTPS, Lets-enrypt, STUN/TURN, ...). That way you add about 5 lines to a config file and get STUN and TURN with little effort. I like that simplicity and if you don't, you are still free to use an external STUN/TURN server.
My apologies if it reads that way. They are indeed standardised web technologies used in a whole host of environments. For example, Valve run STUN servers for their (now deprecated, but very much still in use) Steam P2P APIs; which are backed by a custom UDP protocol.
The STUN one is easy really but the TURN servers need to be able to proxy traffic. Twillio provides STUN/TURN servers and they are fairly cheap. I am sure there are others.
IPv6 was designed back in the 90s, where we had much less concerns about tracking, privacy and anonymity. But IPv4 is even more ancient, and, make no mistake, CGNAT is not designed to make you more anonymous, so you could be fooled by a false sense of security.
Maybe we need to sell IPv6 as a way to track users to boost its adoption? And develop those privacy-conscious networks on top of it? Regardless, IPv4 and NAT are not consumer-friendly when it comes to self-hosting and p2p, so as long as these are the norm, software silos will be at an advantage.
I didn't see anything about encryption based on a cursory read. Does WebRTC have some built in or is this unencrypted?
The signaling for call setup / maintenance / teardown can be whatever, though. It looks like this uses socket.io with HTTPS long-polling. So you need a publicly-addressable IP (or a proxy service like ngrok) in order to host this yourself.
Assuming you host it yourself, all clients should have encrypted signaling with your server, and encrypted voice/video packets between each other. If someone else hosts it for you, they can see your signaling traffic, but -- unless they change the code to forward the media stream elsewhere in order to MitM it -- your media should still be encrypted.
However, this does make use of public STUN servers to help with NAT/firewall hole punching, which will leak your IP addresses and possibly make it so an adversary can figure out the IPs of the people who were talking to each other, based on the timing of connections. There are also public TURN servers listed which will act as media relays if the NAT traversal fails. Usually the TURN servers are dumb packet forwarders and won't terminate the DTLS sessions, so that should be fine, but I don't remember how the protocol works well enough (it's been a good 7 years since I was knee-deep in this stuff) to say that for certain.
DTLS is however fully used for SCTP, which the protocol for data channels alongside the media streams.
I often host online lectures and with screen sharing, it always uses up 100% of my CPU resources.
Interesting experiment would be to run locally on high speed LAN with LOTS of participants. Whats the limit to what the browser can handle?
Thanks for building, vasanthv ;)
Only problem I had with it was one client that couldn't get his mic working. I'm assuming it was a browser permissions issue but I called him on Signal instead rather than trying to track it down remotely.
Did not know that.
There's now support for third-party auth (i.e. oauth). It's not really fool-proof on its own, however you can at least then disable access to those who abuse the system. However, for this to work you need to have an oauth provider i.e. sign-in, which may be non-desirable.
Might just be an example.
Disclaimer: I am coworker with the people that write OpenVidu. Check it out!
From a non-technical persons perspective, my piano teacher was using Zoom but ended up having to switch to Jitsi Meet because (anecdotally) the audio quality was better and more consistent.
> The sweet number is somewhere around 6 to 8 people in an average high-speed connection.
Obviously I haven't audited the code, but it's pretty small, and wouldn't be hard to thoroughly audit.
It's surprisingly light on dependencies, just pulling in express as the webapp server, and socket.io for the call signaling. Personally I'd probably roll my own signaling over websocket to avoid that dependency (socket.io is an awful protocol, though for generally understandable reasons), but that would likely more than double the amount of code the author would've had to write.
Here's an article I found, with opinions on the good and the bad.
https://dzone.com/articles/socketio-the-good-the-bad-and-the... (2018)
Documentation is sparse, and when you find some docs, half the time it isn't clear which (incompatible) version of the protocol they're talking about.
Source: I was working on writing an interoperable server implementation of socket.io a couple years ago, and it was way more work than it should have been.
(P.S:: Didn't check out the source code. It's just a general sentiment of mine lately.)