Peer-to-Peer Communication Across Network Address Translators (2005)
brynosaurus.com
brynosaurus.com
We found that renewing the NAT entry at a rate of once per 20 seconds was good enough. This made it work through all the routers and gateways we came across.
EDIT: SIP is short for Session Initiation Protocol. Its the protocol used to connect VOIP phone calls. We used UDP. SIP's standard port is 5060 and 5061. We had to switch to port 10,000 (any port other than the standard) because too many high end routers and gateways would manipulate our packets. These routers were trying to implement SIP through NAT on their own. I applauded their efforts but it caused problems for us. I still remember the day that I took it upon myself to change production to use port 10,000. I was really scared I was going to break something but everything worked out in the end.
I had always wondered how you guy's were doing it, but I had always thought that the router Vonage handed out was something special that used its own reserved ports or something that wasn't easily visible. I never looked into it deeper. It's interesting that instead something like this "hole punching" technique was used.
I've heard about the technique before, but never investigated it before...
- SIP end points register every so often. Each registration had a TTL. We set the TTL to 20 seconds. This actually leads to a registration every 16 seconds, since SIP end points register at 80% into their TTL.
- When a SIP end point reigsters, do not register the end point at the IP:port that it is advertising int he SIP Registration message. Instead use the source IP and port of the IP/UDP packet/datagram.
- Inbound calls to that SIP end point must go through the server that the SIP end point registered with and it must use the destination port (10,000) as the source port. We need to match the NAT entry in the router.
- When a phone call comes in for that end point, send the SIP Invite message to source IP:port of the SIP end point. Since the device is registering every 16 sec, there is sure to be a NAT entry in the router to allow our packets into the network and routed to the SIP end point. It is vital that we use the same server that the SIP end point registered with.
Another "fix" to SIP - a protocol that works just as successfully over NAT as ftp (mmm active or passive? 1:1 NAT bodge/fix? etc etc) - is to mix the control and data into one stream: eg IAX(2). IAX uses a single well known destination port (4569/udp) which is as easy to implement as DNS or NTP, which are both generally UDP but DNS can require TCP, through a NAT. IAX can make a fine
Sometimes folk decide to run SIP over TCP instead of UDP and may even decide to run that inside a TCP tunnel eg OpenVPN over TCP. That's fine until you discover how TCP inside TCP can result in exponential standoffs when network conditions are sub-optimal. That can really bugger up voice traffic that really needs sub 100ms point to point latency to be acceptable. Then there is dealing with packet fragmentation if changes of media require it and of course the fact that TCP packet streams can arrive out of order and be re-assembled. That latency wont worry your web browser but it makes voice sound horrid.
I don't think NAT had an awful lot to do with that.
For anyone interested in thorough analysis instead of opinions I recommend this really good article from 1994:
https://github.com/papers-we-love/papers-we-love/blob/master...
You might enjoy this for a laugh, though (Aug. '96 vintage).
> The main progress has been WebRTC
Yes, it is a kind of combination of the best approaches to NAT traversal from the SIP world, supporting both STUN & TURN and ICE with the Trickle addon.Some very good references on the topic I found recently: https://webrtchacks.com/ and https://www.html5rocks.com/en/tutorials/webrtc/infrastructur....
The techniques were used in many contexts at the time.
CORBA "et al" lost because they are fragile, not because of network issues. SOAP, for instance, is no more difficult from a connectivity perspective than REST; it rides on the same TCP/HTTP rails as REST. Yet hardly anyone, given the choice, will choose SOAP today.
Not that REST (really just HTTP APIs, mostly) is all that robust; it just asks far less of everyone involved to specify, coordinate and troubleshoot. You can hit an endpoint with your browser, pop open the browser developer tools and figure out what is going on. That _easy of use_, which CORBA, SOAP, "et al" completely forego, lowers cost, speeds development and causes an order of magnitude less nausea for developers. The market has proven the value of these benefits, despite what holdouts wish to believe.
We have also implemented that in our game OpenLieroX. We call it UDP NAT traversal if you search for it in the code.
Main code: https://github.com/albertz/openlierox/
UDP master server: https://github.com/albertz/openlierox/tree/0.59/tools/UDPMas...
WebRTC just uses hole punching, and its implementation (ICE+STUN+TURN) is as over-engineered as you would expect from a modern protocol stack.
1999: http://alumnus.caltech.edu/~dank/peer-nat.html
I have seen reliable, fast overlay networks accomplished with a relatively small amount of code. As a result, this is where I set the bar. The more voluminous the project, the more quickly I lose interest.
There's the standard ways -- hole punching as described there, uPnP/NAT-PMP, etc. -- and then there is a long tail of hacks for getting to that 90%+ success rate you want. Here's an incomplete list:
* Port prediction for symmetric NATs that increment ports
* Creating multiple sockets to work around buggy NATs that can't hole punch if multiple devices have the same local port, etc. There are also buggy NATs where hole punching breaks if you also open a port with NAT-PMP or uPnP. There are a lot of buggy NATs. (Factoid: we have found zero correlation between the cost of a NAT device and its buggitude.)
* Getting the timing as tight as possible -- we use a triangular three-party rendezvous method to try to get it so both sides see a send before the "reply."
* Sending UDP packets with incrementing TTLs ahead of the real hole punch message to open multiple layers of NAT, and randomizing the few-byte payload of those low-TTL packets because some NATs care about this. Yes people plug NATs into NATs into NATs. It's pretty common. Ugh. As long as none are symmetric it works.
* If at first you don't succeed try try again... forever. For long-lived links (common with ZeroTier since they are virtual LANs) if you keep trying and include some random port trials you will eventually punch a symmetric NAT. There are only 65535 ports.
ZeroTier still has to relay though if both parties are behind hostile NATs, but we can provide free relaying for everyone since it's only a small percentage of traffic and cloud bandwidth is cheap.
Edit: there are more sophisticated methods that work more of the time, but those unfortunately can trip intrusion detection systems. We try to walk the line with ZeroTier between high success and not tripping your IDS.
Edit #2: wow there's a lot of outright nonsense on this thread. I continue to be saddened by how few developers (even top notch ones) know so little about networking and have so many misconceptions about it. Networking is still such a black art. Of course if it wasn't maybe there would not be a niche for us. :)
The client enumerates all the machines it can see on its local LAN, and for each one, assembles a list of information about it, probably something like IP address, machine name, assigned user, maybe MAC address, etc. Hash this using SHA256 or something (unlike the suggestion in the ZeroTier article about IP addresses, this approach has enough information that brute force reversal is out of the question). Assemble a list of, say, 20 of these hashes - the client's own hash is first, followed by the rest of the LAN. If you have less than 20 nodes, fill out the list with random numbers; if you have more than 20, include yourself and 19 others at random (it's important that you pick them at random, not just the first 19 in your LAN listing). Send the list to the server - this reveals nothing about your LAN to it.
The server compares the hash lists sent by clients. If any two lists have any hashes in common, send a message to each of the clients that sent those lists: "You seem to be on the same LAN as the client with hash XYZ." The client can then look up the hash in its local table, and try to contact the corresponding node. (The list intersection check is why it's important to pick your sample at random - the birthday paradox is now working for us, and there's a good chance of at least one match even if your LAN has hundreds of nodes.)
https://hn.algolia.com/?query=pwnat&type=comment&sort=byDate
It doesn't deal with the "long tail" of NAT devices that increment/decrement ports in (somewhat) predictable manner and _most importantly_ it describes hole punching as a client-driven process.
The latter is the crucial point. By tasking a dedicated (mediating) server with coordinating the punching sequence it becomes possible to time the process much more precisely and to help predictions to actually match the reality. Combined with a bit smarter port predication it brings the success rate from 80% to 95-97%... or at least it did 10 years ago when I was using this for Hamachi P2P VPN, though I suspect that very little has changed in terms of NAT type distribution since then.
[1] http://copilotco.com/mail-archives/p2p-hackers.2005/msg00126...
What if the private IP from the other peer is in another NAT, but in your local NAT, you have another peer with that same private IP? He would answer you, and would would establish communication with the wrong peer?
Probably would need an extra step to validate the peer's public IP address is also the same?
As brilliant a hack as hole punching is I hope someday we can move to a world of IPv6 where it is no longer necessary.
UDP hole punching is when both parties are behind NAT and it's used almost exclusively for P2P: torrents, VoIP phones, video chat, multiplayer games, ZeroTier, WebRTC, IPFS, etc.