My ISP Is Killing My Idle SSH Sessions
anderstrier.dk
anderstrier.dk
Without NAT the only 2 parties that need to know anything about a TCP connection are client and server.
With NAT you have this problem where the router now also has to keep track of opened TCP connections.
E.g if you have a router with local IP 10.0.0.1 and external IP 30.0.0.1 and you are 10.0.0.2:55000 connecting to 230.0.0.1:443 router will have to allocate a port on it's external interface (let's say 56000) and remember it (this is the key part). So the connection will look like this:
10.0.0.2:55000 <-> NATing router 10.0.0.1 - 30.0.0.1:56000 <-> 230.0.0.1:443
When router receives packets on 30.0.0.1:56000 it has to remember to redirect them to 10.0.0.2:55000.
Memory is a limited resource so you can't just have an unlimited number of these opened connections floating around. This also makes your router vulnerable to an attack where an attacker can just open a bunch of connections and never close them, making your router eventually run out of memory.
So the classic solution to this problem is to use an LRU cache. So when your router is close to running out of space you just drop the connection that has been idling the longest.
Unfortunately, a) some routers are less sophisticated and will still drop your connections even if you do keep-alives and such, b) no matter what you do, memory is a finite resource and if the router doesn't have a lot of RAM, connections will be dropped.
¯\_(ツ)_/¯
www.example.com:443
From source port 12345, and you or your isp has a firewall that blocks everything that isn't explicitly allowed (this is common in corporate networks), the response could be allowed using firewall rules such as
iptables -A INPUT -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
This has the benefit of being general, but the drawback is that the firewall now needs to track the connection, with similar consequences to the NAT example you have.
It's also more likely for the firewalls to time out connections rather than use some kind of LRU scheme. In my opinion the time-based eviction is more predictable, so I prefer it. (Of course once you run out of memory you still need to evict "live" connections)
The big difference here, though, is carrier-grade NAT. That means the firewall is not under your control and might have a tiny state table. NAT is bad enough as it is, but CGN should never have happened. It's just depressing to think about, to be honest.
Even with IPv6 many ISPs are still doing it wrong. They'll give subscribers dynamic prefixes which means having to use unique local addresses (ULAs) in addition to their Internet routable addresses because the latter keep changing. This kind of stupidity makes people at home want to hang on to their IPv4 LANs because they seem more under their control.
If only I could get an ISP like Hurricane Electric to provide me with a DSL line at home for a reasonable price. Consumer-grade ones are all hopelessly bad.
While it is true that most NAT arrangements are provided by firewalls, it is quite possible for a device to provide NAT with no other firewalling features at all, so not be considered a firewall. In this case the device would just be a router that provides NAT.
Some confuse NAT and firewalling because NAT effectively implements a default-deny-all-not-initiated-here rule in one direction which is what most home users want in a firewall.
To make it even more confusing what most people are confusing with firewalling is actually NAPT which is the specific type of NAT described in this thread. There are other types of NAT which don't require keeping track of state and which don't provide the default-deny-all-not-initiated-here rule side benefit.
Yes. I should be clearer myself as just referring to NAT this way could serve to increase the confusion.
What most people just call NAT, what is offered by simple home/office routers (or APs when not in bridge mode or similar) and phones in tethered wireless mode, is actually NAPT (Network Address Port Translation), which is a subset of SNAT (Source Network Address Translation), which is in turn a subset of NAT.
Beware when you're doing pure NAT, it doesn't always do what you think!
Oh well, yes. I agree.
I was thinking more from the point of view of the behavior it causes: essentially establishing some sort of look-up table to verify if an incoming packet corresponds to a previously outgoing one. Where, if the entry in that table gets deleted, the incoming packets suddenly start being rejected.
EDIT: okay I'm wrong. It's 8192 connections, not 1024 connections. But still ridiculously low
Throughput both ways actually gets really close to what I am paying for with this configuration, where as before with the default gateway (regardless of configuration), I was lucky to see 1/2 of the gigabit speeds I have been paying for.
I didn't know their box even had certs, or what "ONT" is. Is there like... a written series of steps I could follow?
Instructions: https://medium.com/@mrtcve/at-t-gigabit-fiber-modem-bypass-u... Github project that makes it possible: https://github.com/jaysoffian/eap_proxy
It's definitely not plug and play but I've been using this setup for a year and a half and I get my full 1gb bandwidth throughout my network with lots of hosts.
If you are stuck with the newer XG-PON hardware, it looks like you might be out of luck for now.
(I have an actual end-to-end-connectable public IP from my ISP, which from the general discussion seems like an increasingly rare thing --- they keep pestering me to "upgrade" to outrageously faster yet slightly cheaper plans with a "free router included", so I suspect they are trying to get me to give up that IP...)
The other issue is ISP provided gateways that handle authentication onto the ISP network, like ATT fiber. These devices contain the certificate/keys to gain access to the network. Unfortunately theses devices also try to be more than just an auth device/gateway. In ATT’s case the gateway also handles some Uverse/IP TV services so they don’t have a true bridge mode where they send all traffic to another device. This approach then causes issues like update downtime or NAT table issues.
Either of these issues shouldn’t be caused simply by an ISP provided router. If an ISP wants to implement either approach they will do so without your approval.
I had the same SSH dropout problem, asked my ISP[1] to switch me from CGNAT to dedicated IPv4; they did, and it's fixed.
[1] Aussie Broadband, a smaller ISP in Australia renowned for good customer service.
This is true. Your options look like:
1. Get a new ISP
2. Get a VPN that supplies you with a public IP (these exist)
3. Hope you can do whatever you need on IPv6 instead
Do you literally mean a website? Using a browser? What’s an example website that would go past that?
UDP is connectionless, but typically a UDP communication is bidirectional. This means a NAT needs to inspect UDP packets and retain a mapping to direct incoming UDP packets to the right place. With no connection information this can only be done as an LRU cache or similar.
TCP is connection oriented, and a NAT might rapidly free up resources when a connection is closed (ie, when the final FIN has been ACKed). But if there's no FIN, the NAT is in the same case as it is with UDP. Making a lot of connections without closing them fills up NAT buffers.
When you have a home NAT and a carrier-grade NAT you may get an impedance mismatch of sorts. The CGNAT might have insufficient ports allocated to your service to keep up with your home NAT, resulting in timeouts or dropped mappings. Your home NAT will have one set of mappings and the CGNAT another, and the two sets probably won't be exactly the same. This means some portion of the mappings held in memory are useless.
As a specific example, many years ago Google Maps would routinely trigger failures. Using Maps would load many tile images, which could overwhelm a NAT or CGNAT. The result was a map with holes in it where some tiles failed to load.
Browsers have long had limits on concurrent connections per domain. Total concurrent connection limits are also old news, but are not quite as old as per-domain limits. You probably can't make a NAT choke with just simple web requests (even AJAX) any more. You might be able to do it using, eg, WebRTC APIs, though I would be surprised if those aren't also subject to limits.
Just curious, were these image resources all hosted on the same domain?
How has the world changed.
https://meetings.apnic.net/32/pdf/Miyakawa-APNIC-KEYNOTE-IPv...
Slides 12-15 show the degradation of Maps in action. 20 connections per user is a heavily over-committed CGNAT, but that level of port sharing does happen.
I bet you could put 10k people behind each IP and never even get close to an issue of this type.
You can't put 10k people behind each IP and not have problems. That's 6.5 ports per person, you need one for each connection. Pretty much any website will have issues with that little connectivity.
It doesn't have to. 100:1 would work just fine. With IPs being about $25 each that's an acquisition cost of less than a dollar per user.
> That's 6.5 ports per person, you need one for each connection.
That's not how connections work. Each user could make a million connections as long as they're spread around different servers. The 65k limit applies to simultaneous connections to a single webserver. Only the most-connected server matters, so probably something at google/youtube/facebook, and even then most of those servers have multiple IPs.
A malicious website could open up 256 websockets and as many HTTP connections as the browser allows, and that might be enough to swamp cheaper NATs.
See https://bugs.chromium.org/p/chromium/issues/detail?id=12066 for some 2009 discussion about people having troubles using the web when background tabs held connections open for polling. That wasn't a NAT issue, but it does highlight that a decade or two ago we all thought we only needed to manage tens of connections for a host to be online but that rapidly spiralled into hundreds.
I've been working with NATs for years and your comment helped me "click" and understand them at a different level.
Stumbling upon a great blog post that makes something click is always a pleasant experience.
Even without NAT there may be multiple devices between a client and server which need to know about the TCP connection. Stateful firewalls, WAN accelerators, and load balancers are some examples.
set ServerAliveInterval in your ssh_config to avoid this (default uset)
I really prefer screen's functionality, instead. Each pane can be independently switched to any of the pool of underlying ptys being managed. The exact layout of panes is more of an ephemeral thing. On wide terminals I can split side-by-side, on narrow ones I can split top-above-bottom (and these can even be different for different clients connected to the server). It handles the multiple attached clients by only make
I think tmux's model is due to being too closely modeled on GUI virtual desktops.
That rant aside, tmux does have some nice features available by default that take a fair bit of configuration to get working well in screen.
I actually did try to get the behavior I wanted out of nested tmux invocations and a hairy mess of scripts. There was no 'next-session' command (though now it looks like 'switch-client -l/-n/-p' would work. Hmm.)
I couldn't get it to work right, only a half-done approximation, and it involved way too many tmux interpositions.
Screen: client -> per client view of P panes -> P of entire pool of T terminals in session.
Tmux: client -> 1 of S sessions -> 1 of W windows -> P panes
To get splitting at the top level I need one top-level tmux client and its associated top-level session managed by one server. It can actually be locked to one window for what I care about.
In each of its panes, I need to run another tmux client so as to be able to actually change what final pty each pane will display. And each of these need to be separate sessions, so that I can display separate things in each of them. Each of these separate sessions will generally only have one window, with its implicit single pane. I should, of course, run all of these inner tmux sessions as a separate server from the top one.
Now I just have to make restricted keybindings for both the top and inner clients, and make make sure that any binding for the inner clients are replicated as "self-insert" in the top-level client. And remember that there are multiple places I can direct the command-lines (so need two bindings).
This results in the top level server having: client -> session -> 1 window -> P Panes -> P clients
And the bottom server having: P clients -> T sessions -> T windows -> T panes
And this doesn't yet let you have multiple top-level clients attached the way screen does.
I'm sure I could eventually sand off all the sharp edges where it doesn't do what I want and make this work.
Or I could just use screen.
if what matters is that different clients can look at different windows, then using multiple sessions with one window each will get you exactly that.
you may of course need to get used to tmux way of switching sessions (which is not different than tmux way of switching windows). however, i just checked, there is a command to switch to the next session, and you can bind that to a key. add another keybinding to create a new session, and that should get you to about 90% of your expected behaviour without nesting tmux.
The fact that it also doesn't have my preferred behavior for multiple clients connecting is just a minor nit-pick at that point. And I did say that the misnamed 'switch-client' would likely work.
Tmux has some very nice features: a nice commmand language, xterm-style mouse support, including both event binding and sending to client (pass through only), well thought out client-server separation, the ability for other programs to fairly easily drive it, visual identification of panes, menu-popups, and the default status-line is nicer.
All of that is merely nice to have, not actually needed though. It fundamentally doesn't have a model that works well for me. I wish it did, or that screen gains such things.
That could be solved by not relying on tmux for splitting, but on a terminal that has that feature, one ssh+tmux connection by panel.
In any case, it seems that you have put a lot of thinking about this, and probably already considered this solution.
I've been in the situation of trying to make my workflow perfectly fit in my new software/os/laptop/etc and not being able to make it work 100%. It's... sometimes exhausting. Nowadays I take another approach: make it fit well enough.
well that would be like putting two terminals next to each other.
doing that inside the terminal instead has the advantage that it works on remote terminals too.
there is splitvt, but it doesn't seem to be actively maintained anymore
Pretty much! Not a perfect solution for sure.
I used to only rely on tmux, but nowadays I often prefer opening more terminals windows. I found out that by overly relying on tmux, I kept opened a way too high number of shells (browser tabs, anyone?).
Now for most tasks I open a terminal without tmux, that way I'm forced to close it if I don't want to have a cluttered desktop (I never minimize or hide windows, I don't even know how to do it on my wm).
that does make sense and i can see that's a useful way to work actually.
so how does screen do that?
Exactly. Or at least that's one way to describe it, though it might also equally describe other things.
> so how does screen do that?
It's screen's native model.
'C-a |' (split -v) splits side-by-side and 'C-a S' (split) splits top-and-bottom. Screen calls these "regions". 'C-a X' (remove) will remove the current region, letting the sibling take the entire space again. 'C-a Q' (only) will replace the entire layout with the current region. 'C-a tab' (focus) will jump to the next region (it can take directional arguments to move up, down, left, and right, as well as 'prev' to go the opposite way of 'next', but these are not bound by default).
What tmux users would normally think of as window changing commands just switch the current region in the layout between viewing the entire pool of running commands.
Yes, using a terminal emulator that natively support tmux control mode (like iTerm2) is really nice because you can easily resize/rearrange the windows and panes via GUI actions, but absolutely suck when you have to reattach using a traditional terminal emulator because you now have to adjust all those windows and panes with keyboard actions.
You generally just need to do:
set -g mouse on
... by only making the ptys that are visible in multiple clients have the same size.
defscrollback 500000
scrollback 500000
termcapinfo xterm* ti@:te@
There you go, normal scrolling!https://github.com/mobile-shell/mosh/issues/122#issuecomment...
http://web.mit.edu/keithw/www/Winstein-Balakrishnan-Mosh.pdf (Section 6)
I'm pretty happy I got over that hurdle.
Adding byobu made it effectively impervious.
Story time: I had tmux sessions active on all our servers and would simply ssh into my work laptop from home, so the sessions were never really closed. One day, I decided to upgrade from 1404 to 1604 and closed out all my sessions (including on the servers because I was pushing out a new tmux config). After 5 minutes we started getting smss that the system was down and couldn't write to disk. One of our production servers had been set up with an encrypted home folder and when my session closed out, it closed the encrypted folder. Unfortunately, the ssh folder wasn't outside the encrypted portion, so we had to use IPMI to restore access. That's the story about how we started joking that closing my laptop is a great way to break the production system.
I also use it when I know I'll need to finish something on a different machine.
You're right about the keys if I'd be syncing them. But for my specific use case I happen to use passwords for these servers.
[Edit: ] Or in this case, an ISP in Denmark that is trying to minimize ipv4 cost by using LSN (carrier grade nat) which also has many other drawbacks.
A SSH session does not generate any traffic
This does not have to be true. You can enable TCP keepalive in the server and client configuration.
Client via ~/.ssh/config:
TCPKeepAlive yes
ServerAliveInterval 60
ServerAliveCountMax 2
Server via /etc/ssh/sshd_config: TCPKeepAlive yes
ClientAliveInterval 60
ClientAliveCountMax 2
Why are the TCP keepalives only sent after 2 hours?Each OS has a default time set for keepalives. If you do not specify it in the ssh config, it will use the OS default. In Linux, you can set this in /etc/sysctl.conf:
net.ipv4.tcp_keepalive_time = 60
net.ipv4.tcp_keepalive_intvl = 60
net.ipv4.tcp_keepalive_probes = 2
After adding this, run sysctl -e -pYou can see the timers on your established connections with:
ss -emoian | grep tim
Note: TCP timers are not the same as ssh client and sshd server tcp keepalive packets. These are two distinctly different mechanisms that can accomplish the same thing. Not all applications support TCP socket keepalive. You can wrap applications with a library called libkeepalived to add support without code changes by using LD_PRELOAD.In Windows this is set in the registry [1] On mac this is set via sysctl similar to Linux.
After you have adjusted your client and server config, restarted sshd on the server, then ssh to your server using the flag -vv and you will see the keepalive packets.
[1] - https://serverfault.com/questions/735515/tcp-timeout-for-est...
I usually just do client keepalives as they are easiest to set up. Server keepalives are good if you are worried about “forgotten” clients. TCP keepalives are usually not worth it IMHO.
This way you solve it where the issue occurs, and with the added benefit that it works for all TCP connections, not just SSH.
However I haven't had this issue. My isp is pretty ok in this regard and I supply my own router. So I don't know if there's issues with this in real life.
Making my ISP fix the underlying issue - that their TCP connection idle-timeout is too short - will make sure all their customers won't have to encounter this problem.
For wired connections I think its only the small newish ISPs + stofa that does CGN, the rest like tdc and telenor provides IPv4 to the CPE equipment.
I have hiper, they do CGN by default but if customers ask for it they can get a dynamic IPv4 for free or a fixed one for a small fee.
I had the same problem and did the ~/.ssh/config trick years ago. Interested in contacting my ISP so that they fix the problem for all users (although it might be fixed now, idk).
I've had YouSee since I moved here, and I have a single public IPv4. I didn't realise that was not standard.
Here in Finland, the situation is similar; when using mobile broadband you usually end up behind CGNAT.
Luckily, most ISP's will happily provide a static IPv4 for you for a small fee.
Any ISP using LSN will have low NAT timeouts because it takes memory on their routers to track sessions and state. I would be surprised if your ISP remove timeouts unless they are letting it fall back to FIFO pruning on your segment. Did they tell you what they are changing?
For the rest of the customers that don't pay extra for a public IP, all the crappy things you mention do apply.
Hopefully, the ISP does native IPv6?
And, while 60 minute timeouts violate the RFC, it's a whole lot better than I expected. Usually CGN timeouts are around 15 minutes for nice ones, and I've seen 10 seconds at the bottom end.
I wish the longer ones would probe both ends of the connection to see if it's still live a minute or so before they intend to kill it.
In the end, it's an imperfect solution for a real problem that mostly works well enough.
If you care about that you probably should use mosh as that does solve that by design and not by random chance.
On the other hand using VPN with fixed tunneled endpoint IPs causes idle TCP connections going through it to remain connected pretty much indefinitely.
$ screen -x
If you don't do this or something similar, then perhaps you should start now.Tried tmux a few times, but I'm not gonna re-learn everything every 15 years, come on! Also luckily when I switched jobs all admins used and installed screen everywhere, so that was an easy fit.
I am actually shocked that: a) the ISP has an email or any sort of asynchronous communication. Most in the US have at best a "talk to a bot" functionality b) they acknowledge the behavior and did not simply respond "have you tried rebooting the router?"
EDIT: Also see Mosh (https://mosh.org/)
I spoke to Hyperoptic, they said I could have a fixed (non CGNAT) IP for five months for free and then £5 per month thereafter. After five months I noticed they had started charging me £1.25 per month. I am not sure if that will increase to £5 at some point, but either way, I am happy to pay to have a "proper" internet connection.
I share this as perhaps others will be in similar situations but not realize some ISPs will let you escape the CGNAT. They switched me after a five minute chat and within 2 hours (considerably less but they said up to 2 hours) I restarted my router and I was good to go.
Example: you want to access a home camera (or other IOT device) from your phone. Right now I don't see how to build this as FOSS without any third party or privacy concerns. With static IPv6 addresses it should be pretty easy.
I'm on Comcast in California, and I found that they're providing IPv6 (no CGNAT that I can see) through to my (personally-owned) router (an Asus RT-AC68U). So all my systems at home are getting an IPv6 (or multiple) using the /64 dynamically allocated by my ISP.
And today I just discovered that my parents, who get service from Cincinnati Bell FTTH, are also getting IPv6! They're using an ISP-provided router, and everything is just working.
I am really happy that things are rolling out, albeit slowly.
A dynamic /64 is still not proper Internet though.
As a residential customer, that would be a static /56 at least :
https://www.ripe.net/publications/docs/ripe-690
> /64 is not sustainable, it doesn't allow customer subnetting, and it doesn't follow IETF recommendations of “at least” multiple /64s per customer.
(Why are ISPs being skimpy on IPv6 addresses?? Doesn't this imply that they will need to do extra work in the future to move those /64 customers to /56 or /48 ?)
> An alternative is to reserve a /48 for residential customers, but actually assign them just the first /56. If subsequently required, they can then be upgraded to the required prefix size without the need to renumber, or the spare prefixes can be used for new customers if it is not possible to obtain a new allocation from your RIR (which should not happen according to current IPv6 policies).
IPv6 was finalized in 2017.
In 2020 Europe ran out of IPv4 addresses, and many Asian countries never had enough of them to start with (so quite a bit of people are effectively IPv6-only already).
An "I"SP that doesn't provide a /48 or /56 IPv6, shouldn't be legally allowed to advertise that they are providing "Internet" (and technically/historically, they're actually providing ARPANET, IPv4 having been supposed to be only a temporary, experimental version.).
And just like it was done for obsolete TV technologies, laws should be put in place first outlawing hardware that isn't compatible with IPv6, then later hardware compatible with IPv4.
From what I've seen, it looks like /64 are thought of as a vlan, within which clients can perform SLAAC.
For static IPs, I usually concatenate /56 + :id: + :suffix:.
Like: home computers on /56 + ::1 + SLAAC. Most OSes will dynamically change their IPs for privacy reasons.
My servers are on /56 + ::0 + :100,101,102, etc. I generally pick these suffixes to match with the IPv4 addresses, but you can allocate one per service, and get rid of reverse proxies (easier migration, you can just move the service to a new machine).
So, to take a specific example, 2a01:cb14:d6e:2000/56 is my ISP prefix, which can be thought of as the external IP, and 2a01:cb14:d6e:2000::11 is my server. 2a01:cb14:d6e:2001::/64 could be computers. I don't always follow the above scheme, IPv6 is big enough to get away with a lot of things, but it helps having something to default to.
My point is: you don't have to remember the prefix anyway, since every computer in the network will share it. Now, if you need static, easy to remember IPs instead of SLAAC, use static IPs or DHCPv6, or even better, mDNS to resolve .local addresses to IPs.
Looking at the above, this assumes a certain level of trust on the local network, which is fine at home or within a network dedicated to servers, but might not be at a company? mDNS can lie, someone else might advertise the same IP. These problems are not exclusive to IPv6, but they are a product of the era. Nowadays, I wish we just used crypto key routing (like yggdrasil does, and maybe TOR) on a planetwide mesh network, but we'll need IPv6 in the meantime :)
(Some people even advocate that consumer router IPv6 firewalls should be opt-in – which millions of them still are – and as you can guess with how opt-in works with consumers, the overwhelming majority of them therefore use IPv6 without a firewall.)
If I don't like the hostname (some IoT devices don't allow changing it) I can map a different name to that MAC address.
(I still use v4 and have no need to remember more than 2 IPs. Should make migration to v6 much easier.)
I can tell you the prefixes on my home ADSL connection, but not necessarily the ipv4 subnet, just because I work with the V6 addresses so much more often.
host -6 www.facebook.com www.facebook.com is an alias for star-mini.c10r.facebook.com. star-mini.c10r.facebook.com has IPv6 address 2a03:2880:f158:82:face:b00c:0:25de
I'll keep my IPv4 thank you very much.
while true; do ./runserver; sleep 10; done
in a shell script that would start a new server with one window per process. It got called from /etc/rc.local on boot IIRC.Deploys meant pulling the new code on the server (or maybe just editing it in vim right there), then just Ctrl-C in every terminal (or when I got lazier, `killall runserver`).
This was back in 2010 or so... I have professionally come a long way in the intervening decade, and no longer have any PROD services running in tmux panes, but I definitely learned to love that tool.
while true;
I have, while [[ ! -f "$DIRECTORY"/backup_exit.stop ]]; do
sleep 3600 # sleep an hour
time /home/ubuntu/.pyenv/versions/backup/bin/python main.py
done
Whenever I want the process to stop, `touch backup_exit.stop` . This waits until the run finishes and exists on the next loop.Anyway, that is saved to a 'backup.sh'
Then I have also,
#!/bin/bash
session="test"
backup="/path/to/backup.sh"
#create detached session named test
tmux new-session -d -s ${session}
# Create windows
tmux rename-window -t :0 'backup' #rename the first one
tmux new-window -n 'htop'
# Run processes
tmux send-keys -t 'backup' "$backup" ENTER
tmux send-keys -t 'htop' 'htop' ENTER
And this is run automatically on @reboot via cron.It also protects against flaky network connections and accidentally closing the local terminal emulator.
The tradeoff is between a network disconnect-reconnect killing your connection when it shouldn't (because you've noticed the disconnect but you wouldn't if you didn't send heartbeats), and discovering the network disconnect or dead peer (which you wouldn't if you didn't send anything, thinking it's still there).
Personally, I prefer to know the connection is dead sooner than later even if it comes back on its own.
Ping it.
Ping it
Ping it
All you got to do is ping it.
https://stackoverflow.com/questions/13628517/is-it-possible-...
Solution: I’ve found that simply executing “top” generates enough activity to keep the session alive until you get back to it.
As somebody who has experienced the described behaviour, I can empathise because it is super annoying. But this isn't a particular good point with which to illustrate it. If you're doing any sort of work on one or more remote servers that shouldn't be accidentally terminated like a large data transfer, you should be backgrounding it with screen or similar. Even for users without the idel connection dropping as described, the Internet can still fail.
EDIT: Yep, there's `draft-bider-ssh-quic-09` and a localhost proxycommand that offers both-sides OpenSSH integration:
I think SIP+RTP over QUIC might be a contender for that title. No more NAT traversal wackiness caused by needing both a TCP connection and a UDP connection. I still can't figure out why SIP didn't use TCP-over-UDP for the low-bandwidth control packets.
QUIC has protocol-level support for running both a reliable (TCP-like) and datagram (UDP-like) substream over the same QUIC connection. Cannot wait for SIP-over-QUIC!
https://cr.yp.to/djbdns/ipv6mess.html
Twenty years later, it still hasn't displaced IPv4. IPv6 was designed by a bunch of telephone-company guys who were used to Ma Bell being able to declare a flag day. That doesn't scale up to the Internet, and we're all suffering for their failure to understand backward compatibility.
The solution is for the ISP to fix their misconfigured NAT.
The solution is IPv6. Then you don't need your ISP to maintain a stateful connection table.
Nothing stops you from filling your car with orange juice either.
Filling your car with orange juice presumably stops it from working and is likely to cause damage, all while your parents question where things went wrong.
1. Just stop making new connections until some ports are freed. That'll make people happy...
2. Kill the connections that have least recently seen activity even if they have sent/received packets within the usual timeout.
3. Kill the longest running connections that aren't from a whitelist of target ports like 80 & 443 (P2P and VPN systems will just reconnect, the user will hopefully see no more than a short blip SSH will not fair so well).
But the real fix is to push your ISP to deploy IPv6. No need for the ISP to run carrier-grade NAT and you can host as many services at home as you want.
I even had it too when I used cell Internet, but cell ISPs have a better excuse.
We generally found someone else able to host, except for some 1vs1 situations.
After a lot of resources invested in debugging, switching out hardware, and interminable log files we suspected the ISP is to blame, so we built an MVP that was connected directly to what the ISP provides, and a client on a different ISP that wasn't closing connections. We saw the connection close at exactly the same time as before.
If this happened to me then I'd change ISP. NAT on your home Internet? What is this, the US?
I thought if you were in Europe you'd pretty much always be able to get public IPv4, IPv6, and IPv6-PD. I know I can.
As you might guess, the ISPs that are forced to use CGNAT due to the lack of IPv4 addresses, are also the ones who tend to be the first on IPv6, for obvious reasons.
I know of a big ISP, a relative latecomer, infamous due to its CGNAT issues, who last summer boasted reaching 99% IPv6 coverage.
(And at the same time, for some reason they had to be forced by the government threats, that otherwise they would not be allowed to use 5G, to add opt-in IPv6 on cellular, which they did at the last moment last month...)
If any of those points do not apply to your ISP, that's not because of CGNAT, it's just a shit company. I've used four different ISPs in the past four years and getting rid of CGNAT has not been a problem once. IIRC only one of them used public addresses by default, two offered free dynamic IPs upon request and my current ISP offers paid static IPs for $3.
Tons of my friends and family have had CGNAT and never known. It's just not a big deal for most content consumers.
setsid somecommand --blah &
I recommend it.
Every time I need to set it up, I go through this series of very short and digestible articles.
https://news.ycombinator.com/item?id=10937277 (discussion for the third in a 4-part series).
Never thought to look for an HN discussion on it, but it turns out that there is one.
IIRC, one thing the series doesn't mention is that a particular option needs to be explicitly specified in order to maintain port tunnels. By default, as long as the SSH connection succeeds during an AutoSSH reconnect, it will chug along happily without the port-forward if it's still blocked from the previous connection before the drop.
Keepalives are a basic requirement for persistent connections. It's literally what they're for. Harmless enough solution.
No clue why the LAN drops connections without keepalive traffic. I need to get up to speed on using WireShark to diagnose dropped connections one of these days, as I actually have a couple of dropped-connection issues that need troubleshooting.
Typically I'll get a connection timeout every few days, even with keepalives, and even with a hardwired DHCP address for the MAC in question. Connections with no traffic tend to get 'cleaned up' by somebody after a few hours at most.
after the TCP and SSH sessions have been established, no more packages are sent for a long time
."