Every day at the same time, my internet dies for 1 minute. How do I investigate?
How could I begin investigating this? I have a spare Raspberry Pi and and old Android phone at my disposal, and some programming competency.
How could I begin investigating this? I have a spare Raspberry Pi and and old Android phone at my disposal, and some programming competency.
My experience partially confirms that. I had an ISP who would redirect every syn packet on 80/443 to internal “low money” site in the night, but I had to reset my router after sending them money, cause otherwise that crippled session got stuck at their hardware until next autoreset.
1 minute may be a technical pppoe cleanup/reconnect timeout. Some shitty ISPs reset this timeout even for unsuccesful attempts, and this deadlock can last for half an hour until e.g. a router decides to back off for a while. Could be cured by turning it off for 3-5 minutes to cool off the ISP side. It’s rare these days.
i'd second this guess because its a common setup. OP could check the external ip before/after the reconnection and verify this way if it changed.
would also be helpful to know connection technology (dsl/cable/wifi/fibre..) and name of the isp
2. Do a ping 8.8.8.8 from your Linux machine (Raspberry Pi) or ping -t 8.8.8.8 from windows, and watch what happens at 3:40 PM
3. As others have said, turn off uPnP
- Some ISPs assign statically (odd, but true) meaning you will always get the same public IP
- The DHCP lease time is different for many ISPs: 5min to weeks, meaning you’d have to be offline for at least that long in order to be assigned a new IP when connecting
- Some ISPs lease it to the modem and not the router or anything internal (does not apply to what you said, but another tactic: changing the MAC address)
B) OP asked how to investigate. If getting a new IP dodges the issue and it doesn’t happen again, then there’s basically zero chance that the actual cause will be known
Although I suspect your ISP has a 24h lease and your modem renegociates at 3:40 everyday
But I’ve not experienced similar in 20 years of dynamic ips.
The way DHCP 'should' work is that when a client is part-way through the lease it should try to renew the IP it already has:
* https://en.wikipedia.org/wiki/Dynamic_Host_Configuration_Pro...
* https://tools.ietf.org/html/rfc2131#section-4.4.5
This allows the client to (hopefully) keep extending the IP it has so that a re-IP does not cause current connections to drop. If the modem is not doing this (can the OP log into it in someway to see logs?), then it is acting as a 'non-ideal' DHCP client.
Until this is sorted out, try rebooting the modem an an 'odd' hour that will not disrupt you during the day. This is no guarantee though: the DHCP server may remember the old/current DHCP lease and simply re-issue it with the same expiration time.
I'm on DSL/PPPoE, so this may not apply, but: my Asus router has a setting that allows it to automatically reboot. I do this at ~04h00 to get a new IP every day so help with privacy concerns. I generally surf with cookies disabled, so these two things help with the low-hanging fruit of simple tracking techniques. (My DSL modem is bridged.)
The lease would be dropped and renegotiated every few hours, we could notice it because our SSH tunnels would get torn down.
What was annoying was that the business plan had a static IP associated with it, presumably managed by their DHCP setup.
It took weeks to convince Comcast the issue wasn’t in our building, and moments for them to fix once they “got it”.
> The Gileadites seized the fords of the Jordan before the Ephraimites arrived. And when any Ephraimite who escaped said, “Let me cross over,” the men of Gilead would say to him, “Are you an Ephraimite?” If he said, “No,” then they would say to him, “Then say, ‘Shibboleth’!” And he would say, “Sibboleth,” for he could not pronounce it right. Then they would take him and kill him at the fords of the Jordan. There fell at that time forty-two thousand Ephraimites.
As OP said, mtr will tell you where in the line it’s happening.
mtr combines the functionality of the traceroute and ping programs in a single network diagnostic tool.
As mtr starts, it investigates the network connection between the host mtr runs on and HOSTNAME by sending packets with purposely low TTLs. It continues to send packets with low TTL, noting the response time of the intervening routers. This allows mtr to print the response percentage and response times of the internet route to HOSTNAME. A sudden increase in packet loss or response time is often an indication of a bad (or simply overloaded) link.
The results are usually reported as round-trip-response times in milliseconds and the percentage of packetloss.
This actually makes me wonder about whether there's a command that'll let me see short summaries of what every file under /usr/bin is, in the form of a list. Definitely wasn't aware of mtr as an average web developer up until now.My router is currently showing a lease time of 3 Min. I think most router tends to renew their IP at 50% of ISP's lease time ( in this case 6 min )
What are the purpose of these ridiculously low lease time? I remember in the old days they tend to go for 48 hours if not longer.
And apology for asking an off topic question. I guess that is what the downvotes are coming from.
But the other part is money... why give something away for free when you can make it a "feature" to pay for?
Tons of other reasons in various articles...
https://www.online-tech-tips.com/computer-tips/ott-explains-...
https://superuser.com/questions/590391/why-do-isps-change-yo...
Id be glad to give them away for free but it is all downside for us and over 99% of residential customers dont want one to start with.
I think we have 3 residential customers out of 1500 that have statics.
I'm actually glad I don't deal with that stuff on a day to day... but I like learning about it.
(or if I set up my own router could I just use the all?)
Sorry for the all caps. I just scrolled through my texts and saw the “internet out?” texts are at the same time.
Hopefully that means it’s a provider issue behind us. Wonder what we have in common?
- South Lake Tahoe, CA
- Spectrum Gig internet
- Asus Zenwifi mesh
Also Spectrum, but no gigabit here.
Because the exact same thing was happening to me at the exact same time. I’m no networking expert but when I checked the logs it seemed that every day Comcast would try to “fix” my (not broken) DNS settings. I had been using mullvad DNS and open DNS as backup but since switching to nextDNS the issue has stopped.
I don’t know enough about networking or DNS to offer any explanation as to why.
It went away after I called them and had them "de-provision" and "re-provision" my modem. Basically deleting it from their system and readding it. For whatever reason, whenever I'm having chronic comcast weirdness, having them do this solves it. It sucks that this is the way things are.
Anyway, if the ISP does this, then you have to reconnect once every day (or sometimes every two days), and if you don't do so manually they will just cut your connection on their side and make you reconnect.
My (German) ISP does that too, the one I had before did it, the one before did it as well.
You might also check the router's logs for that time, perhaps there's something useful information.
Also, it made it harder to run servers until we had services like DynDNS.
On newer connections it isn't common anymore, especially if you have a voice service.
I think the OP is so eager to get back online he sees 3:40 every day and is using that timepoint as his troubleshooting clue.
However If he were to disconnect 3:10 and then just be offline an hour until plugging the device back he would probably see the reset happen 4:10 going forward.
In my experience, even if I discover what is the problem, if it's on their side I still have to go through support to get it fixed, so I just wasted my time. And in most cases the lowest level of support doesn't even write down the information I give in the ticket, even if I send an email.
2. If there are packet losses, then you have your answer: the issue is from your ISP.
3. If there are no packet losses, then you'll need to look closer at your network. Check to see if some hardware on your network might be performing a reboot at the 3:40pm mark.
Switching to 5Ghz will help if this is the case
Chances are they experienced it there too. But I had an issue like that with Cablevision at 3-4am when they would push out firmware updates to the modem.
I had a similar experience to this years ago while working at my mother's kitchen table for while. The network would keep going dead and I had no idea what was causing it.
Eventually, I discovered that it happened every time she used the microwave to heat up a drink or something.
Changing to a different WiFi channel fixed it. :)
If that doesn't fix it, below is how I'd proceed.
Use the trace route networking tool (google it, different syntax on different OS) to some known good public IP address. I usually use 1.1.1.1 or 8.8.8.8.
This tool shows you the path your traffic takes:
$ traceroute 1.1.1.1
traceroute to 1.1.1.1 (1.1.1.1), 30 hops max, 60 byte packets
1 192.168.1.1 (192.168.1.1) 1.286 ms 1.386 ms 1.963 ms
2 192.168.11.1 (192.168.11.1) 2.591 ms 4.342 ms 5.219 ms
...snip...
16 0.ae19.GW8.CHI13.ALTER.NET (140.222.230.223) 42.888 ms 35.742 ms 0.ae20.GW8.CHI13.ALTER.NET (140.222.230.225) 41.060 ms
17 152.179.105.202 (152.179.105.202) 51.913 ms 50.681 ms 55.442 ms
18 one.one.one.one (1.1.1.1) 70.087 ms 65.130 ms 53.607 ms
The snippet above shows that I have two local devices (192.168.1.1, 192.168.11.1) my traffic passes through and 18 devices total before my traffic hits the public IP.Run this command at 3:39. Then run it again at 3:40. Then again at 3:41. Compare results and that should give you a good idea which device is malfunctioning at that time.
Some routers also have a ping command baked into them. If you can make it constantly ping 8.8.8.8, while you run the test above, it may help you. In the case where your traceroute stops at your own router, you can look at the ping from the router itself, and see if it also stopped:
A) If the traceroute stops at your router, but the router ping is not interrupted, then something in your router is flaking out.
B) If the router ping is also interrupted, then the problem is "upstream" from you. Maybe the cable modem (or whatever device) that is upstream from your router.
If you decide the problem is upstream, contact your ISP. If it's local, investigate and/or replace the flaky device.
HTH.
Still don't know what exactly caused this and I don't live there anymore but because I was the only one in my family with these problems, and only my room was facing the church, I want to believe the bells rang at a frequency that messed with my old routers Wi-Fi.
If someone could tell me if this Is possible, I could finally get some closure on this.
https://www.brgproducts.com/carillon_pricing/index.php
So it seems entirely possible the wireless part of the church bell system was interfering with your router. If the church did have real bells it seems very unlikely they would generate strong EM radiation at 2.4GHz or 5GHz. They might induce some vibration if they are really loud, which could trigger an intermittent connection in the router, though that seems pretty far fetched also.
> My devices remain connected to the network, but all traffic dies.
Can the devices send packets within the network, but not the internet? Use ping to double check. if no, do you have any wired devices to confirm rather or not its wireless interference vs the router doing some sort of maintenance task.
Does this impact all protocols? tcp/udp/imcp. I've seen random network hiccups only impact tcp before.
ping tests imcp, dns (nslookup on windows/host on linux) tests udp (and can also test tcp) http tests udp and tcp, depending on browser and service. curl can let you confirm the test is going over tcp.
MTR is likely your final solution for investigating this. It is like a traceroute that rapid fires out to get second by second details about all the hops in a network path. The types of errors you get can also tell you why.
If it ends up being tcp only and internet based, things get harder, some mtr clients support using tcp instead of imcp to find hops where tcp breaks, but your isp routers might also refused to reply to such packets directly.
Finally, watching the lights on the router and modem (if seperate) can be illuminating. First, get a feel for normal light operation, normal activity blink speed, etc, then starting a few minutes before the event normally happens, just observe the lights for changes until the event ends while also using your phone to know when the event has started and ended.
On that note, look up how to access web consoles on the router (and modem if seperate). They may have event logs that tell you if anything is happening, like ip renewals, or if they are getting commands from the isp to do things.
Second Check the logs of the pppoe daemon in your router. (If the router provides such logs).
After 20 years of network/system/software engineering across a wide variety of network sizes and levels of complexity, I teach people the following method to troubleshoot complex network issues:
1. Write down the symptoms that you're seeing. How do you know something is wrong? What do you see happening?
2. Draw a diagram that includes all of the components and nodes involved. (this is hard)
3. Develop some hypotheses that could be worth testing. (this is hard and highly variable based on knowledge/experience)
4. Establish a "test plan" that allows you to prove/disprove the hypotheses while making minimal changes to the system. Start from one source device and work your way out to the farthest component you identified in your diagram. Start at the lowest OSI layer and work your way up.
5. Methodically test components, step by step, along the diagram.
6. As you test, record your results. Add new hypothesis to the list as you gather more information but don't start testing them right away! You can develop a new test plan after you finish the first one that proves/disproves your new hypotheses.
7. Repeat all steps, adding new information, until you find a solution.
I know this seems like a lot of steps when you're just troubleshooting a WiFi issue at home--it may be overkill. However, this framework applies to that scenario or when you're diagnosing any complex network issue. You'll learn more about the systems, protocols, and devices that truly make up the Internet than you could ever imagine. Through repetition and experience, these steps will get easier and some of them can happen quickly and in your head.
To help you get started, here's some info for steps 1, 2, and 3:
1. From your post: "3:40pm every day my wifi loses internet access. My devices remain connected to the network, but all traffic dies. Almost exactly 1 minute later everything is resumed"
2. Start your diagram off as an "equipment string diagram". This will include all physical devices between you and the Internet. Along the way, you may need to modify the equipment string diagram to include "virtual" devices, network segments, various protocols/servers, etc. Your diagram should include at least the following items:
- Your laptop/desktop. If it's happening to all devices, then pick one. Identify your IP and MAC addresses
- The WiFi access point that device is connected to. Identify the IP and MAC addresses
- Any switches, firewalls, modems, etc that connect from your WiFi access point to your Internet Service Provider (ISP) and their IP and MAC addresses.
- The "next hop" from your modem into the ISP's network. You can use a cloud to represent the ISP's network that you don't know/understand, but always identify the IP address of the device that your modem first reaches. If you can find the MAC address (or other OSI Layer 2 address) as well, even better.
That's the equipment string between you and "The Internet". For now, you can ignore the complexity inside the ISP's network and beyond. You might have to add more of that later, but start small.We know there are other components involved in making the Internet work that introduce complexity and can cause issues along the way. Let's list them on your diagram and identify what servers are used and where they might be located.
- DHCP (local to your network and also between your modem and ISP)
- DNS (could be local to your network and often is a third party service either run by your ISP or not)
- Encryption (VPNs, SSL certificates, network device clock settings, etc)
- IP routing (what devices do IP routing? Hint: all devices that operate at Layer 3, using IP, do IP routing--including your workstations)
3. Some hypotheses (some were identified in the comments): - Is there a device in the equipment string that is rebooting every day?
- Is DNS intermittently failing?
- Is DHCP releasing/renewing your IP address assignment? This could be from your device -> local DHCP server OR your modem -> ISP DHCP server
- Is there an upstream connectivity issue with your ISP?
- Is your WiFi access point losing connectivity?If you consider 4 time zones in the us it's actually more common than birthdays since you'd divide by four again and people would suspect both a company wide tx synced issue and a per tx issue.
With 23 people about 1/3rd of the people commenting at this point we get a chance of 2 people colliding as:
Full day: 16% 12 Hour: 29% Across time zones: 79%
Hardwire the Pi to the network. Start a verbose traceroute and ping loop on a device on WiFi, as well as the Pi. Once it fails, observe the lights on your router/wifi device(s). Note time to the second of the start and finish of the outage, and the behavior of the lights. Check the results of your traceroutes/pings. Start digging through the logs of each network device.
Also, you can never quite rule out power, even if the lights stay on. A power conditioning UPS is useful here.
This would let me pick up on disconnects, DHCP lease issues, and a variety of other things in one hit. And if it wasn't any of those, would allow me to explore some other possibilities such as the problem being at the ISP's end.
I am with a broadband provider in the UK called Zen, at 23:15 every Sunday our PPPoE session is terminated, and it takes a couple of minutes before it successfully resumes. Occasionally it will happen mid-week, however still at 23:15. At one point I suspected the provided FRTIZBox router, however the problem persisted once we moved to an MT992 modem and a Unifi USG.
Honestly, this should not cause anything to lose its connection (the renewed IP address is the same), but it did, so I worked around it.
In doing some of my own network investigating recently, I discovered that services like Suricata will (by default) take down network interfaces and restart them when they get new rules. I wonder if something software on your router is doing something similar: Are you running Suricata or Snort?
Also: Logs are your friend. If you have access, review the logs of your firewall / router / etc.
Disclosure: I'm smarter now.
You can do that with cheap socket timer.
Not sure why would anyone set it to the middle of the day. Maybe error in setting times or clock or it was set intentionally so that you can service it at normal hours if it doesn't come back on.
The problem is you can put in 'off' time and 'on' time but they have to be at least 1 minute apart.
Edit: https://en.wikipedia.org/wiki/Dynamic_frequency_selection
On xDSL services, some monitored alarms may cause line problems when they "check in". In these cases an in-line DSL filter should do the trick.
Another possibility is that some modem/routers have a "self healing feature" that is just really a restart. See if there's anything like this in the admin/system part of the configuration settings.
Still haven't resolved it, but it is likely Fing on Big Sur.
HAVE a separate modem and router.
Or maybe it is related, in which case I'll eat at a hat.
I need to watch more Futurama. That was great!