Cellular outage in U.S. hits AT&T, T-Mobile and Verizon users
cnbc.com
cnbc.com
In the entire history of electromechanical switching in the Bell System, no central office was ever out of service for more than 30 minutes for any reason other than a natural disaster, or, on one single occasion, a major fire in NYC.
The AT&T Long Lines system in the 1960s and 1970s had ten regional centers, all independent and heavily interconnected. There was a control center, in Bedminster, NJ, but it just monitored and sent out routing updates every 15 minutes or so. All switches could revert to default routing if needed, which meant that some calls would not get through under heavy load. Most calls would still work.
Funny when they billed extra for long distance calls even though all calls were routed through one place for a huge geographic area. Calling your neighbour could be a hundreds of miles round trip over mobile.
Yes. Too much of routing is centralized. Since phone numbers are no longer locative (the area code and exchange number don't map to physical equipment) all calls require a lookup. It's not that big a table by modern standards. Tens of gigabytes. All switches should have a database slave of each telco's phone number routing list, to allow most local calls if external database connectivity is lost. It may be behind, and some roaming phones won't work. But most would get through.
You’ll see numerous COs in a big city, but they are also pretty widely dispersed throughout the suburbs and rural areas.
I mean COs are integral in AT&T's architecture, it's where every fiber connection lands on an OLT
In towns, we generally tried to keep loop lengths under 30k feet, but in rural areas that simply wasn’t possible. You’d often find remnants of party line systems in those areas and definitely load coils out the wazzoo. It was “fun” unwinding all that crap to install ISDN circuits and later DSL.
I remember the old hats at the time laughing about VDSL saying “leave it to the nerds to dream up some unrealistic shit where the loop length can be at most 2k feet, where does that exist!?” not realizing a few years later RTs and DSLAMs would mean a significant portion of city and suburban customers would be.
Then again. That only affected long-distance service.
"A History of Science and Engineering in the Bell System - Switching Technology 1925-1975" is a readable reference. The Internet Archive has it.[1]
More hardcore: "No. 5 Crossbar"[2]
The Connections Museum in Seattle still has a #5 Crossbar working.[3] Long distance used toll switches, "#4 Crossbar", and there were 202 of them.
#4 and #5 Crossbar machines are collections of stateless microservices, implemented from electromechanical components. The terminology is used in the old books is completely different, but that's what they are. Each service always has at least two servers. The parts that do have state are distributed. The crossbar switches that make actual connections have state, but are dumb - they are told what to do by "markers", which are stateless but can read the state of the crossbars and of other components. Failure of a single crossbar unit can take down less than a hundred lines at most. Other than the crossbars to external lines, everything had alternate routes. Everything has fault detection, with lights and alarm bells.
Error rates were fairly high. In the previous "step by step" system, a good central office misdirected about 1% of calls. With bad maintenance (and those things were high maintenance) that could get much worse. Crossbar was better, maybe 0.1% misdirected calls.
Routing tables in crossbar were mostly static ROMs of one kind or another. Routing consisted of trying a predetermined set of routes, in order. Clunky, but reliable.
Modern systems need a backdown to that mode.
[1] https://archive.org/details/historyofenginee0000unse_q0d8
[2] https://archive.org/details/bellsystem-no-5-crossbar-blr
[3] http://www.telcomhistory.org/connections-museum-seattle-exhi...
Anecdotally, I remember more electricity interruptions and plumbing issues when I was a kid, but that could be location dependent and I couldn’t quickly find good numbers going back that far.
Edit: While the phone network didn’t necessarily go down, I frequently got “all circuits busy” when I was a kid. I don’t remember the last time that happened.
Children have few responsibilities and are shielded by their caretakers. They simply do not notice much of the things that happen.
I think the average person views infrastructure improvements as improvements in the roads, airports, or air traffic control.
In USA for example, road design AND vehicle design is directly linked and beholden to NHSTA regulations and policies.
Infrastructure (IE State route roads) are ever trending towards wider lanes and more gentle shoulder. This is precisely due to vehicle industry requirements requiring more vehicle safety features (and thus width and length), height, and shoulder level clearance @ windows.
All these infrastructure and endpoint changes are driven "organically" by the USA trend towards SUVs, but mainly driven by insurance requirements. Insurance and gov't "make out" on safer/roads/vehicles due to (perceivabally) less accidents and road maintenance.
I can't speak to airplanes, but I imagine the fact that far, far, far more people are able to fly today than even 25 years ago should show that the infrastructure has drastically improved.
When it rained, I could pick up my phone and hear conversations from my neighbor on my landline and talk to them without calling.
Not to mention if you were in the same house, you could surreptitiously here conversations by just picking up the phone or getting a device from radio shack that didn’t have a microphone, that you could plug in to another phone outlet.
With analog cellular, you could also buy a receiver from Radio Shack and hack it to pick up the unencrypted signals from cell phones.
The roads in many US cities arent built to those standards and are grandfathered from them. New York City highways areas horrible.
Just purely in terms of total downtime itself we are less realible. Purely because of the complexity involves.
Didn't take out the entire Chicago area but was probably the worst case scenario for a suburban switch. Hinsdale handled the airports, FAA ATC offices, and the emerging mobile/cellular network. Long distance and 411 was down for whole counties.
This is a pretty good USENET archive/digest of the event, if you can get past the Web0.1 formatting:
http://telecom.csail.mit.edu/TELECOM_Digest_Online/1309.html
May be it doesn't even need to be 6G. With 5.5G and above and OpenRAN there is another opportunity to radically reduce complexity I could only hope there is enough of push towards this.
What is/were the cascading effects of this, particularly for drivers?
Many people in buildings were unaffected, as they could fallback to wifi. But I imagine this had a pretty broad impact to drivers.
Just a few things I can think of:
- Packages delayed (UPS, FedEx, Amazon, truck drivers, etc.) for drivers that relied on their phone's mapping apps to get them to their deliveries
- Uber/Lyft/taxi/etc. drivers not able to get directions to their pickups/dropoffs
- Traffic worsened because drivers weren't able to optimize their routes, or even get directions to their destination
Maybe larger companies have their own infra for this, or have redundancy in place (e.g. their own GPS devices)?
I'm curious to hear thoughts on whether these (and others) were impacted, or if there are ways they're able to get around this.
Also, unrelated to drivers, I can imagine there is/was a higher risk of not getting treated for emergencies due to not being able to make calls (I'm not sure whether/how emergency calling was impacted).
I have Google offline maps downloaded for areas I end up in just in this case. Gotta do traffic rerouting the old fashioned way though.
Or have an old-school GPS map thingy in your glovebox.
(Also have kiwix and a whole archive of Wikipedia on my phone).
I wonder if meshtastic communicators sales took off during this. How’s LoRa traffic these days?
Thankfully I’m in Canada where it’s not impossible to end up in the sticks with no service.
Chewing through your handful of gigabytes/month of data wasn’t hard. Only in the past year or so have double digit gigabyte/month data plans become cost-effective.
And our roaming prices are extortionate, so for jaunts over the border (or internationally), I’ll sometimes go “naked”.
Do they? I know there are a lot of old units out there but I figure people would have tossed them.
At least I’ve found Waze has been pretty good at starting off with wifi and loading the map of the whole journey after coverage was lost with some resilience for stops/detours.
You click a few buttons to download OSM tiles and then it does routing. The latest OSM even has a decent amount of stores, restaurants, etc., listed.
But even if it doesn’t, there are a ton of offline map apps that use OpenStreetMap data.
You can also install Organic Maps on your phone.
I make sure to have this around my usual area and anytime I travel to an area with poor coverage, plus my Garmin watch has offline maps and GPS everywhere, but this is not typical.
OSMand usage is even less common.
I think it's off by default, and I'm guessing most people haven't thought to turn it on, or are even aware of it.
Anecdotally, I’ve made it to a remote destination using Maps, then hopped back in the car an hour later (with no signal), and it couldn’t load anything. This seems to happen quite often.
https://en.wikipedia.org/wiki/Assisted_GNSS
Edit: You can think of it as a CDN for the GPS almanac.
My Garmin watch gets a GPS lock in way less than 5 minutes without any cellular connection.
I think that because Huami/Amazfit/Xiaomi smartwatches already do that. We know this from reverse engineering efforts in Gadgetbridge, but support for Garmin is still new and so there isn't as much info about it; either way it probably works in the same way.
If the receiver has recently (last few days) gotten a fix and hasn't moved too much from that fix, it'll be in at least warm start mode. It still needs to download ephemeris data, but this usually takes 30ish seconds to fix.
If the receiver has seen a fix very recently (last few hours) or a recent network connection, it can fix from hot start like you saw, which only takes a few seconds and may not even be observably slow depending on how the system is implemented. Phones go to great lengths to minimize the apparent latency.
https://www.bevhoward.com/TripMate.htm (Not me)
Back then, just getting a GPS fix at all was exciting. Then driving around with it propped on the dashboard or rear window.
It's not as simple as you think.
TOTP or stronger, please.
TOTP just "feels" more secure.
TOTP is more secure because it isn't tied to a phone number. You're right that it's still phishable but that's not the point.
In both cases, the primary benefit to the general population is to have a rotating credential that, if one website is hacked, is useless on another website.
You fully control how to store the TOTP seed and how you compute the value, so it is far more secure.
Yes, it can be phished if you fall for that, but it removes several attack vectors.
Sorta. The seed still needs to be issued to you in some way.
How was the first factor (the password) compromised?
Assuming the user is using site-unique passwords, in 99% of cases where an attacker obtains a functional password they can get at least one TOTP code or the seed in the same manner. (ie, if I can steal your password DB, odds are pretty good for me stealing your TOTP seed DB as well.)
The outcome of a single successful authentication is a longer-lived session cookie. Once an attacker has that they can reset your creds (usually just requiring re-entering the password) and the account is theirs.
IMO, the only 2nd factor that matters are those that mutually authenticate like PassKeys / FIDO keys.
One of the biggest weaknesses with TOTP apps I've tried using is that you have to remember to transfer them to a new phone before you get rid of your old phone. I once got locked out of a domain registrar because I set up TOTP on an old phone many years back. That was long gone by the time I wanted to do something with my domain.
TOTP is fine, but always give me recovery codes I can print and out and keep with my other important documents. Too many services don't do that.
TOTP -> Time-based One Time Password SMS -> Delivery mechanism.
You can deliver TOTP over SMS.
Obviously, SMS shouldn't be used, but I was under the impression that the code generation mechanism and the code generation algorithms are completely disparate concepts.
(Well I could be happier if they supports TOTP, but I'm not holding my breath)
Offline road maps are subject to construction/seasonal/holiday route closures/deviations too, and so is transit.
During Canada’s Rogers outage in 2022:
> In Toronto there was some dependency on Rogers. One quarter of all traffic signals relied on their cellular network for signal timing changes. The Rogers GSM network was also used to remotely monitor fire alarms and sprinklers in municipal buildings. Public parking payments and public bike services were also unavailable.
https://en.m.wikipedia.org/wiki/2022_Rogers_Communications_o...
As it was summer, I recall some park programming for kids had to be cancelled because the employees were required to have a phone capable of calling 9-1-1 (but sounds like that at least still worked here)
Their ops are critical enough you'd expect better from them.
Not the kind of shortcut Canadian banking takes for core stuff.
Unfortunately they failed to notice that this was a reseller for Rogers lines.
I'm not sure that is a thing. The vast majority of drivers are on familiar routes and are not navigating via electronic means.
Better question: How are the autonomous cars doing? Are they parked by the side of the road unable to navigate without cell coverage.
If I need to drive 20 minutes with most of it on the expressway, and they’re prone to accidents and there are multiple viable routes, I’m 100% going to load it up on Maps every trip, if it will save me being delayed 10-60 minutes every few weeks.
But if I’m going mostly backroads, probably not worth it, since you can more easily go around accidents, and they’re less common.
But again, I’m guessing more city expressway commuters use navigation daily than you think.
I've been rerouted due to an accident many times, and I've seen the detours get backed up because of people taking more optimal routes (without traffic being redirected via other means).
I'd be curious to see more data on it, but I would speculate it's less than the "vast majority".
> Better question: How are the autonomous cars doing? Are they parked by the side of the road unable to navigate without cell coverage.
Yeah, that falls under my point about Uber/Lyft/taxis. I would speculate there is broader impact from those vs. autonomous cars (that are probably still relatively uncommon).
I wish for an economic system in which all causes could be backpropagated to the source and the source be held responsible.
If for example I lost 2 hours of my time today because I had to fight with Comcast, Comcast should be charged for 2 hours worth of my hourly salary.
If I lost a job offer because of bad interview performance because of heating issues because of bad maintainence on part of landlord, landlord should be charged for the difference in time until I get my next job offer or the difference in salary until the next job offer.
If I had to fight health insurance for 5 hours on the phone due to incorrect bill and that caused me additional stress that caused my condition to worsen, health insurance should be held liable for the delta effects of that stress.
In this case the cellular operators in question would be held liable for the lost incomes of those drivers plus the lost incomes of passengers who lost money because they couldn't get to their destinations on time or missed flights and had to rebook them.
I know this level of backpropagation is hard to implement in the real world but it would be awesome if the entire world were one big PyTorch model and liabilities could be calculated by evaluating gradients.
Just having that map on the wall isn't going to do any good since without regular use, no one's going to be able to use it effectively. And it's doubtful that people can be forced into using it.
Modern trucks have cell modems tied to a private APN that are used for updating vehicle firmware & doing telematics. They also typically have a route to the internet that provides a WiFi hotspot in the cab.
Depending where the fault was in the telco stack, that APN may have still been functional
Not saying this was a significant resolution, but at least a possibility.
This might explain a huge random traffic jam I hit in the middle of my town this morning.
I had no idea any kind of an outage was happening because I've intentionally scaled back my dependence on my phone. I always used to automatically pull up Google Maps to navigate no matter how short the trip. At some point I realized I was losing my ability to travel without being completely dependent on some company tracking my location and telling me what to do, so as part of my phone de-Googlification I switched to Organic Maps. And even then I try to navigate on my own without any GPS assistance as often as possible. I feel like navigating is a skill you can actually lose if you don't practice doing it.
After running an errand across town this morning, I decided to try getting back home via the biggest arterial through the city that I know about, and I immediately hit a huge westbound backup stretching at least a mile. It was a total standstill. I peeked ahead trying to see if there was some kind of accident or something and didn't see anything. Everyone was just sitting in this traffic jam, and I couldn't for the life of me figure out why.
I immediately flipped a u-turn and went 3/4 of a mile north to another westbound road I knew about. That one was completely clear of any traffic at all, and I was able to drive the speed limit all the way back.
The most-used navigation apps I know of suggest alternate routes when there's congestion, so why were all those people just sitting there in that jam while a parallel road less than a mile away was clear? Maybe it was this cascading effect of too many people conditioned into being told what to do by their phones while their phones couldn't tell them to take the other route.
radio is pretty good with traffic news, but how many people would even think of local radio?
Yeah, I get that there can also be a bit of a sunk cost thing along with regret minimization going on too. I think game theory suggests that you should switch routes the instant you hit significant congestion though, because P(congestion on the current route)=1 as soon as you hit it.
There is a lot of people who couldn't navigate to a neighboring street without a direct directions even if their life depended on it.
Add to that what the most people doesn't have a slightest idea where are they, where are the cardinal directions and what they need to get from point A to point B.
And I'm only going off of examples I have heard. These outages are very damaging.
I think Sears around '04-06 (?) was the last time I saw one of those used. I think I bought a dehumidifier or air purifier.
When they started rolling out credit and debit cards without the raised numbers I thought fondly of those and how they were definitely done for now.
- A new route injected in the network caused the routing engines on a type of cellular specific equipment to crash nationwide. This took down internet access only from cell devices nationwide. But most people didn't notice because it happened at 2AM maintenance window and was fortunately discovered and reversed before business hours why the routing engine was in a crash loop.
- A tech plugged in some new routers, and the existing core routers crashed and rebooted. While the news worthy impact was just a regional outage for something like 20 minutes, we discovered bugs and side effects from the Pacific to Atlantic coasts over the next 12 hours. So when you say you're impacted at location x, that data point could be everyone is down in the area, many people are having issues, or only one or two people have issues spilled over to some other region. This is why seeing it does or doesn't work in location x is limited value, as almost every outage I've investigated could result in some people still having service for various reasons. The question is in a particular area is it 100% impact, 50% impact, or 0.001% impact.
- A messaging relay ran into it's configured rate limit. Retries in the protocol increased the messaging rate, so we effectively had a congestion collapse in that particular protocol. Because this was a congestion issue on passing state around, there were nationwide impacts, but you still had x% chance of completing your message flows and getting service.
And then there was the famous Rogers outage where I don't remember them admitting to the full root cause. It's speculated that they did an upgrade on change on their routing network, which also had the side effect of the problem booting all the technicians from the network. Then recovery was difficult because the issue took out the nationwide network and broke the ability for employees to coordinate (you know because they use the same network as all the customers who also can't get service). All the CRTC filings I reviewed had all the useful information redacted though, so there isn't much we can learn from it.
So it's fun to speculate, but here's hoping at the end of the day ATT is more transparent then we are in Canada, so the rest of the industry can learn from the experience.
Of course, was fun to see yet another huge org have no back-out/failure plan for their potential enterprise-breaking changes. No/limited IT 101 stuff here.
The only positive thing we learned was that the big 3 (really 2) telcos thought it would be a good idea to give eachother emergency backup sims for the other network to key employees in case their network went down. They did that in 2015, but better late than never.
Fun that Rogers used the same core for wireless and wired connections, so many of us were in total blackout, even if we used a 3rd party internet provider that ran over Rogers. Like, everything including their website was down, corp circuits, everything with non-existent comms from Rogers.
Thankfully my org was multi-homed and switched over its circuits at 6am so on-site mostly continued without issue.
Also fun where the towers remained just powered on enough for phones to stick to them but not be able to do anything, so 9-1-1 calls would just fail, instead of failing-over to other networks. Seems like a deficiency in the GSM spec (or Rogers SIM programming?) that I don’t think was actioned on.
https://en.m.wikipedia.org/wiki/2022_Rogers_Communications_o...
If it ran over Rogers circuits then why wouldn't it go down too? Isn't that the case everywhere?
The 3rd party providers aren’t white-label resellers, but there’s obviously some overlapping susceptibilities to going down when Rogers breaks something. Depends what they break, and in this case, it took them down too.
Actually, I think this is going to change after the Rogers outage, it's just slowly happening behind the scenes so it's not getting much attention these days. The government has mandated a lot of industry response to failover between providers... we'll see where they land after all the lobbying happens. I do think implementations are changing a bit around this, mostly in the phones so that they give up and go into a network scan if the emergency call is failing.
I worked mostly on core network stuff, so I was a layer removed from the towers, but if they hadn't lost management access they would've been able to tell the tower to stop advertising the network and 911 service. I do understand the question of from a vendor implementation perspective of how automatic this should be though... because automation in this regard does have some of it's own risks and could complicate some types of outages or inadvertently trigger and confuse recovery of problems.
I'm with you though there should be an automatic mechanism to fail over to other network operators, I just haven't thought through all the risks with it and I hope the industry is taking their time to think through the implications.
It seems like this is a global problem, since all Rogers-subscribed devices in a Rogers reception area couldn’t make 9-1-1 calls. But could be a SIM coding issue and not afflict other providers elsewhere.
I just always imagined the GSM spec was so resilient that you could always make a 9-1-1 call if a working network was available but this outage proved that wrong. Surprising to learn in 2022.
Of course it’s Canada, so I agree with them that the thought of letting users failover to a partner for everything would thrash the partner’s networks. Even though Canadian subscriber plans are laughably low in monthly data and population density is low (per the telecom’s usual excuse for our high prices) it turns out the telecoms still underbuilt their networks to have less capacity than what other networks internationally built out to support plans available on the international market (e.g. close to truly unlimited data/free long distance calls)
The X is broken but claims it isn't stops failover pattern is strong all over networking. It's not unusual to see it in telco root cause analysis.
As I recall it is slightly more nuanced than this and was particular to the failure mode, and has a couple of different things aligning to create the failure mode.
If you're phone is just blank, no sim card. To make an emergency call, it has to just start scanning all the supported frequencies. This is very slow, tune radio, wait for the scheduled information block that described the network on the radio protocol. See if it has the emergency services bit enabled. If not, tune to next frequency and try again. I used to remember all the timers, but almost a decade later I can't remember all the network timers for the information blocks.
The sim card interaction, is say you're at home and you boot up your phone with 100% clean state. You don't want to wait for this scan to complete, so the SIM card gives the phone hints about which frequencies the carrier uses, so start on frequency x to find the network. But if you roam internationally, it can take alot longer to find a partner network, and there are some other techs around steering to preferred partners, but I don't know that those come into play here. I don't know but would be surprised if there is a SIM option to try and pin the emergency calls to a network, I think it's more likely the interaction is this hint on where to start the scan.
The way the rogers network failed, it appears to me it caused the towers to stay in a state where they advertised in their radio block the network was there, and the 911 bit was enabled so the network could be used for emergency calls. This is where I don't really have the details since they haven't been public about it, how much of their network was still available internally. Maybe the cell towers could all see each other, that network layer was OK, and the signalling equipment was all talking to each other as well. That's the part I don't really know and have to speculate, as well as the tower side since I was a core person. So because the towers had enough service to never wilt themselves, they kept advertising the network, along with the 911 support. But then when you try to activate an emergency call, somewhere in the signalling path, as you get from tower to signalling system, to the voip equipment, to the circuits to the emergency center the outage knocked something out. Oh and for all these pieces of 911 equipment, there are two of everything for redundancy... two network paths, two pieces of equipment, etc.
And because they lost admin access to their management network, no one could go in manually and tell the towers to wilt themselves either.
If the towers had just stopped advertising 911 services, the phone would fall back into the network search mode as I described when you have no sim card. It just starts scanning the frequencies until it see's an information block for a network it can talk with the emergency support advertised to and does an emergency attach to the network that the carriers will all accept (An unauthenticated attach for the sole purpose of contacting an emergency center).
So my suspicion is because carriers are so used to we have two of everything, and all emergency calls are marked for priority handling at all layers of the equipment (they get high priority bits on all the network packets and priority CPU scheduling in all the equipment), this particular failure mode where there was a fault somewhere down the line, and they lost control of the towers to tell them to stop advertising 911 services all sort of played together to create the failure mode.
0) At the network terminal level (mobile phone): at least for emergency calls if a given network fails to connect, fail over and try other networks. Even if the preferred networks claim to provide service.
1) At the network level: failure thresholds should be present. If those thresholds are crossed enter a fail-safe state. This should include entering a soft offline / overloaded response state.
2) Where possible critical data paths should cross-route. Infra Command and Control and Emergency calls in this case. Though if Roger's issue was expired certs or something the plans for handling that get complicated.
Days later, Rogers said you might be able to pull out/disable your SIM card to call 9-1-1, but then it depends: if Rogers is the strongest network, you might end up in the same predicament anyway.
And yet, like everyone else, I genuinely feel that I'm probably right
The Rogers outage in Canada took out the nationwide debit card payment network because that infra depended on Rogers. Credit cards still worked, but depends on your station’s access to make the transaction. And no shortage of shops running their POS “in the cloud” and needing to close if they lose internet access. I actually did have to lend cash to a colleague to buy gas to get home during that Rogers outage.
All it takes is for one pipeline valve to depend on a cellular connection for billing to get the whole line shutdown.
And ugh, we hope for a botched software upgrade too, but a corp cyberattack is so much harder to recover from so can’t be discounted from the realm of possibilities. I know that’s where my mind went with Rogers given how thorough their outage was.
Was kinda unimaginable for a total outage to happen with no org comms ready to go in the pipeline. Your plans are supposed to have those comms ready for a bad update that you’ve been planning for weeks. It’s a cyberattack where you may stay silent. But I know Rogers isn’t going to admit fault until they find someone else to blame.
yeah, a lot of orgs just don't enable that (or don't have a process to enable it as required, and have difficulty pushing out a notice to do so if the network is down!).
Also can only do offline credit card transactions. Can't with our Interac (Canadian-only) debit network. Unsure about Visa/Mastercard debit transactions.
AIUI, the debit card itself enforces online confirmation, even if the transaction goes through the credit card rail.
Projecting, biased.
You are welcome to infer as to why I’m thinking this way!
This is the thing with black swan events. The more pedestrian explanations are almost always true, but then there's a tiny fraction of the time where you're much, much better off having taken a bit of an alarmist view.
We are wired that way for a reason. Until you personally see conflicting evidence you have to make an assumption or you would spend your life paralyzed or ignorant.
Biology rewards action more than accuracy.
>A temporary network disruption that affected AT&T customers in the U.S. Thursday was caused by a software update, the company said.
>AT&T told ABC News in a statement ABC News that the outage was not a cyberattack but caused by "the application and execution of an incorrect process used as we were expanding our network."
https://abcnews.go.com/US/att-outage-impacting-us-customers-...
I assume roaming being one of the top reasons no?
Is this common in the industry?
You don't need to register to allow 911 calls. You 'register' (it's not a regular registration) at the moment you are placing the actual 911/112 call. At least that was in 2G/3G networks, doubt it changed.
There is always some amount of terminals without SIM or without working SIM, there is no need for them to bang every available network in the vicinity just in case there could be an emergency call.
https://www.wired.com/story/att-landline-california-complain...
Status: Restored AT&T FINAL, Service Degradation, Global Smart Messaging Suite AT&T Global Smart Messaging Suite
Event description: FINAL, Service Degradation Impacted Services: MMS MT Start time: 02-21-2024 22:00 Eastern, 21:00 Central, 19:00 Pacific End time: 02-22-2024 11:00 Eastern, 10:00 Central, 08:00 Pacific
Downtime: 780 minutes
Dear Customer, We are writing to inform you that Global Smart Messaging Suite is now available. The MMS MT service has been restored and our team is currently monitoring Thank you.
AT&T Business Solutions Kind Regards,
The AT&T SMS Service Administrator
Maybe im reading to much into it but it bothers me that thats not in the communication.
They need to be accurate. At&t status claims everything is fine.
My wireless service is down. Down detector has tens of thousands of reports, so clearly everything is not fine.
Either they automatically update based on automatic tests (like some of the Internet backbone health tests) or they’re manually updated.
If they’re automatic, they’re almost always internal and not public. If they’re manual, they’re almost always delayed and not updated until after the outage is posted to HN anyway.
It’s most annoying when you have something like recently - known maintenance work on my upstream home fiber connection that was resulting in service degradation (but not complete loss, my fiber line was back to DSL or dialup). The chat lady could see that my area was affected, but the issue lookup system couldn’t.
If the issue lookup had told me there as an issue I’d’ve gone on my merry way.
I even checked a few more times until it was resolved; the issue never appeared in the issue lookup system.
Making this decision easy is a fight I fight for my customers every day. :)
Quite often you see automated tests that check how well your cache/in memory data are working. But when some other customer that isn't in the hot path tries to access their request times out. I've seen a lot of people making automated checking systems fail at things like this.
>Service Alert: Some of our customers are experiencing wireless service interruptions this morning. We are working urgently to restore service to them. We will provide updates as they are available.
It would be nice if the FTC mandated this. It is exhausting when the status page is taken over by the marketing department (the infamous green check with the little "i").
I'm not familiar, what are some examples?
https://mailman.nanog.org/pipermail/nanog/2024-February/2250...
I could log into my AT&T account just fine and all phones showed up correctly.
(I’m submitting this from an AT&T 5G connection, no WiFi nearby)
The cause of a failure of the HSS could be manifold, ranging from router failures to software bugs to cyber attack (databases of 100M+ users being a juicy target).
One slightly scary observation from NANOG was that FirstNet, the network that ATT built for first responders, was down. That would be ugly if true and I'd expect the FCC to be very interested in getting to the bottom of it.
[1] https://www.cbsnews.com/news/outage-map-att-where-cell-phone...
It's interesting how naked I feel without access to the internet. I reach for it way more often than I would have ever guessed, something you only notice when it's not there. Last March my area saw large wind storms that knocked out power for almost a week (I'm not in a rural area). I can work around the loss of power but the cell tower(s) that service my area could not handle the load and/or the signal in my house was weak and I was unable to load anything. Not having internet was way worse than not having power and I ended up driving a few hours away to my parent's house instead of staying home.
Now my computer is insanely more powerful but without an Internet connection it feels dead and useless.
This happened to us with the recent storms a month or two back, some places didn't have power restored for 2 weeks+
[1] https://www.cisa.gov/news-events/cybersecurity-advisories/aa...
[2] https://twitter.com/CISACyber/status/1758495005176447361
> All clear! No outages to report.
> We didn’t find any outages in your area. Still having issues?
https://www.att.com/outages/If you tried to make a 9-1-1 call, it would just fail. It wouldn’t fail over to another network because the towers were still powered up but unable to do anything, and Rogers couldn’t power them down because their internal stuff was all down.
Like a day later they said you could remove your SIM card to do a 9-1-1 call. Thanks guys.
Of course, no real info from the provider during the outage. Turns out they did an enterprise-risking upgrade on a Friday morning and nobody at the org seemed to have a “what if this fails plan”. CTO was on vacation and roaming phones were black too and he thought it was just an issue for him.
https://en.m.wikipedia.org/wiki/2022_Rogers_Communications_o...
It could be entirely coincidental and unrelated to the stuff with other networks, but the timing was odd and I have never ever seen anything like this outage from them. I can think of one time it was out for around 2 hours in the last 5 years, and it was with a very specific infrastructural upgrade they knew about.
Still down for me though.
[1] - [https://www.cnn.com/2024/02/22/tech/att-cell-service-outage/...
This isn't telling of anything, right? Wouldn't CISA be involved with anything that impacts Public Infrastructure at this level?
> “Everybody’s incentives are aligned,” the former official said. “The FCC is going to want to know what caused it so that lessons can be learned. And if they find malfeasance or bad actions or, just poor quality of oversight of the network, they have the latitude to act.”
If AT&T gets to decide if they are at fault, they will, of course, never be at fault. So a third-party investigation makes a lot of sense.
I would also suspect that the FCC would not be as well versed in determining if there was a hack or even who did it, which is why I feel like CISA would need to get involved in the investigation.
like, you could commit a dumb BGP config and break lots of stuff. have done that in the past, actually...
but any time a national-tier ISP has a national-level outage, that warrants a look from multiple orgs. and given the number of threat actors like china, NK, iran, and russia, who are, and have, made aggressive efforts in this space -- and have strong reasons to do so now -- its not crazy for the US fed'gov to want to know a little more, and offer to help. but again, entirely possible it's unrelated.
I wonder if the MVNOs that piggyback on AT&T are showing down also. If not, it’s some AT&T service authorization system that exploded.
Your iPhone will instruct you on where to point and help you track an emergency satellite that is manned by live humans who will take your emergency request and relay it to the proper people.
More specific info here: https://support.apple.com/en-us/104992
If there is no cell service then it's SOS with a little picture of a satellite next to it.
So the radio bands may play into it although I would think with latest iPhones, they can use any of the bands from AT&T although I could be wrong.
I suppose I can't speak to likelihood with a sample size of one.
My T-Mobile phone hasn't had any problems, knock on wood.
My Pixel 4XL was working at 2am as I placed it on the charger (even though it had about 80%). Noticed zero bars. Shrugged and went to sleep.
Woke up at 9am and the phone is totally non operational.
Yes I know sometimes phones just break, but less rarely while not moving, not wet, not dropped, or due to a charging issue.
What a coincidence.
It's basically almost as good as the pixel 4xl although picture quality isn't as good.
Regarding the cellular based authentication, it is perfectly doable to be securely authenticated even without the any connectivity and there is a solution to this based on combination of MFA and OTP [2].
[1] Martin Kleppmann talk on local-first (LoFi):
https://news.ycombinator.com/item?id=39444519
[2] A lightweight and secure online/offline cross-domain authentication scheme for VANET systems in Industrial IoT:
>postmortem
Let's hope that everyone lived so a postmortem won't be necessary :)
The AT&T land lines (CAMA trunks provided by the ILEC Frontier) that handle the 911 service did not fail. Only mobile service failed.
How do you actually report, for example, that .003% of your customers are having a really bad day but the rest are just fine?
secondly, there are financial, management and information security pressures to NOT REPORT reality to the public. This happens VERY OFTEN in real business. In fact, that is why legal enforcement actions and real consequences are crucial versus Big Business.
Small outages happen all the time, and are difficult to report accurately.
I think AWS has pivoted to trying to report status in each individual customer's support portal, so that they can give a dashboard that reports the state of the cloud from that customer's perspective. That way a rack down that only affects a few customers is only reported to those customers, and the dashboard doesn't have to always be red for everyone (or green for everyone, even those affected).
Because those cloud provider internal tools are often wrong about who's impacted and who's not.
Before I could complete my report, the call dropped, and for the next 10 or so minutes I had no service, only "SOS".
I'm on Verizon, and the timing doesn't match up with this headline, but now I'm suspicious.
Note: This site does not like to entertain contentious topics and has rate limited me over my prior post about factory and foundry installed malware in Chinese manufactured equipment because I did not respond to a suspicious demand for proof, instead opting for a sarcastic reply. I can live without being associated with Garry Tan's company, YCombinator. (I thought to post this link while posting to another site, not because I hangout here.)
>AT&T told ABC News in a statement ABC News that the outage was not a cyberattack but caused by "the application and execution of an incorrect process used as we were expanding our network."
https://abcnews.go.com/US/att-outage-impacting-us-customers-...
Do I believe that? No clue. I believe it more than people speculating the timing corresponds with the solar flare or nation state taking it out.
Resets futile.
Shows "Mobile Broadband Disconnected"
However can still ping google, with excellent under 10ms response.
You would think that a nation wide cell service would be more distributed and that it wouldn't all go down at once like this.
But maybe there are some good enough reasons that there needs to be some central components or systems that make it all work.
In those cases, there were mitigations, but the failures cascaded.
What's also odd is I'm not even able to log into my ATT.com account.
When I try to log in, it states "want to pay your bill".
So my immediate concern was, did my credit card expire and my service was turned off.
I'm still hoping that's not the case, and it's just that I'm impacted by this outage (and not something else).
So not a network outage, but account related?
I think your experience is just that it's not 100% outages, some people still have service.
Even if its not a hack I'd love to see the root cause on this one! Communications is critical infrastructure, so I'm gonna guess the government will demand a full report.
Back in my circuit switched days, we lost half of our US long distance routes, because a farmer in Wyoming dig up an unmarked fiber link. Hence, backhoe fade.
Sounds like only one had an outage
Me with a Pixel 7 Pro (dev preview android) on the same plan: not affected.
Strange.
Strangely sending to that phone from an iPhone looks like it sent, but nothing was received
Calling from Android to iPhone gives 'your phone is not registered', iPhone to Android plays a 'call could not be completed'
iMessage between iPhones seems okay
Thankfully, WiFi calling seems to be functioning.
I can't wait to see a postmortem on this outage.
[1] - https://www.thousandeyes.com/outages/
Edit: I’m getting 1 bar when I usually get 4 in southern Ontario. But I see no broad reports of issues.
Could be a hack Could be a single point of failure Could be a config change that borked the system
Could be other things that make more sense we will have to wait for more info.
We did have Two Class X flares from sunspot 3590 but they did not result in a coronal mass ejection.
There was a CME from another sunspot not visible but it is not aimed at the earth.
https://www.swpc.noaa.gov/news/two-major-solar-flares-effect...
Edit: And back again a few minutes later.
How can one carrier going down affect your ability to make emergency calls?
I like that they used this 1850's era tech during the outage: https://www.wbur.org/news/2018/12/30/911-outage-fire-boxes-b...
Telecoms, airlines, insurance scam..ahem companies...
Must be nice. I wonder if I could just stop showing up to my job and keep collecting my check. No? Why not? We allow the entities that write our laws and have politicians rubber stamp them to do it. Seems a mite unfair.
Maybe we should let the federal government run it, like it does the Social Security Administration - which btw, shuts down its website for 4 hours every weekday night and even longer on the weekends.
** They shut down a freaking website for almost 40 hours a week ** because they don't have the technology or skills to keep it running 24/7...in 2024.
No thanks, the reason that this is even in the news is because widespread outages are so rare.
Is this only effecting non priority accounts?
Seems like they are doing drone testing. Hmm, should we use drone network that will interfere with mobile phone network?
I use a carrier that leases from a major (Verizon). There were clearly disconnects between the systems that even tier-3 support expected to automatically resolve overnight (i.e., be eventually consistent). Still, after 2 such overnight waiting periods failed, I got someone to spec a fix and see it through over the course of 3 hours. They were clearly surprised that both their system and the major system were not reporting correctly, and the solution came only after making changes while ignoring status.
I got a sickening feeling that enterprise software layers mated with eventually-consistent web persistence across two outsourcing organizations to produce a situation where truly no one understood what was happening -- or could even figure it out. This resolve only when some youngish person had the guts to just make changes that should work -- so we're back to the days of heroes.
Many systems issues are mistakenly thought by non-technical users to be "hacking".
Just imagining the dropped emergency calls today, etc?
(Genuine question)
Knowing most megacorps, they'll blame some midlevel engineer doing a "bad" thing, when it was ordered from above.
(charitably) - FirstNet is mostly now just a type of billing plan and priority level, on AT&T's network.
CNN reported that AT&T confirmed it has not been impacted throughout this event.
https://www.cnn.com/business/live-news/att-outage-02-22-24/i...
https://www.ooma.com/blog/home-phone/landline-home-phone-is-...
Carriers do get fined for E911 downtime IIRC. It is taken very seriously, as a result.
People/companies that actually care already have a second line and are taking very good note of this.
But in this case, it's redundancy from the various providers - all providers must let you make a 911 call, no matter your phone or contract or provider.
So if AT&T's towers go down, your phone can still make 911 calls via Verizon or whomever else has a tower in your area, anything that can hear your phone will respond to a 911 call request.
No no no, see, we have those. But we only use them to make sure that our for-profit prison complex stays massively profitable and the people in power retain that power.
... isn't that pretty much everybody? Not sure if I know of any other actual plant owners