How Southwest Airlines melted down
wsj.com
wsj.com
Anecdotally, I flew Southwest just before Christmas. The network was already buckling and we had major delays, but we were lucky and made it through. Despite the stress, the SWA crews were helpful, empathetic, and polite. They handled it better than I would have if I had been in their shoes.
"In one example during the storm, the system assigned a pilot to deadhead on a flight from Baltimore to Manchester, N.H., and then back to Baltimore the next day, without ever flying a plane"*
* The article defines deadheading as sending a pilot as a passenger to get to another location.
It would be interesting to look at what the system is trying to optimize for to make such choices.
Or (more likely, I think) the first deadhead was planned to have the pilot take over a flight which was then cancelled after the pilot was already underway.
Very few software shops I am aware of that do anything like this.
It's quite feasible to get within a factor of 2 of the optimal solution (with much less processing power), which sounds great from a CS algorithm analysis standpoint, but a factor of 2 looks like an awful schedule from a human standpoint.
Having a (relative) bunch of surplus critical-task workers, forever being shuffled to where they system guesses they're most likely suddenly be needed, makes perfect sense.
(Yes, you'd have to be a bit more sophisticated, so your surplus pilots got enough flying time to stay certified. And didn't get pissed and quit. Assume that I know why RAID 5 is better than RAID 4:)
That could also be the result of thrashing after they were really far into the mess. That is, the solver maybe did output something sensible, but that solution has to get into the system that's used at runtime. Someone may have run another solution, or did manual updates while it was too overwhelmed to get all the changes posted. The solvers also depend on other context that's changing underneath them.
No idea what SkySolver actually does in totality, I'm sure it's complicated but I would think a flight crew could indicate where they are right now and then it could maybe pickup the next possible course they could perform. Not sure why the phone lines "jam up" exactly, don't you have a hierarchical management structure for this sort of thing? Or do 1000 pilots all report to one person?
They've got like 1000 planes and like 100-150 destinations, it's not the traveling sales man problem, an optimal plan isn't needed now so much as a functional one.
Of course, it's easy to bitch about it not being hard when I've never seen the code. Maybe is also tracks hours and does payroll and a dozen other functions.
If the duty time rules only count time after takeoff, or only time after doors are closed, those both offer a different set of times.
The resumption itself.
> Not sure why the phone lines "jam up" exactly, don't you have a hierarchical management structure for this sort of thing? Or do 1000 pilots all report to one person?
It's not a problem of management (well it is a problem of 15 years of management fuckup which is the root cause), it's a problem of the scheduling system needing manual updates when things didn't go to plan.
At this point it's completely screwed, so it needs to be completely reset, as in the entire scheduling system needs to be reconfigured from empty, more or less.
And because SouthWest operates entirely on point to point, the same cascading properties which led to its complete collapse mean it needs to restart in a somewhat synchronised manner, otherwise you fly 3 planes, there's no followup, and you're hosed again.
> They've got like 1000 planes and like 100-150 destinations, it's not the traveling sales man problem, an optimal plan isn't needed now so much as a functional one.
The system completely lost track of crews, so all of them need to be relocated, their work cycles reworked, flight slots need to be reallocated, flights need to be re-encoded.
Then next earnings release - when they post shareholders profits - that’s exactly how much they were willing to hold back to fix this, trading profit for people stuck for days.
Accrue a few billion for refunds, depreciate your software to zero, and start to set aside a few billion for new cloud whatever.
Basically, take this as a huge loss and communicate that you’ve lost a lot of goodwill and trying to make passengers whole is #1 priority (after safety, of course).
https://www.forbes.com/sites/advisor/2020/07/15/how-airlines...
And loyalty programs are bigger business for the higher end airlines that cater to business travels and the more affluent.
On top of that, the top 3 airlines make a lot of money off a first class and business travelers who are using other people’s money.
If they have operating income of eight to ten billion a year, spending a third of a billion to upgrade systems seems like it would have been a reasonable investment instead of buying back shares and increasing the dividend.
- use a tree traversal algorithm with a lot of built-in optimizations and pruning. Google OR-Tools is one well-known solver.
- use a set of meta-heuristics (Tabu search, simulated annealing, etc.) against arbitrary predicate expressions. The best example of this kind of silver is OptaPlanner.
The advantage of tree traversal is that it is exhaustive and guaranteed to be optimal; but it requires a fixed computing budget for a given set of constraints. Given a large compute cluster it's ideal, but on individual machines/servers these solutions tend to take hours or days. Southwest's system would likely require a significant supercomputer to run its scheduling through an exhaustive optimizing constraint solver, since it would have to re-run the entire solution (or significant numbers of subtrees) when any variable changes (which happens likely several times per second).
Meta-heuristics are more flexible and allow for all sorts of interesting, convenient features that can be very helpful when your compute budget is less than a Tiobe-500 supercomputer. They offer time-constrained solving, real-time monitoring of a solution in progress (you can watch it get better as more of the solution space is searched), and over-constrained solving (where the requested solution is impossible, but we need to make a best-effort attempt).
Fun Fact: google-ortools was written by the same people that wrote ILOG. You used to have to "bring your own search" when using ortools, which I used to assume was to avoid conflict with IBM. I haven't used it in a while, but looks like there's a one-size-fits-all search algorithm now. I'd be curious to try it sometime.
I use to write field service software for ruggedized Windows CE devices.
My understanding is that the full system reboot wouldn't have taken all that long, it's just the the company was trying to do a major fix while keeping whatever was still sort-of-working running. As any sysop will tell you, patching a running system is all kinds of crazy risky.
That being said, the software sucks. Southwest may have lost track of where their employees are. The ground crews are quitting. I wouldn’t be utterly shocked if management doesn’t even have a good overview of their the planes are.
(Obviously anyone halfway competent could hack up a script to find all the planes based on ADS-B data in a few hours. And it wouldn’t be terribly hard to text a link to all crew asking them to fill out a simple form with their location, nearest airport, and when they can get there. But this requires competence and agility.)
I'm confused, maybe my calibration of the scale involved here is wrong, but how do you not see this process taking days? Even if you had perfect information of where all the planes and crews were and where they should be, just coordinating with various airports to schedule the flights would take days; and that's assuming every plane is (1) flying empty, (2) fueled and (3) isn't carrying any passenger luggage.
The problem is nothing like writing a script to scrape ADS-B data. This is the classic fallacy of a programmer thinking the most cartoonish imagination of the technology is the problem while being completely blind to the fundamental difficulty of organizing large groups of humans in some activity.
Some of the observations here are correct, but there's a big missing piece. Southwest went too far in with an overoptimistic schedule. They won't say that because then there's no easy scapegoat, it's just a pure management thing.
The problem of recovery is also extensively studied. But it seems like SW did not put enough effort into having fault tolerant restart.
This whole event is a management decision.
As for ADS-B, that's hopelessly naive. ADS-B tracks planes, not people. The problem Southwest is having is trying to figure out who is legal to work and where they are. Tracking equipment is so trivial in comparison it's largely a secondary concern.
Once that complexity tipping point is reached (and other comments suppose that SA competitors were more aggressive at cancellations, avoiding that point in their systems) the system takes on a life of its own.
What follows, for those running the system, is called a "Significant Emotional Event".
It's simple to say, "oh it would be this easy..." and not consider a single real life scenario.
> problem could be simplified to getting every plain and every crew home to where they would spend the night under a normal schedule
This alone would take days under your own recommendation.
He's ignorant and naive and disrespectful of the workers involved.
If you are running a little video game that plots flight charts, sure perhaps you could write such a script. You cannot, however, script a bunch of logisitics support together across disparate, independent, multinational airports.
Add the obligitary paperwork (we don't really want to be shipping people around in coffins, enabling human trafficking or illegal substance smuggling, trust me), and add your standard bit of managerial incompetence managed from an excel document and sharepoint, and it's actually downright impressive if it's only a couple days to completely reboot the system.
Arguably, management doing the standard "oh that will never happen" is probably why it's not even better - you would think the airports would be able to automatically "fix" themselves with a self healing protocol, but that was probably deemed too expensive and left out of the feature set.
Of course this will introduce complexity. But it will also need decently less than a normal day’s number of operations.
If I were designing a solver for where to move aircraft and crews (which is apparently what SkySolver does), I would test it on random and malicious inputs. If it’s initialized with every crew member and every plane at an independent random airport in the US (where Southwest has service) with a few broken for good measure and a few airports shut down, it should still find a solution. Waiting 12-24 hours to start so none of the crew is timed out could be part of a valid solution.
And you can't just say "go to your usual starting airport". Flights are 24/7. Days of the week matter. Holidays matter. Even your local burger joint doesn't have cast-in-stone schedules. Much less a huge interconnected network that's trying to reboot. Even if 80% of the crew could go to a "usual airport", how would you know if you were in the 20%? Oh... everybody has to phone in, or the system has to send a huge number of notifications out. System crash. (Remember, normally the system depends on scheduling in advance, and only a few last-minute changes need to be handled.)
And how do these what, 20k+ people get where they should be? If they aren't already at that airport, then normally by hopping on a SWA flight. Which aren't happening. So now what? Book flights on other airlines? Who's going to be doing the booking? How does it get paid? How many employees have a company credit card (any?) Use the employee's cards? How many are maxed out for Christmas?
So let's suppose you're scheduled everyone and everybody knows where there's supposed to be and when. What happens when 10% call back in (without crashing the system!) and say they can't get there on time. (Remember, many SWA customers are stuck and are having a hard time arraning alternate travel. Crew without access to SWA flights would have the same trouble.) You're going to either apply massive changes to the existing schedule, or start all over and reschedule everyone. Which is exactly the problem they're currently having. They can't handle massive corrections. What's worse, a crew member may think they're good, and then their travel gets delayed somehow.
How sensitive is the scheduling? If 10% of the crew are no-show or delayed then about 9% of the planes are affected; every 8 hours 90 planes will have the following 8 hours of flights cancelled or delayed (throwing further monkey wrenches in the schedule). About 40 of those planes will be missing a pilot, which means they can't even be deadheaded to where they're supposed to be next.
So how to reboot? They've had to cancel two-thirds of their flights, so apparently they're able to keep 1/3 of them going. Keep those flying so they can shuttle crew around. You're initially only scheduling 1/3 of the crew, so the poor overloaded system can handle it. 2/3 of the flights are just outright cancelled, days in advance, so the customer support load is reduced. ("I'm sorry, your flight has been cancelled. We can't reschedule you until x days from now." vs. much back-and-forth trying to find something with an overloaded system.) Slowly add additional crew and flights so the number of phone-ins is kept manageable.
But this is HN, and, from an optimization / constraint satisfaction problem perspective / scale perspective, it’s really not a very large problem. The “going home” problem is a standard programming contest problem. This is more constrained, but a globally optimal solution isn’t needed. Southwest literally has a program called SkySolver. It should work.
So you need to choose destinations for a few thousand planes? Pick the place you would normally have them at 4 AM two days hence. Need captains familiar with the airports? Surely the captains who would normally fly from a given airport at 4 AM two days hence are familiar with those airports. This gives a good starting solution for further optimization.
You need to get a message to ~20k people? Great, Twilio will do that with minimal effort.
Will a budget of a couple million dollars, this is not hard. There are good SMT solvers available for free. CPLEX is expensive but not on this scale. Managing test cases on this scale is straightforward.
Even in an emergency, something good enough to get to the point where SkySolver starts working again should be doable.
People who manage electric grids have protocols for “black starts”. These protocols are messy and complicated, but they are developed in advance, and they work. Airlines should have the equivalent capability.
Finally, even with perfect communication there's the problem of getting people where they need to be when the transport system is very flaky. I think that's probably the issue when the CEO said 'they would get things manually set up, and then something would happen and we'd have to start all over again.'
The pilot's union and others have been critizing SWA for not investing in infrastructure. The previous CEO, responding to the previous melt-down, said something about 'you can't test for these kind of scenarios.' That's wrong. You can't do it easily or cheaply, but it is possible. They haven't done it, they haven't developed robust systems, and now they can't instantly resolve the problem.
i can't even begin to imagine what the utility function _alone_ must look like.
the solver can't just produce a blank slate solution every time, it's got to optimize... something, subject to some constraints about not changing too much from the previous plan it came up with. and it's probably got to do this every time there's new real world input.
> The “going home” problem is a standard programming contest problem.
the hubris and handwaving on here is hitting epic levels.
this is certainly a complicated situation.
uhh it’s really not a very large problem
Ok. So you need to choose destinations for a few thousand planes?
No. You need to figure out which planes are where, which of those planes are legal to fly out of which airports, which planes can be made legal to fly with low effort, and which planes need to go in for service (potentially requiring a non-revenue flight).Ostensibly Southwest operates a fleet entirely of Boeing 737s with two subfleets (ETOPS and non-ETOPS). The ETOPS planes can be used anywhere but the others cannot be used to Hawaii. Some of the ETOPS routes require the range of the MAX and some do not. Some of the airports Southwest flies into can only handle the smallest planes. It's entirely possible that everything stuck in Oakland is too big to fly into Burbank for instance. Or its entirely possible that all of the ETOPS fleet is stuck in Hawaii (so all of the routes too Hawaii are no-gos).
Then you've got to figure out what's broken on each plane. Depending on what's inoperative a plane may not be able to fly into a specific airport if the weather is anything but perfect. If enough stuff is broken the plane may be stuck away from a maintenance base and illegal to fly in revenue service. There goes a plane and a flight crew. These are laws, not optimizations.
And finally you need to figure out where all the diverted flights went. No idea what's normal but Southwest had to divert two flights today (one due to mechanical issues and one due to an unruly passenger). So there's at least one plane that's out of position and unable to fly until it's been repaired.
Need captains familiar with the airports?
No. In fact I'm pretty sure that unlike RyanAir and EasyJet, Southwest doesn't fly into any airports that categorically require specific familiarity and training beyond the ETOPS certs required for flight and maintenance crew on the Hawaii routes. You do need to figure out who's legal to fly in revenue service and who might be legal to fly on a ferry permit though. Southwest was pretty late to the autopilot game but at other airlines some captains are qualified to land in worse weather than others. So even if you have a flight crew ready to go they may be unable to complete your desired route today – and that's been a problem. The Southwest scheduling software assumes that every pilot completes their assigned flights successfully. Whoops.Then you've gotta do the same with the cabin crew. And then you've gotta make sure none of this runs afoul of the various CBAs. The baggage handlers? They've been working without a contract for nearly three years and were threatened with immediate termination by Southwest's VP of ops. Wanna guess how strictly the ramp rats will stick to the letter of the contract?
And you haven't even addressed the issue of out of position luggage. It's not just that luggage is piling up at airports, but Southwest's actually flown some of the luggage without the passengers. The hand waving I've seen suggests that Southwest still does a lot of the luggage tracking manually.
When that's all said and done you're still going to need to handle everything that goes tits up as service resumes. As I pointed out earlier, Southwest had at least two diversions today. So now you have two more out of position planes, crews, and more out of position luggage and passengers.
It's a massively complex, dynamic problem. If a couple million bucks would solve things, Southwest would've spent it already. List price on a single 737 is around $40 million. Even if Southwest got half off you're still talking tens of millions of dollars per plane. Put another way, Southwest has so far bought two airlines only to junk their entire fleet. A few million here and there is nothing, especially if it could've headed off this chaos.
People who manage electric grids have protocols for “black starts”. These protocols are
messy and complicated, but they are developed in advance, and they work. Airlines should
have the equivalent capability.
Most airlines do but, based on the anecdotes from Southwest crew, Southwest does not. Again it's not really that Southwest flies point-to-point, and it's not just a software issue at this point. Procedures that worked when Southwest was a scrappy little airline simply don't scale and Southwest management is ill-equipped to handle this. It's all compounded by the employee contempt that Kelly's managed to foment.You still have to contend with the laws of physics. AKA the speed of aircraft and distance between airports. Compound that thousands of planes making potentially cross country trips and it will take days to do what you are suggesting, leading to the same problem SWA is already facing.
And getting into a tech company and having high compensation doesn’t require “top talent”. Just the ability to memorize DS&A (junior to mid) and memorize system design and being able to regurgitate answers to behavioral questions showing “scope” and “impact”.
It takes knowing two or three well known algorithms to implement the scheduling problem they are trying to solve and system design chops to know how to scale it.
But, as far as having to be on hold for hours, that’s a staffing issue.
Most of their issues are symptoms of poor management - not software engineers.
Like I said earlier in the thread - there are many bad engineers that are overpaid and many good engineers that are underpaid. On average though, compensation and quality are correlated. In the absence of additional information, compensation is the best metric we have to asses the quality of the underlying engineering talent. It also passes the sniff test - try using the Southwest Airlines app!
And it doesn’t take a $400K senior developer that can reverse a binary tree on the whiteboard while juggling bowling balls while riding a unicycle on a tightrope to make a mobile app.
How much do you think Delta developers make in Atlanta (hint: they aren’t making over 200K). Their app is excellent.
It never ceases to amaze me how little most people on HN know about compensation throughout the industry.
2) COLA matters. $174k in, say, Southwest HQ in Dallas is roughly equated to $422k in SV.
I don’t think COLA matters all that much. The reason SV pays such high salaries reflects the underlying higher productivity of those engineers. The fact that much of that income goes to landlords in the Bay Area is orthogonal to the underlying quality of the engineers.
Like so many strong claims made online, it's an overly simplified mental model for something that is much more complex in reality. Like you said, we need more information before making strong claims.
It seems to me that just like pre-staged inventory helps in logistics management, that extra planes and crews in the rotation could improve operations under these circumstances.
I don't know how this is handled at Southwest, who does not fly hub-and-spoke and thus doesn't have a bunch of pilots sitting reserve around a base at, say, Atlanta.
That lack of slack is hell. It makes any disruption, even minor ones, require the other workers to work more time. Major disruptions mean soul-crushing crunch level hours to just get by.
Slack _must_ be planned into a system, otherwise there won't be any safety / recovery margin, and you're seeing the results live with Southwest's implosion.
Interesting. You don't say how far before Christmas you were traveling. Had this crazy weather system already started moving from West to East at that point? Or was the system buckling just from passenger volume at the point i.e similar to the Summer meltdown that Southwest had?
We had days like that in the UK but with the rail system. And it when it happens it’s also due to snow. We’ve seen it on global scale recently due to covid putting ships in the wrong places so that optimised shipping routes become a mess.
Naively, I'd assumed these kinds of things were handled in some sort of mission-control center with warnings from rule engines blinking on some big screen and a team of crack operators mapping out what needed done. But clearly that wasn't so: they were just making things up as they went along. Sounds like Southwest is in a similar spot, but this time on a much bigger scale.
I've been on a delayed plane where one of the pilots timed out while sitting at the gate, so we had another second delay finding another pilot.
You also need sufficient open gates of the right type in the right place during the right time window or you idle a plane (which can then idle a crew). It's pretty easy to imagine how such a system which may work at a certain level of load can be tipped into a runaway scenario.
As a systems person I'd be an interested reader if Southwest were to someday release an incident analysis and post-mortem of the type that good IT organizations do.
https://www.theonion.com/smart-qualified-people-behind-the-s...
In the end, there is no substitute for human judgement in the situation itself.
Yes? I suspect they have playbooks and/or run gameday exercises for other aspects of disaster preparedness and business continuity. Why not do the same for these sorts of issues that affect the travel network?
The network of airports, planes, and crews is a graph (or multiple graphs). What happens when a node disappears?
That may be a hard question to answer, but it seems cheaper to answer, and plan around, it before the situation occurs rather than after.
Where did you get your information? I have experience in the industry and scheduling logistics is clearly not how you describe it. The issue is that to optimize for profit you sacrifice the ability to maintain service through catastrophic events and can end up in a bit of a dominoes situation.
Or are you asking why I think they're making it up as they go along? That is my conclusion, to be sure. It's one thing for dominoes to fall, but if you are in a situation where the dominoes are falling and you are not able to predict which dominoes will fall next and respond accordingly, you are making things up as you go along. I'd have expected that almost any decision they could even hypothetically make would run through a system that checked it for violations of constraints, which would have told them way ahead of time that they needed a fresh crew, and they'd have had the entire flight of the replacement plane to get one in (IIRC it flew Miami->Boston just to get us the plane).
As for your conclusion, it's just not accurate. Systems _are_ used for scheduling and allocation of equipment and crew resources. You just have to realize that getting you, as an individual, to your destination on time on any given day isn't the number one priority of the company. There are a lot more concerns they are considering. If the airline industry is really as dysfunctional as you think, with such shortcomings in areas as important as equipment and crew scheduling and allocation, you should drop whatever you're doing and start an airline. You'll soon be filthy rich and do the world a great service.
Also your generalizations fly directly in the face of the information emerging from this Southwest fiasco, which indicates they don’t have systems in place to even track where their crews are (just where they are scheduled to be), so the shortcomings seem to be quite real.
Sounds like an argument for nationalization.
Airlines run on very thin margins, so need to optimize heavily.
We just need to accept that we can either get "affordable" tickets with occasional meltdowns or we go back to 30 years ago and pay double price for tickets but have more resilience.
The market as a whole has already resoundingly chosen the former.
Safely flying humans in metal structures is just really expensive and the current ticket prices pretty much reflect the real cost (and even ignore the climate externalities)
Who cares? The point of nationalization is costs don't matter, we just see it as a cost we bear because transportation is important for civilization. The primary mode (for better or worse) in the US is public (roads) and that "cost" is not covered by anyone.
No, it sounds like an argument to never fly SWA again.
> Unlike many rival airlines, Southwest’s planes generally hop from one city to another, rather than orbiting a major hub. That approach lets Southwest maximize use of its planes and crew, but the daisy chain structure also makes its network more delicate—problems in one corner of the country can be difficult to contain
That's in part what doomed the A380, which was popular with airlines still going strong with hub-and-spoke (Emirates being by far the most prominent one) but is worthless in a point-to-point model.
My worst nightmare, was pulling into a gate, and seeing one of those puppies arriving on the same concourse.
Immigration queues are bad enough, with 747s, but the A380 is much worse.
Major airports are major airports even in a point to point system. If everybody wants to go into or come out of NY, NY's airports are going to be slot constrained.
Also when I say "winding down" I don't mean "killed entirely", most airlines still have home bases, and mix point to point with regional hubs (especially for international flights).
But at the height of the hub-and-spokes model, unless you departed from or went to a hub you'd always need a layover.
From a net utility standpoint, NYC should just pay for "port packages" for nearby houses, add the cost to landing fees, and remove source/destination restrictions on La Guardia.
Further alot of major Hubs where imacted by the Storm, yet those airlines where able to transition. Why?
Well most airlines have mobile apps, and web portals and other techonology so their crews can be reassigned in almost real time, (just like I would get auto booked on a new flight before I even knew my flight was canceled via the mobile app)
Instead Southwest has systems from the 80's they require crews and customers to call and talk to an actual live human...
From what I understand (from /r/flying) it's the reporting / fixing of issues which is manual. This means when the reporting / fixing is overloaded and not given a respite it can't catch up, keeps getting more overloaded, and ultimately the scheduling system completely fell over (it lost track of crews entirely is what I understand).
"Mobile apps, and web portals and other technology" is orthogonal to the issue at hand.
The problem is that their disaster recovery plan (assuming that there is one) is highly centralized. Southwest has each and every body calling the scheduling department one at a time. The proper design (and what most other airlines actually do) is to decentralize a wee bit and have, for instance, crews report to station managers who in turn call the scheduling department.
The manual "reporting/fixing" of issues, which was done over the phone, includes pilots and crew members reporting "hey, I didn't actually get on the flight I was scheduled to". If the system did not receive that message, it was built to assume that the crew members made it to their next location and would be available for their next scheduled flight. But once a certain threshold of flights got canceled because of the storm, the limiting factor became how quickly staff members could even report that they didn't make it to their next location - they'd stay on hold for hours, until their next flight in their next city was scheduled to leave, and the system would still assume they were on that flight because it hadn't yet heard otherwise.
Lack of webapps certainly wasn't the only issue here - as the article notes, their whole point-to-point scheduling model is more vulnerable to cascading failure as well, vs. a hub-and-spoke model. But the phone-based system did certainly cause a feedback loop that kept the disruption happening for much longer than it should have. And moving away from point-to-point would have many more tradeoffs/downsides for consumers than simply modernizing their tech system.
And explaining that FAA or someone banned flying from this airport is much simpler.
https://www.insidehook.com/daily_brief/travel/airlines-fewes...
The pattern I've seen over the years is that upstart airlines, as they grow, end up having to look--for various reasons--a lot more like legacy carriers over time. Whether that's flying into areas with seasonal bad weather, flying out of default airports, instituting various forms of passenger status, etc.
1. On their flight search page, it shows the prices on a per-leg basis. Every other airline shows round trip prices. That way, the initial result on Southwest looks great and it’s only when you get to the final payment page that you realize it’s actually twice as expensive as you thought.
2. They refuse to share their data with any of the third party flight search engines like Google Flights or Expedia. Again, so people don’t realize how expensive they actually are.
b. other airlines will advertise prices that don't exist: it'll be the base price for the seat, but all (remaining) seats will be of the upcharged variety, so you can't pay the base price.
When you make the adjustments to compare apples-to-apples prices, SWA's get a bit closer. (But I do think SWA does tend to not be the cheapest, unless they're having a sale.)
I actually prefer the per-flight¹ pricing, too. That's what I'm buying: two flights.
¹IIRC, the pricing is per-flight, not per-leg/segment. (But it isn't, as you say, round-trip.)
And that’s why they refuse to share their prices with Expedia and Google Flights. If people could more easily compare prices, Southwest would lose a lot of sales.
SW has a high number of delays and cancellations BECAUSE THEY FLY A HIGH VOLUME OF PEOPLE. By percentage, they are in the middle of the pack for both delays and cancellations, which isn't great, but they are not the worst by any means.
How are HN people so consistently bad with basic information?
Here is the data which backs up the majority of these low effort cancellation related news articles.
https://public.tableau.com/app/profile/flightaware/viz/Airli...
Perhaps it's true nonetheless, but the numbers there won't tell you.
(And IME, it's perhaps true that SWA is often delayed … but by tolerable amounts. Compared to delays I've endured with Delta, where, e.g., a flight was delayed longer than the time it would take the plane to drive at highway speeds, from where it was coming from. Or … also Delta … where I was cancelled on twice in the same flight. They wanted to go 0-3 but I gave up and bought a ticket on … SWA.)
In the aviation arena, high reliability is maintained in part by careful analysis of "near failures": lessons are extracted and improvements are made to aircraft designs, procedures, etc.
By contrast, the "near failures" of the SWA system as a whole don't appear to have been utilized to motivate system improvements.
Frontlines have been emitting concerns and warnings for years but management didn’t care until an ops CEO (Bob Jordan) got in recently (as in 2022), but now there’s 20 years of ops neglect to deal with, and less than a year is nowhere near enough to start enacting real changes.
https://www.seat31b.com/2022/12/the-great-southwest-meltdown...
That said, I feel like these sorts of catastrophic ultra-fragile McKinsey-consulted-to-death failures we keep seeing in various industries are basically a giant signal to any adversaries that say "Hi! Check out how easy it would be to grind this entire industry to a halt!"
Resiliency is literally the opposite of efficiency. These systems need to have slack, aka inefficiency built into them. Unfortunately the business culture has moved towards ultra fragile, ultra efficient thinking.
Skysolver is a GE Flight Services trademark - there’s a video here showing how it works and SW planes. Contrary to the reddit claim, it does appear to use a predictive algorithm.
Highlight quote from the video:
“It is humanly impossible when there’s a major disruption for somebody to figure out what the optimal approach is to get them back on schedule”
I know nothing about the SkySolver or leadership at Southwest, but that seems like a very likely scenario based on what I've seen of what corporate types expect from software and rank-and-file employees.
Multiple things can be true here.
Manage the entire fleet. All tails, flight legs, and crew.
It was for a much smaller US airline but still a household name.
The crew aren't union so that helped. But it was tricky to manage the crew part.
If we assume the algo rocks (because you have operations research veterans), what stage of planning does your system address? What is the experience required and cost of manual intervention?
Fun problem. Human factors and edge cases abound. So much is at stake for a system that already works (to any predictable degree).
I'd love to hear more though.
The startup focusses on assigning goods to transport. (As in its not part of the commercial sales process.. only execution of the required capacity planning that comes after it) .Something that has been tried 20 times and failed. It currently is done manually.
So yeah, its going to be a challenge.
Sadly, it was a people problem, not a tech or algo problem.
Last I heard, the airline scrapped the entire project to continue doing things the old way because they felt it was too much work on their end. ¯\_(ツ)_/¯
Would love to pick your brain because I can relate to what youre saying.
Southwest's seating may be chaotic, but it's not United where you have to pay an extra $60 to not get a middle seat, likely in the back of the plane.
Do their planes still have smiles painted on them?
What I don’t understand is how come SW couldn’t enlist help to get customers rebooked on other airlines - their phone lines were slammed (I waited 3 hours) just to get a refund since their app wouldn’t allow me to choose to rebook/cancel.
If I was as customer focused as they say they are - I would’ve contacted AMEX global travel and gotten their entire network of booking agents to backfill and rebook customers on other flights.
"Sure, it's held together with rubber bands and is a mess, but it would cost hundreds of millions to fix, and it's working, isn't it? So the programmers complain a bit, that's their job."
Which works until conditions change in some way and it catastrophically does not.
I think a lot of our society is now run on unreliable fragile software. I expect to see a lot more of this. "Automation" is especially cost-savings when you don't min it being a fragile unreliable time-bomb.
I would expect any interview candidate to spot the issue within a minute.
For hub systems, ready crews are either at the hub, or at a spoke, ready to come back. That gives the hub a queue of ready crews, and each spoke can return a crew-plane combination to the hub when available. So with natural queue's, there's no delay cascade: it's all a function of whether and crew/plane readiness.
For point-to-point systems, crew-plane's are scattered, and the next flight opportunity might not be the next flight need. There is no buffer anywhere. Furthermore, any greedy/opportunistic strategy at one point can block a superior global solution.
That's the point-to-point trade-off taken by SWA. In the common case of good weather, you avoid the extra miles from going via hubs. But in the rare case of global weather shutdowns, there is no good recovery.
The only real question is whether SWA had any obligation to communicate this to investors and passengers. So far, Apple stock has gone down more than SouthWest's in this period, and passengers are remaining loyal, so no damage done.
Most ERP platforms are like this. over the decades it is not uncommon for little of the original platform to be used.
My org.s ERP is like that, we use probably 10% of the commercial code, and 90% of functions are custom in house written code.
It's worse than a framework because most frameworks are designed for people to do some form of software engineering with. But with some "configurable" tool, you end up with something that you need to do software engineering but it doesn't have the tools, power or affordances to do so.
I've done some Jira consulting and worked in MDM and I'm of the firm belief that "buy vs. build" isn't the real dichotomy: it's "build" vs. "buy and build", and the effort of building on top of the platform is usually underestimated and undervalued.
Companies are allergic to headcount
Working in enterprise, there is a lot of “modifiable off-the-shelf” software; ideally, this provides basic functionality out of the box, but also the alignment to custom business needs of in-house software.
(Often, it seems like it combines the up-front and ongoing external licensing cost of COTS software, with the internal development and maintenance costs of in-house software, and the combined problems of both.)
Good luck replacing it, on an operational airline that can’t shutdown. How you would do this is to me, just mind-boggling.
The article actually says "SkySolver, an off-the-shelf piece of software that Southwest has customized and updated".
I think a more obvious reasons is because of staffing issues brought on by covid, layoffs, and the vaccine mandates. They lost experienced employees who were able to wrangle the bad scheduling software. Throughout 2022, Southwest was having hiring issues because they were still mandating the vaccine through at least the summer for new employees. Their pilots association warned about this causing disruptions after a bunch of summer cancellations. Do people forget how flaky Southwest was during summer 2022? Southwest just recently reached staffing levels that matched their 2019 high. This "inadequate technology" narrative just seems like a convenient scapegoat.
My brother has worked in the white and blue collar unions (he prefers his ramp job). It's not like there's some impermeable cover of secrecy. These are just regular people who you can talk to. It's a combination of computer problems and regulatory controls (sleep blocks) leaving insufficient staff (and mechanical dangers) due to weather. The ramp teams were sitting at almost quad pay with no planes to service out of Minneapolis for a significant part of the weekend. This same situation has occurred, to some degree, every year.
Due to the inevitable Guld Stream collapse, this will be a routine problem until SWA triages it.
Employees could quit because they don't want to get vaccinated, but they also could have like just died from COVID too, or won the lottery, etc.
So to me this still points to the technology as something of a root cause. Your tech is as brittle as the number of people who know how to use it. Losing the people who make your sunk-cost old tech actually work, and not planning for the "bus factor' still makes it your fault for not addressing.
"Dieing from covid is like winning $1000 in the lottery"
// It sure isn't like winning $1 Million. // And it sure isn't a $2 winner // Somewhere in between.
If, on the other hand, it's not at all about software failures as many comments here suggest ("company management lost track of crews" notion), then does it have something to do with software at all?
Even less if commercial airplane pilot it seems. 100h is for transport
This is going to be a prime example CIO's can use as what happens when a company fails to invest in IT, treating as a cost center to alway gut budgets from
Somewhere there are Sysadmins, and Devs at southwest just face palming at the years and years they have beggged for resources to improve things
> Southwest said it will begin scanning bags on the tarmac this year to improve accuracy. Other large U.S. airlines already use this technology. Southwest's contract with bag handlers did not include a provision for using scanners until 2016. The airline then spent time evaluating technology and testing scanners at a few airports.
https://www.keranews.org/news/2019-02-20/southwest-airlines-...
SW decides what their priorities are when negotiating these contracts and deciding the terms. Other airlines obviously prioritized bag scanner tech and made the necessary tradeoffs to make it happen while SW didn't. Half the blame at the very least should be apportioned to SW.
These sorts of things rarely benefit customers.
I realized this on my travel this holiday, just considering the entire security theater apparatus in the US. It's not just that the lines are longer and slower; they added differentiated tiers of service for the line, and the longer waits and ban on liquids encourage more purchases once through the gate. All the players involved in the racket will look at these add-ons and say, "sounds good for my bottom line," even though it acts as a tax on service quality.
And most of these things will likely remain because the technical considerations and the political realm exist in an equilibrium; giving the consumer more options can't be done within the logistical and safety requirements of large-scale air travel, thus it's up to the whims of the political system to coordinate and come to an agreement. A "sky highway" technology could automate away a lot of the coordination in flight, and smaller/more efficient aerial vehicles with VTOL and low noise could present a lower footprint for air travel, but these disruptions to the model are still a long ways away from reality.
Then you tack on a hugely understaffed call center and problems with their phone systems and you have a hell-storm.
http://resources1.news.com.au/images/2013/08/02/1226690/4683...
But… why? What actually happened?
Essentially the system never know if a flight was successful or not, it just assume it went well, in the case of a delay or a change some employee has to input data manually into the system. For a few flights that's ok, but for a exceptional event like this the problem snowballs and now the system doesn't know were planes, crew and passengers are.
Quoting the top comment by 4Sammich here for those who refuse to touch Reddit:
> I have friends in CS and the hotel assignment side too. There were 2 specific problems, the software for scheduling is woefully antiquated by at least 20 years. No app/internet options, all manual entry and it has settings that you DO NOT CHANGE for fear of crashing it. Those settings create the automated flow as a crewmember is moving about their day, it doesn’t know you flew the leg DAL-MCO it just assumes it and moves your piece forward.
> In the event of a disruption you call scheduling and they manually adjust you. It does work, it just works for an airline 1/3 the size of SWA.
> So the storm came and it impacted ground ops so bad that many many crews were now “unaccounted” for and the system in place couldn’t keep up. Then it happened for several more days. By Xmas evening the CS department had essentially reached the inability to do anything but simple, one off assignments. And to make matters worse, the phone system was updated not too long ago and it was not working well.
> Last nite they did a web form and had planned to get the system up as much as possible with what communication they could muster, however it was too much to keep up on and ultimately the method for tracking crews failed again.
> This 100% is at the feet of all management who refused to invest in technology updates because it is the southwest way to be stuck in 1993. Heck, they still do 35 min turns on a -700 and 45 on an -800 frequently with only 2 man gates. But the good news is HDQ has a pickle ball court now.
> Edit: I just realized I never added the 2nd issue. Staffing. When the weather hit all those stations at once the ramp crews had to work in shifts to not become injured due to the cold. That slowed down the turns and backed up the planes. Many many ramp staff quit because of the management harassment (Denver) and just over it. So many rampers are new and making around 17/hr. Once they lost so much staff the crew scheduling software inputs couldn’t keep up because CS is also woefully understaffed and it became what we have today.
> This 100% is at the feet of all management who refused to invest in technology updates because it is the southwest way to be stuck in 1993.
https://www.reddit.com/r/SouthwestAirlines/comments/zxg6op/t...
The unnecessary, axe-to-grind editorializing undermines some of the credibility here, imo. There's no way this event hinged on a simple, binary "invest in tech" or "don't invest in tech," and that execs thumbed their nose at the latter.
There doesn't exist an automated way for the crew to log in the system that the flight is canceled.
The crew of the cancelled flight calls in to a call center in a central office. There's a person there on the other end. The crew tells them that SW1234 or whatever was canceled. The person manually enters into SkySolver that the flight was canceled. SkySolver runs the algorithms or whatever and reschedules/reshuffles planes and crews to ensure that all passengers get to where they're going. SkySolver automatically sends out the corrections to all the crews and agents what the new crew/aircraft/flight assignments are.
This works great if the number of canceled flights is less than the capacity of the call center. But if the number of canceled flights exceeds the capacity of the call center by 1 aircraft, SkySolver does not know that the canceled flight did not arrive at its destination, and is unable to perform its next leg. It does not know that there's still an aircraft taking ramp space, so when/if the next flight arrives they have nowhere to park. So your 1 cancelled flight turns into 3 cancelled flights. But we're already over the capacity of the call center; so those 3 cancelled flights turns into 9 cancelled flights. 27 cancelled flights. 81 cancelled flights. 243 cancelled flights.
All cancelled flights.
The "manual entry" system is already there, so not much seems to be needed other than adding this seemingly simple front end. You don't muck around with old business critical software, you add to it whenever possible. And this seems to be exactly like that. The existing manual entry system feels like it should be able to accept self-reports from a mobile app or web site with minimal changes to the existing system.
I feel for the employees
Southwest Airlines is a domestic carrier flying a point to point schedule, using the same fleet type to eliminate the need to train pilots on multiple aircraft type and also reduce training costs for transitional training from a narrow to a wide-body jet. All pilots can (theoretically) fly all planes in their fleet, tremendous training and labor savings. When flight crews "bid" on their flights, having only one aircraft type also reduces the complexities of bid lines down to basic seniority. (Let's ignore over water to Hawaii, m'kay?)
So, getting SWA back in the air should be over simplified compared to every carrier who got their planes and crews back flying the same or next day - get it? SWA has a single fleet and all of their crew is qualified to operate those jets or crew the cabin - so "wat da problem is"??? US Gov, inquiring minds want to know!!!
At all airlines, crews and planes are scheduled, and optimized, using decision support models which take into account how many hours the crew (pilots and flight attendants) can fly by law and contract; how long each tail# can fly until it needs to be flown into a maintenance base for an A/B/C/D check for scheduled maintenance, and how to get the maximum airtime out of the asset per day - in perfect weather conditions.
There are also decision support systems that monitor pricing of competitors every second of the day and why your airfares change from one browser refresh to another, yield management models which run overnight taking input from industry load data from the past year in the market SWA flies to help predict the passenger "load" which sets fares and also permits manual inputs for special events such as a World Series, or Super Bowl etc which would spike demand and drive airfares higher.
And there are Air Operations systems which are similar to the named product which take into account weather and crew events and help an airline re-plan based on where crew is currently. These Ops systems should have interfaces built to the crew (pilot and flight attendant) systems to know where they are located as well as where the jets are. Those values, along with the number of hours the crew has worked, would be used to re-calculate the crew and fleet assignments with the associated fleet and crew scheduling decision support models. These DSS ran on either big iron multiple CPU Unix servers or multiple CPU Linux servers - point being the computational power was outstanding and yes, CPLEX was typically a library utilized by our PhD's. There was a lot of money spent on the hardware and the people who developed these models - it wasn't cheap but then none of the clients ever experienced this type of problem with scheduling and yes, the clients are named in the weather impact article alongside SWA, but they are up and running either same or next day...
My world had a common database and data model where all data was integrated from various systems regardless of if it was our system or not because a decision support system without current and accurate data is like having corn cobs for toilet paper vs. Charmin... painful and not a lot of value. This is the one time I'm looking forward to the gov looking under the hood of private industry and revealing where the technical and management issues are. You can't blame the technical teams as they're only paid what SWA pays to hire both FTEs and contractors, and why I've never had a phone call that lasted beyond finding out what the position paid. We all know, you get what you pay for, but again, my opinion, and I'm greedy!!!
I just worry about what hell we’re creating. Software basically captures miscommunication and poor understanding and executes it.
Except when edge cases happen, and then we have problems.
Some economist should figure out what percentage of our productivity gains weren't really productivity gains at all, but rather risk acceptance.
E.g. if we are currently producing 1000 units per year, then a new plan in which we produce 1010 units 95% of the time, 1000 units 4.9% of the time, and then 1% of the time we'll produce only 400. And just as people do price comparisons of unreliable electricity generation to reliable generation without taking the reliability into account, the various executives that signed off on this assumed (correctly) that the public would see a gain and give them the appropriate bonuses. Wall Street did a similar thing during the Great Financial Crisis.
This type of analysis has already been done by Dean Baker for productivity gains due to trade deficits[1] - e.g. instead of building an iPhone, we just design the iPhone, and then call up China and ask "make us 100 million of these", and because the designers can capture so much value of the finished phone, it appears like an explosion of labor productivity. But it's all contingent on having good terms of trade and strong IP enforcement. If these were to decline, we'd capture much less value and suddenly the productivity of those designers would plummet, even though they are still doing the same work per unit of time, because productivity is measured in terms of value creation. So it could well be that on an apples to apples comparison, total productivity growth in the US from 2000 to 2020 has been flat or negative, and we've merely hidden the losses elsewhere, for example in unexpected outages, labor force problems, or periods of rapidly rising prices (revaluations).
[1] https://www.cepr.net/documents/publications/productivity_200...
Who or where are these people? (For any industry / system.)
In my experience, people understand things at a local level but quickly lose understanding at a system level. And I think there is a subtle difference from figuring out what went wrong once something did go wrong from understanding things prior to such an occurrence. Humans are pretty good at figuring out things after the fact, but we have a pretty poor track record of systematic understanding and preventing failures.
If you hear someone on a huge incident conference call leading the capturing of new information, threading their way from one subsystem to another (most of which are outside the person's usual silo) based upon what they see, paging in teams as necessary to get the answers they need, giving a blow-by-blow account of the troubleshooting picture they're currently painting, with reasoning behind what they think might be happening, how to test that, what possible next steps might be, instantly discarding anything that doesn't move everyone closer to resolution, crafting ad hoc "good enough" tooling on the spot as needed to get enough information to make a decision but not enough to know details to the last degree, freely admitting if they're on a subsystem they know nothing about yet able to quickly grasp the essentials in minutes from domain experts to work with those domain experts to pull just the information they need...you've found one of those people. Extremely rare to find in one person, usually you see 2-3 people in such incidents splitting such a dynamic.
They don't know "exactly what's going on". But they sufficiently grasp how systems work to approximate that outcome close enough for most business purposes.
It’s rare to find much teaching on complex systems in school so the hiring pipeline is fraught with peril and low tier companies understand the buzz but not the job. In my experience SRE is mostly people wander in and realize they’re in the right place, rather than setting out to be good at the required skills like managing complex systems and large scale incident response.
This is almost certainly a localized phenomenon to some specific West Coast cities.
What gives me a kick is the longtime sysadmin in my department who grumbles when I discuss introducing new software because "who will maintain it?"
Buddy, no one even understands the current software (except me a little bit)!
If I end up leaving a Ball of Mud behind, it is in total compliance with the organization's engineering practices.
Yup.
Instead what people do is find the lowest fare possible, all other considerations be damned. It’s not like this is the first time this has happened to SW.
Southwest did lose 10% of its market value overnight, so it seems the market is punishing them. Given how much executive compensation is in stock, a lot of senior level folks at Southwest are likely feeling some pain (or, least, as much pain as any very rich person ever feels).
> Instead what people do is find the lowest fare possible, all other considerations be damned.
That's an over-simplification. Yes, passengers are very price sensitive. But whenever I talk to friends and family about travel, it's become increasingly clear that they factor predicted quality of service in to their choice.
I don't fly United and try to avoid American. I'm certain that many many people will hesitate to fly on Southwest after this. People aren't stupid and no one wants to roll the dice on an entire vacation just to save twenty bucks.
A couple of VW execs went to prison and it cost the company 30 billion, that feels like getting hit hard enough by governments. What more do you want?
I can only hope the car I replaced it with also cheated on the emissions, because I'd love to sell it back for more than it's worth, too.
Yes, you can choose a more expensive airline but let’s not pretend Delta, United etc don’t also have huge delays and failures from time to time. Up until now Southwest’s reputation really hasn’t been bad, they’re no Frontier. No matter, there should be avenues for compensation.
https://www.faa.gov/about/office_org/headquarters_offices/at...
Presumably the consequence of regulators making canceling and delaying flights expensive is that airlines pay more to reduce the likelihood of that happening and fares go up as a result.
Whether that's good or bad is a matter of perspective I guess. Regulators could implement a lot of rules that would make flying more pleasant. But prices would go up.
See EU 261: https://help.ryanair.com/hc/en-us/articles/360017825538-EU-2...
See Indonesia: https://www.balidiscovery.com/delayed-flights-know-your-righ...
AFAIK they still need to make sure you get food and a place to sleep, though.
i'd like to see all airlines required to provide the sort of compensation that southwest is providing here
Further Southwest famously does not have interline or code sharing agreements with other carriers that would allow their agents to rebook passengers on a different airline. This is yet another detail of Southwest operations that is exacerbating the situation.
https://en.wikipedia.org/wiki/Flight_Compensation_Regulation
There’s a reason we give airlines a pass for unsafe weather: because we don’t want them killing people.
The point is that it is irresponsible to be running so lean that their entire system is tightly coupled like this, and we need to make doing so completely unprofitable to airlines.
The way I read it, you kind of are, just without saying it out loud:
> we could just make airlines always liable for any situation where they fail to get you to where they said they were going to get you by when they said they were going to get you. Like literally any service that is not delivered as promised.
What incentives does that set up when terrible (or even questionable) weather exists at the departure or destination airports or widespread en route?
This is a pretty huge incentive to get it right, and you can be sure every other airline will be looking into how they can prevent something similar from happening to them over the next few months.
Very coincidentally, their fleet size is 737.
The US should definitely have more rail--and higher speed rail--but that is only practical in select areas with enough population density. Many of the people in those areas could drive to where rail would service.
On the other hand, a high speed train can do 200 mph (320 km/h) in actual operation. That would be a six hour train ride. A plane is faster, but once you add all the airport overhead, the train might actually be competitive.
Their inability to get themselves out of it was clearly a failure that article discusses.
Anyone can quit, die, or become ill at any time. If they had to contract every employee into a binding contract where they must fly a given route on a given day at a given time, we'd be howling about how they're abusing their employees, and in fact that's not far off from what everyone's been complaining about the railroad industry doing.
Predicting the future, even a few days at a time, is always best-effort.