A corrupt file led to the FAA ground stoppage – also found in backup system
cnn.com
cnn.com
I, for one, appreciate the work that the FAA does to keep us safe in the air and appreciate that this was handled appropriately. Everything breaks. It's just a matter of how and when and what we do with it when it does. The FAA handled this outage appropriately and in a timely manner. I feel for the engineers who had to work on this incident.
It’s absolutely a corruption issue, in that the government prefers to pay 2-3x what they would to solve things internally to contractors who then perform a poor job and lobby to keep whatever they build in place for decades.
Being serious, I wish I could find one of these 2-3x multiple payouts in government. Every time I've looked at anything government related (including direct contractors) the pay is garbage. Usually 15% to 50% of what the private market pays for my skill set.
Govt > contractor > sub-contractor > employee was pretty common. I never knew of anyone being a direct contractor to the agency since the contracts were so large and involved a lot of employees.
The actual employer can pay the employed contractor whatever they want, but the rates are published and if they underpay too much the employee will be poached by a competitor. Because the rates are public info the employee can look up how much profit their employer is making any time they want.
The old answer used to be for the government to hire directly, but that's been hamstrung for like 40 years by now.
(Also, OP is probably underestimating the full-sheet cost of a federal FTE, as well as the complexities of fund-based budgeting and forecasting.)
If you were the FAA and started up an internal startup to find only the best to replace systems that have been running forever you would face a lot of problems.
1. Nobody would give you the funding until something like what happened yesterday did
2. Your developers getting paid more than the FAA directors will get lots of political attention
3. Even a great internal engineering team would most likely take ages to do something like this. This isn't move fast and break stuff with a greenfield, it is painful deconstruction and analysis of a very complicated system.
4. Many people won't want to work on this no matter how much you are paying.
It is also out of the wheel house of an agency tasked with flight safety. So expensive contractors arise. Do they do a poor job often? Yes. But these are jobs that are really hard to scope and execute. If it is really a case of overpaid contractors coming in and not doing the work, people here should start a startup and hire a SEAL team six of 1970's computer system rip and replacers and make a lot of money.
Seems FAA is pretty effective at reducing risk.
1) mostly thank the airline operators for that.
2) the fact the 737 Max disasters happened outside the US was a matter of chance not exemplary policies by the FAA. FAA policies did not prevent the deployment and roll out of this dangerous aircraft in the United States and did not lead to grounding of the aircraft until multiple events occurred.
3) I'm mostly talking about software here, and I believe a big part of this issue is shoe-horning software into the policies and procedures specifically designed for aviation. there are so many little pointless (in most contexts) requirements that cause the engineers working on these systems to lose the forest through the trees. the FAA creates its own complexity which prevents thinking holistically about our systems in a meaningful or effective manner.
in many ways this is a force of entropy. the more lines of code you have to support thousands of requirements and the more revisions you make to those lines of code without a top to bottom refactor the more likely you are to have insidious bugs that pop out like this incident.
This is a 100% proof that you have not the faintest clue what you're talking about.
The #1 priotity of airline operators is profit, not safety.
In theory, if you ask explicitly, of course they do. In practice, they choose the cheapest ticket and maybe avoid airlines that have had a high profile incident recently. They don't have the time or means to actually evaluate an airline's safety culture.
Meanwhile, executive decisions are driven by quarterly earnings reports, and you can cut a lot of corners for quite a number of quarters before your luck runs out and 200 people die.
My statement is, if anything, not emphatic enough.
What was the ultimate reason for the 737 MAX debacle? That airlines want to save money on type rating training.
Look at accident reports, and half the time the airline's safety culture (or lack thereof) is at least a contributing cause.
The FAA may be in many ways dysfunctional, but so are the airlines, and it's sure as hell not them who are pushing for better safety standards, it's the FAA and (especially) the NTSB.
But they do. And they've been so incredibly effective that we have collectively forgotten what risk used to feel like, so we're ready to say we don't need these standards/organizations or that they're not working. Obviously this is not to say they are perfect or free from criticism. But it's not just theatre.
What do you consider to be effective, and how would you differentiate it from theater?
Right up front, it seems necessary to account for the relative safety of flying if that safety has nothing to do with FAA policies.
It also seems that any safety policy that is effective enough will eventually appear indistinguishable from theater as people become more and more disconnected from the possibility of disaster.
Ozone mitigations come immediately to mind.
So Thank airlines for not crashing and do what when Boeing creates a death machine called 737 MAX??
On the surface, the relative safety of air travel and the lack of major stoppages over a span of 22 years seems like a major counter example.
You’re making this statement emphatically and authoritatively, though, so I’m curious to understand where that certainty comes from and how it accounts for the other publicly visible properties of the FAA and air travel.
1. The FAA basically handed their risk-management keys over to Boeing when authorizing the 737-MAX, contributing to those deaths (https://www.newyorker.com/news/our-columnists/how-boeing-and...)
2. The FAA's pilot medical vetting process, while thorough, is behind the times. There are people who took ADHD medicine in high school that are unable to obtain a medical certificate due to the FAA's overly-strict policies on prescription drugs. There are current pilots with serious mental issues who are afraid to see a doctor about them due to fear of losing their medical license (https://www.flyingmag.com/why-pilots-dont-want-to-talk-about...).
Couldn’t this also be interpreted as: when the FAA holds the keys, disasters like the 737-MAX tend not to happen? Obviously this raises questions about how that decision came about in the first place, but as an example, it seems counterproductive to your point, i.e. evidence that shifting away from some long standing policies directly led to harm, implying the original policies might have been better ones.
In a thread that seems eager to move fast and break things, this seems like a big problem, and would seem to indicate the need for a return to founding principles, not the opposite.
> 2. The FAA's pilot medical vetting process
This is an interesting one for sure, but also seems like an incredibly complex issue. Have there been studies about the safety of operating machinery while on those drugs that would obviate the need for a policy change?
The potential risk averted by such a policy would need to be weighed against the negative impacts of the 2nd order undesirable behaviors - obviously it’s bad that the policy discourages much needed mental health support, but how bad this is depends entirely on how effective the initial screening process is.
I’m not saying these mental health policies shouldn’t be changed, but neither do they seem to have obvious or measurably better alternatives at the moment.
And taken in the context of the original claim - that people are grossly misunderstanding the FAA and all of this is theater - they seem like weak examples to use as evidence of broad organizational failure.
Two planes crashed due to a design flaw - hundreds of people killed and one manufacturer and model forever tarnished like McDonald Douglass and their DC-10.
https://en.wikipedia.org/wiki/Professional_Air_Traffic_Contr...
>On August 5, following the PATCO workers' refusal to return to work, the Reagan administration fired the 11,345 striking air traffic controllers who had ignored the order, and banned them from federal service for life. In the wake of the strike and mass firings, the FAA was faced with the difficult task of hiring and training enough controllers to replace those that had been fired. Under normal conditions, it took three years to train new controllers. Until replacements could be trained, the vacant positions were temporarily filled with a mix of non-participating controllers, supervisors, staff personnel, some non-rated personnel, military controllers, and controllers transferred temporarily from other facilities. PATCO was decertified by the Federal Labor Relations Authority on October 22, 1981. The decision was appealed but to no avail, and attempts to use the courts to reverse the firings proved fruitless.
My late friend Ron Reisman worked at NASA Ames Research Center on air traffic control and flight safety, and he hired up a bunch of the professional air traffic controllers who Reagan fired, and taught them to program.
Because it's much easier to teach an air traffic controller how to program, than it is to teach a programmer how to control air traffic.
And we have them to thank for how safe the air traffic control system is today.
https://www.nasa.gov/50th/Folklife/biosPropulsion.html
>Ron Reisman has BA in Philosophy and Classical Greek, and an MS in Computer Science. He joined NASA Ames Research Center in 1988 as one of the original members of the Center Tracon Automation System development team. Since the late 1990s he has worked on traffic flow management research and development. He is currently supporting the Next Generation Air Traffic System research.
I saw him give an earlier version of this talk at the November 1989 Usenix Montery Graphics Conference, where he discussed training air traffic controllers to program, and he subsequently gave me a tour of the flight simulators and air traffic control systems at NASA Ames:
https://www.usenix.org/legacy/events/usenix02/usenix02.pdf
>INTRODUCTION TO AIR TRAFFIC MANAGEMENT SYSTEMS
>Ron Reisman and James Murphy, NASA Ames Research Center, and Rob Savoye, Seneca Software
>This introduction to air traffic control systems summarizes the operational characteristics of the principal Air Traffic Management (ATM) domains (i.e., en route, terminal area, surface control, and strategic traffic flow management) and the challenges of designing ATM decision support tools. The Traffic Flow Automation System (TFAS), a version of the Center TRACON Automation System (CTAS), will be examined. TFAS achieves portability across platforms (Solaris, HP/UX, and Linux) by adherence to software standards (ANSI, ISO, POSIX). Software engineering issues related to design, code reuse, portability, performance, and implementation are discussed.
Based on what I know about Ron's and other people's diligent methodological work on air traffic control and safety, I feel extremely safe and confident flying, and I find it insulting to his memory and the legacy of his work when the armchair architect ex-Facebook employees on this thread (and the GOP) glibly and patronizingly implore the FAA to "move fast and break things", as if they had no idea how many lives and fortunes are at stake.
Here's a video of Ron showing Marvin Minsky the flight simulator, an early AR headset, and the hydraulic lifts:
https://www.youtube.com/watch?v=mOKENF_-z8Y
I previously mentioned his earlier work making Apple ]['s talk with dolphins, and his thoughts on AI and dolphin intelligence:
https://news.ycombinator.com/item?id=32039126
To answer "What is AI?" you first have to answer "What is I?"
Check out the Apple ][ at the Dolphin Research Center in 1982, and Ron Reisman’s thoughts about dolphin intelligence:
https://www.youtube.com/watch?v=CWHCTNztnwQ&t=392s
“It’s gotten to the point that I never say anything about intelligence in general. I don’t know what it means any more. I used to. But then I started trying to test it. And if you think about it for a while, you don’t know what it is.” -Ron Reisman
"The FAA, citing lack of funding and resources, has over the years delegated increasing authority to Boeing to take on more of the work of certifying the safety of its own airplanes."
"There wasn’t a complete and proper review of the documents,” the former engineer added. “Review was rushed to reach certain certification dates.”
Seems like FAA certification was a disaster in the making.
https://www.seattletimes.com/business/boeing-aerospace/faile...
What you have to do is:
1. Not take those meds for at least 1-2 years.
2. Show documentation that getting off the meds hasn't impeded your performance. This generally means showing a stable work history if you've been off them for a long time or documents showing no change in performance between before you stopped taking them and x months after if you recently got off them.
3. Take an FAA ADHD re-evaluation.
Then they'll clear you. It's an annoying process but it's absolutely doable.
This line of reasoning is pure, 100%, cope. There is a reason this was broken and grounded flights nationwide. Providing cover for systems that allowed this to happen is not productive.
What does this even mean?
There’s a meaningful difference between providing cover and reminding people that major systems like the one in question are in a different category than the average web app.
The fact that such a nationwide stoppage hasn’t occurred since 2001 speaks to the stability of these systems, and I’m not sure what you’re suggesting here.
> There is a reason this was broken and grounded flights nationwide. Providing cover for systems that allowed this to happen is not productive.
I’m not trying to sound snarky here, but things do tend to break for reasons. Are you suggesting that there is a cure?
It is not perfect organization, but I think it deserves more credit than it is receiving in this thread.
That ain’t bad considering how infrequent incidents are in light of the immense complexity of the operation.
Much less complex and much more modern systems go down regularly and sometimes don’t come back up for much longer.
It is possible to have all of those mitigations in place and still experience a failure like this.
Post deployment validation is only as good as the validations executed. 99% coverage still leaves the door open to failure.
A DR strategy is just that - a strategy.
A failure of this sort is not an automatic implication that those things do not exist, just that they failed in this particular case.
I would find it incredibly surprising that an organization of that complexity could have survived as long as they did without a major incident if none of those things were in place.
They’d be either incredibly lucky, or incredibly competent, and if they are the latter, they would not operate without such mitigations in place.
It seems far more believable that an organization of the FAA’s age and complexity missed something along the way.
This seems like a bad case of binary thinking, and my point was that the occurrence of an incident like this is not sufficient to support that claim. It’s just as likely that an ancient process that wasn’t accounted for somewhere in the architecture broke down, and this is how it manifested.
Clearly improvements are needed, as is always the case after an outage. That doesn’t justify wild speculation.
Anecdote time: I once worked for a large financial institution that makes money when people swipe their credit cards. The system that authorizes purchases is ancient, battle tested, and undergoes minimal change because the cost of an outage could be measured in the millions of $ per minute.
Every change was scrutinized, reviewed by multiple groups, discussed with executives, and tested thoroughly. The same system underwent regular DR testing that involved quite a lot of involvement from all related teams.
So the day it went down, it was obviously a big deal, and raised all of the natural questions about how such a thing could occur.
Turns out it had an unknown transitive dependency on an internal server - a server that had not been rebooted in literally a decade. When that server was rebooted (I think it was a security group insisting it needed patches despite some strong reasons to avoid that when considering the architecture), some of the services never came back up, and everyone quickly learned that a very old change that predated almost everyone there established this unknown dependency.
The point of this story is really about the unknowability of sufficiently complex legacy enterprise systems.
All of the right processes and procedures won’t necessarily account for that seemingly inconsequential RPC call to an internal system implemented by a grizzled dev shortly before his retirement.
But none of that is really the point. The point is that even with every correct procedure in place, you’ll still encounter failures.
Modern dev teams in companies that build software have more checks and balances in place from the get go that help head off some categories of failure.
But when an organization is built on core tech born of the 80s/90s, there will always be dragons, regardless of the current active policies and procedures.
The problem is that the cost to replace some of these systems was inestimable.
It must be nice to sit behind your keyboard and just have all of the answers all day long! Do you have any tips for how to be so omniscient?
I'm not surprised. FAA does not fly each plane. Government organizational complexity helps ensure the government organization survives through next round of Congressional appropriations.
Org complexity + opaque oversight + 'safety' + 'homeland security' + taxpayer funded = playing around and more budget.
The pilot is responsible for safety. Air travel has rules to avoid collisions (eastbound gets altitude levels different than westbound, pilots shall broadcast on known frequencies) and pilots have distributed intelligence to keep their flight safe.
Yes, somehow there needs to be coordination of runway use. Many ways to provide reservations and queuing.
The only thing I'm left wondering at this point is whether the corrupt data were a config, or state.
Moving faster does not mean moving more effectively. Frantic activity should not be mistaken for progress.
I'm OK with the FAA taking "days to resolve the problem" if it means that nobody dies.
For what it's worth, my team and I are partial to "move carefully, and tend things".
In my experience, this tendency towards frantic is multiplied the larger and more complex the organization and architecture becomes.
The entire point of “move fast” in software circles is to leave behind the constraints of legacy tech and management practices in favor of building something “better”.
In a mature org that grew up before these ideas were mainstream, maybe one or two teams can manage to move faster, but invariably they end up depending on other teams, who in turn depend on deeply ingrained and established company culture and procedures.
We can talk about why those impediments are a Bad Thing, and I wouldn’t advise a consumer startup to adopt those methodologies in 2022, but there’s still the harsh reality that where they exist, “just move faster” doesn’t help much more than telling a depressed person to “just do cardio every day”. There’s often a lot of inner work that’s gotta happen to make way for the new.
The only way I’ve seen this sort of work in a large org is when a brand new “emerging tech” group is spun up and given autonomy to work outside of the legacy norms. This is not perfect either, and seems much better for greenfield projects. When applied to deeply entrenched legacy systems, all of the problems mentioned above come to a head.
This also creates a weird in/out group dynamic which tends to further stratify the old tech and widen the gap between the old practices and the new.
In the context of this particular conversation though, I think the concept of “move fast” has lost all meaning and has little to offer for an org like the FAA.
You should work smarter, not harder. Just turn up your smart knob. But why didn't you ever think of that before? Probably because you had your smart knob turned all the way down.
ESPECIALLY if millions of people's lives and fortunes are at stake, as in this case with air traffic control.
He has now moved on and we sometimes chat, he works for a huge corp, still sucks at writing SQL.
Are you proposing the entire world simply give up air travel, because government regulations and industry standards and the laws of physics and chaos theory prevent you from having the simple easy to understand air traffic control system you envision?
Does CI/CD exist? Does CI even exist? Are deployments automated? Is data sanity checked before loading? Is there a development environment? Do things like hourly snapshots exist? Can you easily provision a replacement system from scratch and restore data from a known good snapshot?
Or, is every process manual, slow, and error prone because there's never been a need to move fast.
Look at this one sentence in the article:
> In the overnight hours of Tuesday into Wednesday, FAA officials decided to shut down and reboot the main NOTAM system -- a significant decision, because the reboot can take about 90 minutes, according to the source.
So they do a reboot, that takes 90 minutes for some reason, and then that didn't even fix the problem. Their system that needs a lot of stability is now broken.
Please, don’t break people.
People are non-fungible.
It's still not working as designed, right now. The workarounds in place are fine, but were made up on the fly. Certainly there's room for improvement.
I'm seeing a lot of misuse of the word "risk". The FAA prioritizes safety over mission. The mishap that resulted in downtime affected the mission. The common-cause failure of the backup system affected the mission. That is not evidence that they're bad at managing safety risk.
Given that there was an article published at the time about removing the dissimilar backup, it's probable that they explicitly accepted this mission risk.
All the checklists in the world to prevent something from happening are fine and dandy until something happens anyway (which it will). And then they hamstring you from actually fixing it.
Instead, if you can move fast consistently, you can minimize the total downtime.
Eventually, but not as frequently.
If you can't move fast when things are working well, you can't move fast when things are broken. Acting like moving slow is going to prevent things from ever breaking is just wishful thinking.
In safety critical software where _a_ failure can result in loss of life, is “total cumulative duration of downtime” really the metric we’re optimizing for?
https://news.ycombinator.com/item?id=34338373
> ERAM’s original design did not include a dedicated backup system. FAA believed that ERAM did not need one due to the redundancy provided by the system’s dual channel design. This design was intended to prevent outages because it allows for seamless switching between the two channels without impacting air traffic control should a problem occur in the active channel. The Agency believed this redundancy would make it unlikely that a problem in one channel could migrate to the other channel. In addition, to provide additional backup capabilities during ERAM’s implementation, FAA planned to temporarily maintain EBUS, its pre-existing backup system, before phasing it out completely beginning in 2015, to rely solely on ERAM’s dual channels.
> However, problems experienced during and since ERAM’s implementation have shown that the system remains susceptible to dual channel failures. As a result, FAA decided to maintain EBUS much longer than intended because air traffic controllers currently rely on EBUS for backup. However, with FAA’s ongoing and planned upgrades to ERAM, which will span the next 7 years, EBUS will soon become incompatible with the new hardware. As such, FAA plans to begin phasing out EBUS in April 2019, leaving ERAM without a backup system to supplement the system’s redundant dual channels.
A $2 billion dollar Lockheed Martin system where a memory overflow takes down BOTH the main system and the backup system. Surely that's not a signal of fragility.
Obviously at a big government agency, the same person is not writing both aviation regulations and software procurement contracts, but the institutional knowledge is there. Nobody thought to be as paranoid about software as they are about planes flying over the ocean, but honestly, paranoia is good if you want reliability.
But most orgs can't realistically have two distinct software systems. How do you create proper isolation or failure mechanisms between them?
I'm guessing this sort of thing is what you mean by their experience with ETOPS.
It’s not bad design, it’s just you have to understand what resiliency you have and plan against each of various such threats according to your risk appetite.
you are pointing out that they prioritized that planes in the air could land safely over allowing more planes to take off. I'm actually quite reassured now.
They have excellent intuition around making things redundant to single pieces of hardware failing but don’t really grok making stuff resilient to wider failures.
Anything involving transaction logs, rollbacks, and plain old backups take a backseat to live hardware-redundant environments. “It’s OK though because we follow the NASA software development process which has a rigorous set of validation steps that prevent bugs.”
I always feel like making single components redundant is a fairly well-defined process -- generally speaking, the mechanisms are the same (1+ redundant components, failover, STONITH, etc), where making things resilient on a higher level is not as well-defined, and often requires bespoke solutions to each unique situation.
BFT state machine replication is well-understood and well-defined: use N of M agreement for inputs and run them through a deterministic state machine. Optionally, do N of M signature of outputs.
OTOH what are properties of failover? "Failover" seems like an attempt to cheat on Byzantine generals' problem: Generals send mail and the confirm results in a Zoom call. But what if Zoom doesn't work? What are the assumptions for 1+ redundant components/failover/STONITH?
It's just that production software essentially never used formal verification to a sufficient extend.
imagine a spreadsheet with 700 lines in it telling you that you need to do ABCDEFG each of those lines is instructing you to write a document detailing a procedure with the chain of custody forms and keys and whatever password rotations etc etc. follow this process and fill out this presentation template wait 3 months and then present it to a board who doesn't give a shit or have any understanding of your project.
it's all an insane amount of work and it gets us the opposite in terms of the goals these processes are intended to achieve.
I have zero doubt that the only thing that will happen as a result of this catastrophic incident is another 50 lines in a spreadsheet somewhere telling you to do stuff that nobody will comprehend or implement correctly or even verify until there is a similar incident causing an investigation into it.
If you're firm in your view that that's reasonable, I'd like to learn more about the proposal as to how the math would work.
I'm not going to code up a simulation because the research hasn't been done to confirm my choice of constants, but I can sketch it. Each workday is a function of the macroeconomic climate and some set of cultural norms during which we exhibit some blend of the following personae. As we'll see, introducing UBI reduces the prevalence of the bureaucrat persona which has knock-on effects leading to surplus.
---
The Missionary - has a mission and is working towards it. Cares more about the mission than prestige.
The Worker - doesn't have a plan, but likes to be a part of something meaningful. Will gamble with prestige in order to ensure that the work stays meaningful.
The Bureaucrat - willing to tolerate or create waste in favor of preserving prestige. Sometimes manages to trick a worker into believing they're a missionary.
---
Obviously people are more complex than this. Also, I'll use dollars to indicate productive output even though I think that most of the time collapsing such things to a single dimension is a slippery slope to somewhere awful. All this to say: gimme a break, it's model.
Here are my totally made up constants, note that X is a parameter which will depend on UBI:
---
Missionary creates 100$ of output always, plus a 1% daily chance to inspire a worker to become a missionary, a 1% chance to inspire a bureaucrat to become a worker, and a 1% chance to burn out and become a worker.
Worker creates 80$ of output if they're following a missionary and -20$ if they're following a bureaucrat because it's likely that they're causing more harm than good. They have an X% chance of burning out and becoming a bureaucrat.
A Bureaucrat creates -$20 of output, because they're definitely doing more harm than good.
Now lets say that everybody consumes $5 each day to stay alive.
---
So X is our worker burn-out rate.
As with most systems of this kind, it's very sensitive to initial conditions. If you start with a high enough concentration of workers and missionaries, your bureaucrat rate will be very low and you'll have a surplus. Too many bureaucrats and most of your workers are doing more harm than good, the system is carried (if it survives at all) by the missionaries and the minority of workers following them.
Critically, X is a function of risk tolerance. The worker becomes a bureaucrat because they cannot tolerate the risk of pointing out the wastefulness of the bureaucrat above them.
Introducing UBI does two things. It makes standing up to your Bureaucrat less risky, reducing X, and it creates a fourth type, the Video Gamer, who consumes $5 to stay alive but doesn't sabotage the output of any workers like the bureaucrat does.
Some percentage of the Bureaucrats will become Video Gamers if UBI is implemented. That percent depends on the size of the surplus. If the surplus gets big enough, UBI can be so comfortable that there's no reason to be a bureaucrat, because it doesn't afford a significant quality of life increase.
---
So to answer your question about the 3M and the 210M, I'd guess that today we've got 213M people living on the positive output of maybe 50M--the rest are bureaucrats or are following bureaucrats. They're busy fighting over their slice of the pie instead of baking it. Bureaucrats sort of expand to consume available resources, so as automation improves worker output, that ratio will get worse unless we find a place to put them.
We'd have to do research to come up with better constants and run that model for real to be sure, but I don't think it's unreasonable to assume that reducing both the bureaucrat concentration and the worker burnout rate by 50% would triple the system's output once you let the personae conventions find a new equilibrium. I'm not sure how much more federal employees will get paid above UBI, but I think there's room for the end result to be that future UBI is cushier than today's government work.
They're just letting it be inflationary and setting the payout to increase over time to adjust for inflation. So maybe you get $5 per week this year and $8 per week next year... This can be balanced so that it amounts to a more or less constant purchasing power.
Personally I prefer the demurrage approach where account balances just have a decay rate--that way you've got a better shot at $5 written down today having the same meaning to people who read it next year, but the economics are the same (more on the theory here: http://en.trm.creationmonetaire.info/ ).
It's gotta be decoupled from the government so that, as discussed in my model, it can act as a safety net while you're ridding yourself of wasteful bureaucracy. It doesn't really work if the bureaucrat you're deposing can threaten to take away your UBI.
"We only hire the absolute best" does not lead to "move fast and break things" (which sounds awful in a FAA context anyway) it leads to people who devoted their lives to being the absolute best at coloring inside the lines and never straying off the path, to being the best follower out there, to the ultimate authoritarians desiring to grow into being the authority.
The heaviest selection pressure usually does not lead to the most efficient system, it generally leads to a system able to endure heavy selection pressure.
There's a sociologist who wrote a famous book about bureaucracy and its in my library at home and the name of the sociologist and his book are at the tip of my tongue but he wasn't near the top of a quick google search; the above is a paraphrase of his book. No its not Douglas Adams or even Scott Adams although those two are correct about the problem in general LOL.
That's one problem with treating the implementer as a machine to run code. The whole procedure can't be tested, so when parts are changed they can break the whole. It relies on the human in the loop to resolve the conflicts, which is not repeatable.
The other problem is the "mind-numbing" part. No-one can maintain 100% perfection all of the time. And in the context of presenting to people who don't know what it all means, I can see why mistakes would be made.
If problems of circularity arise they can try to clarify the issue, or go back to the drawing board.
Generally you need some flexibility to handle slight variations in circumstance whenever making a decision, and at times things come down to judgement calls that can not be turned into an algorithm. But bureaucracies don't like empowering their workers to make decisions, and so you get ever more conoluted instructions to shift the decision making process higher up the ladder.
Where did you acquire this notion?
We all know this from the subject the FAA regulates: flights. New unleaded gas gets forever to approve. Simple changes to instruments takes years of certification, leaving 1960s technology in place when clear improvements have happened in the last 60 years.
I always wondered what would happen if this culture were carried over to another space. We know what it looks like in medicine because the FDA has similar priorities. Rarely do we see it in tech, which is known to "move fast and break stuff". But here we get a glimpse of the dystopian crossover between FAA-procedure-culture and software engineering.
Which also suggests to me that the problem isn't per-se the glacial progress & approval processes, but that they're using the wrong glacial processes.
If this FAA software was developed with something like the Space Shuttle's process it would still take forever to change something, but at least you'd end up with a function that you could mathematically prove would be able to handle any conceivable input.
I think you're never going to convince the government of a "move fast" culture. Even if the overall cost of grounding planes would be less than the cost of more reliability they'd never go for it.
Too many people's asses are on the line, and they're not having to spend their own money. The only thing they have to "pay" for is possible loss of face, or loss of political capital, both of which can be insured against effectively for free with taxpayer dollars.
But you might just be able to convince them that they're using the wrong sort of bureaucracy. You'd still spend a billion or two on something that should cost a million, but at least you'd get actual reliability as a result.
But that team had kind of an "unfair" advantage in that were able to program in assembly code on bare metal with no real software stack. Whereas the rest of us are forced to build on a foundation of sand using multiple layers of low-quality third-party software in order to deliver any useful functionality.
The Space Shuttle’s avionics software was not written in assembly, rather HAL/S, a high-level language invented for the project. Assembly was mainly used for the custom real-time OS kernel. They also maintained their HAL/S toolchain, which was written in XPL-a PL/I dialect which was popular for compiler development in the 1970s. The development environment ran on IBM mainframes, and the main CPUs on the Shuttles were the aerospace derivatives of the IBM S/360 mainframe architecture, System/4pi, model AP-101. The same CPUs were used by USAF (e.g. the B-1 bomber), but USAF mainly used JOVIAL to program theirs. Another big user of JOVIAL was the FAA, who used it to write a lot of their original mainframe-based air traffic control software (FAA HOST).
The Space Shuttle team inventing their own programming language was a byproduct of the time the project started (1970s). If they’d started a decade later, they probably would have used Ada instead. But Ada didn’t exist yet, and they thought inventing their own language was a better choice than JOVIAL
oh and in my experience, government fte's can screw up an infinite amount of times with no risk to their job.
Weird, my impression was federal employment is the opposite: mess up as much as you want and you’ll be retrained and reassigned.
Or a that only in contexts like harassment or drug use?
For routine FTEs, as a sibling post mentioned, you can probably mess up in a lot of ways (short of murder or being really non-PC) and still keep your job.
The closest thing you'll see in tech will be at back-end software like banks, insurance companies, pension funds, investment companies, embedded engineering / SCADA, ERP systems, etc. The "move fast and break things" mindset seems to mainly be a thing in Silicon Valley internet companies / start-ups, and mainly the latter because they're driven to pump up their own value instead of provide reliable software. Because if Twitter is down, it's an inconvenience, but if airplane software fails, lives are on the line.
There is a difference from risk averse ("we require heavy testing before deployment") and dysfunctional ("we are terrified of breaking anything but are unwilling to invest in maintenance").
Take the Air Force. Their risk decision is: can we accomplish a specific mission at hand with ideally minimal loss of warfigher life. If they don't maintain their planes, they cannot achieve the warfighter life loss minimization.
Similar with IT generally. Kicking the can down the road just grows the problem.
Risk aversion only works for an entity that has a forcible monopoly in its space. The FAA and FDA do. Another is the Nuclear Regulatory Commission, which exhibits the same behavior: their job is to prevent accidents, and the surest way to do that is to never approve anything at all.
The FAA is frequently called a "tombstone agency" - who only act long after fatal accidents and the bodies have been buried.
I guess even that wouldnt answer it, since this isnt really a problem in private industry.
About ten years ago, a new manager was brought in to make us act less like a moribund government department and behave more efficiently. As an example of government waste, he pointed to the money we were spending on storage for data back-ups. We wouldn't need back-ups if we stopped making mistakes.
You might think that this no-back-ups policy would be an instant disaster, but it lasted years without issue. When there was a failure, the manager would hand the sys-admin a soldering iron, the admin would fix the hard drive, and we would be back on track. Finally, the sys-admin retired and a new one replaced him. Not long after, a critical system failed and data that we were required by law to maintain was lost. The manager handed the admin a soldering iron and told him to fix the hard drive. The admin said it was impossible and the manager fired him (yes, you can get fired from a government job). Other candidates were interviewed, but no one applying for a $30k job was confident that they could repair a broken hard drive.
Finally, there was talking of hiring the old admin to come out of retirement and fix the drive. Except he explained that it had always been impossible. During his tenure, he'd spent 5% of his salary (gross, not net) paying for back-ups and replacement drives. When the manager gave him a soldering iron, he'd just chuck out the old drive, by a replacement off Newegg with his personal credit card, and load it with the data he'd backed up to his personal S3 storage. His back-up script was still running on the server, but he'd stopped paying for the storage space the moment he retired.
Eventually, the manager was forced to spend a whole year's budget on an expensive data-retrieval firm to collect the data (which was still cost an order of magnitude less than the fine the department would have had to pay if we'd lost the data). He was fired and a new manager brought on board. Because of the money which had been lost on the data retrieval, new measures were put in place to prevent this from ever happening again. This included a new back-up system and audits to ensure that other employees were using personal funds to pay for departmental resources. Of course, this meant rigorously documenting exactly what resources each employee was using...
Six years after the manager was brought in to decrease cost and increase agility, we were now more over budget and tightly controlled than we'd ever been.
And even after $2B it was completely broken and late and required another year and hundreds of millions.
I get the feeling people see thousands of millions of dollars as some abstract thing. You could build an A+ team with 1/100th that cash. Yet it still ends up sucked down a blackhole that's only designed to consume more money.
And there's zero consequences for failure. The same few contractors will get the contract next time.
ERAM is employed the the 23 air route traffic control centers [1] throughout the nation as their primary operating system. If there were a system-wide outage of ERAM, the consequences would be magnitudes more consequential than any NOTAM outage. Basically every flight in the air and not close to a terminal facility would lose radar contact and controllers would be working blind, causing widespread chaos and likely many safety incidents. Non-radar air traffic control is a thing, but generally controllers do not have adequate training or currency to do it safely, and definitely not at anywhere near normal capacity.
That’s not speaking poor of the engineers (which in my experience can be very good) but of the management and innovation culture of these agencies, which is too often terribly broken. They would say they are “risk averse” but as yesterday highlights their poor approach to this creates a ton of risk.
It's like making a crap sandwich and filling out a bunch of forms proving that it's not crap sandwich, rather than just spending the time you would be filling out all those forms on i don't know...not making a crap sandwich.
Obviously the NOTAM system isn't up to scratch, that's why we're talking about it. But "hard problems are easy" isn't constructive nor very realistic. I bet someone could put in a three-hour presentation talking about all the complexity that led us to this conclusion.
Its optimized for inefficiency.
That "owner" is also a federal employee, not a contractor.
https://en.wikipedia.org/wiki/Federal_Information_Security_M...
This stuff is such cringe, it is the type of response you see from jaded boomers more focused on box-ticking and punching out as opposed to doing anything new.
This is a situation where tossing the whole damn thing out and starting over again would be productive. The systems that lead to the creation of these half-fossilized government projects (that still don't work!) will not change and needs to be tossed.
The biggest issues in general relate to the data itself and the processes involved.
You mean the quote that you just made up?
Yesterday should have been 'run this on these 2 old boxes and make sure they still work' regression test. Sounds like either that step is not there, totally skipped, or does not match reality. Building a real five 9s type system means taking each piece and pondering 'what are the different ways this can fail'. Then mitigating each of those. It is mind numbing tedious work than takes a long time to do. Also most of this is probably run by contract houses. Which means the people who use it do not really 'own it'. Which is by design, for CYA. Which costs more and takes more time because it is all paperwork.
Of course you also have "old but still developed, and likely to be supported in near-perpetuity". Things like SQLite fall in this category. Unfortunately there's this other problem of management being inexplicably down on anything open-source. I don't know whether they've been exposed to too much FUD from vendors selling proprietary solutions, or just that the idea that security through obscurity is no security at all has failed to reach them.
The vendor simply changed the procedures to say the user must use a two-fingers procedure to simultaneously flip both power supply switches on or both off at once. It bothers me that we still don't know WHY the original problem happened. How do we know there isn't electrical damage occurring during the brief period between the two switch contacts (since no human can do that perfectly)? If it's a software problem, what is the most time one power supply can be up and one down before the bug is triggered? That's not been characterized, to my knowledge.
But in the context of this thread, what's relevant is to avoid addressing the problem the vendor changed the meaning of "dual power supplies" from an OR condition to an AND condition! They met the letter of the spec while completely violating the spirit of the requirement.
edit: Better reporting of Canada's issue from Canada's CBC (and frankly, better reporting about the US, too). [2] "In Canada, pilots were still able to read NOTAMs, but there was an outage that meant new notices couldn't be entered into the system, NAV Canada said on social media." "NAV Canada said it did not believe the outage was related to the one in the U.S., but it said it was investigating."
[1] https://www.independent.co.uk/news/world/americas/canada-fli... HN thread https://news.ycombinator.com/item?id=34347520
[2] https://www.cbc.ca/news/us-air-travel-chaos-notam-outage-1.6...
Lil' Bobby Tables strikes again!
> Speculation: corrupting input, either international or North American? E.G. UTF-8, SQL escape, CSV quoting.
I read it as filesystem corruption from a bad disk, coupled with redundancy that doesn't actually work.
I would not expect a disk failure to replicate to the backup.
Edit: Should add "won't be silently propagated"
Comprehensive DR testing is really difficult. Many orgs settle for “on paper,” or “in theory” substitutions for real testing.
They do it right; no problem.
Doing it right, though … there’s the rub …
But they said nothing about it being bad drive, just corrupted data file, which very well might be software bug or operator error
It's not typically a performance concern because computing checksums is fast on modern hardware. Besides, historically IO was much slower than CPU.
It is expensive. It might be prohibitive in a very competitive environment. This is hardly the case here. Safety first!
There are databases that maintain redundant copies and can tolerate disk / replica failure. e.g. Cassandra.
That's different from filesystems that do checksumming (zfs, btrfs). Those can detect corruption.
In any case, if you use a database it handles these things by itself (see ACID). However I don't believe they can necessarily detect disk corruption in all cases (like checksumming file systems).
Oh you mean they should be testing/validating the generated backup db file before replicating it to long-term archive ...
No, Restore is a use case.
(Replace "use case" with "requirement" or "user story"...)
That said, I do remember this being an issue even with plain text.
So, if you have a database of some sort, a part of the backup process would be checking that the backup can be used to run an instance of it. If you have images or videos or PDF files, all of those should be validated as well, to make sure that they're not corrupted (e.g. display/playback fails). If you have compression as a part of the backup process, then the compressed archive should also be validated, that it's not corrupt and that the data can also be extracted.
Only then would checksums of the backup be actually useful, once you'd know that you don't just have a bunch of mostly useless binary data on hand. Sadly, the tooling just isn't there (e.g. CLI utilities for validating everything) and there are far more file formats than most people want to concern themselves with.
If you don't have an option to restore the data your backup might as well not exist.
I've been trying to toot this horn at every place I've consulted/worked at, with the most common response from software engineers being: "We're using the AWS backup system of the database, of course it works, no need to test it"
People have way too much trust in systems overall, especially ones they have no inside information about. And people also seems to default to "Of course it works" without testing it, instead of "How would I know it works if I haven't tested it?" or even "I don't know".
1. daily encrypted restic backups to my local NAS, running ZFS raidz2
2. weekly rsyncs of the backup directory to an offsite machine
3. yearly full offsites of the entire NAS
I test my backups frequently if only because I have some odd persistent bug that eats my zsh history file about once a month. I now have a "recover-latest-history" script that pulls that file back down from the latest restic backup. I have just completed an offsite and exchanged it for last year's, so I can test that now too!
Rate my tao :)
FAA NOTAM System Outage - https://news.ycombinator.com/item?id=34337158 - Jan 2023 (206 comments)
US Halts Flights Nationwide After Key FAA System Goes Down - https://news.ycombinator.com/item?id=34337807 - Jan 2023 (36 comments)
Your framework is out of date, go use whatever the latest js framework published today is.
Do you hate JavaScript programmers?
Do you also agree with your friends how of people of different color are all the same?
Do you feel the need to call out other people’s sexual orientation, at least behind their back when only your friends are hearing?
Some of your friends might be JavaScript programmers, you know.
I'll take a book to the observation car and read next to the Mennonites doing their knitting and love every minute of it, no worries from me.
A personal preference.
On the train there's more room, you can move around, and the atmosphere is more relaxed.
Some people like planes. I saw an HN comment where someone said they took lots of "flights to nowhere" over COVID and used the plane as their office. I don't understand that at all. But to each their own.
I will say the views are unbeatable.
It's hard to agree with that. There are some amazing train routes through Austria, Norway, Switzerland, Germany that keep me looking out of the window the whole way.
The ground fading away, the tiny cars, glittering rivers, seeing it in reverse for landing. That is something I'll wax poetic about. But if I could take the train the rest of my life I would.
Clearly the middle ground here is to bring back zeppelins. /s
I was waiting for a Ryanair flight from Friedrichshafen when one of the Zeppelin NT craft flew past, heading over the Bodensee towards Switzerland. It looked simultaneously classic and futuristic, the sort of shiny utopia future of 1950s sci-fi book covers. I envied HARD.
https://www.tsa.gov/travel/security-screening/whatcanibring/...
Final decision vests with the TSA agent at the desk.
But there are needles which are short metal/plastic parts, with a long flexible back (eg for circular knitting) which are very unlikely to cause an issue: the point is smaller than a pen.
But I'd like to see more investment in rail & a better coach experience, absolutely.
Amtrak isn’t allowed to be successful
Either way rather be stuck on a train then a plane anyday
The actual sane solution is to shut down those routes and run buses. American heavy rail outside the NEC is better suited for freight anyway. Build new dedicated corridors for fast passenger services, or stay home.
By contrast, on the railroads, Burlington Northern Santa Fe or whoever purchased the rail from the decrepit husk of bankruptcy that preceded it, and upgraded it, and installed all the new safety equipment, and maintained it over time. The capital involved here is, by and large, private.
So there shouldn’t be any surprise that the situations are different.
I don’t know why you want to drag in manufacturers, though. It seems to muddy the water.
The reason the freight companies own those rails is primarily from the land grants made my the federal government which came with obligations, such as providing passenger service. Amtrak has trackage rights because the freight companies wanted to divest their passenger rail operations. Maybe an unwelcome guest, but essentially a former part of their own operations.
Anyhow, buses are an insufficient substitute for passenger trains and already available from other operators. If they were sufficient people wouldn't be on the trains.
But the argument of “is this justified given the history?!” should take a back seat; certainly Congress is able to force the industry’s hand whether or not it’s justified. The first question should be whether it’s a good idea in the first place, since rail is doing useful things for the US economy, which would ultimately shoulder more costs for it than the railroads themselves as a business.
Damaging the supply chain and raising prices across the economy while putting more trucks on the taxpayer-funded roads emitting more carbon dioxide is a steep price. Incremental improvements to the reliability of seldom-used cross-country routes through the sparsely inhabited West at speeds of about 60mph aren’t worth that price.
If only, one of the best investments for NA would be an actual HSR network outside freight. California is trying and there are so many people who are doing everything they can to stop it. Including musk inventing an impossible alternative he never invented to build[1], hyperloop, as HSR would compete with tesla and hurt his sales.
[1] https://www.fresnobee.com/opinion/editorials/article26445107...
Half a dozen news stories about passengers getting trapped on planes that were sitting on the tarmac from the past 30 days alone.
I'll take being "trapped" on a train that has a cafe car, numerous bathrooms, power outlets, likely cell phone service or wifi, plenty of room to walk around, and considerably more leg room and seat-reclining...over being trapped in a metal tube, crunched into a tiny seat, with a bathroom that probably won't function past a few hours, limited food, no power, and nowhere to get up and walk around.
Not to mention, if you're anywhere near civilization, if push comes to shove: you have at least some possibility of being able to just leave. On a jet airliner in an airport, you are completely trapped.
Federally airlines should be required to deplane passengers after a certain amount of time, or immediately if the plane becomes too hot/cold, runs out of water, or the bathroom stops working....and the flight crew criminally punished if they don't. But that will never happen because of airline industry lobbyists.
The concept of "lint" a file and the concept of verifying a backup are truly ancient and coincidentally are also completely absent from the description of the problem. People that felt no need to do either 20 years ago are certainly not going to start doing it today, especially when the inevitable system failure in the distant future results in yet another lucrative replacement contract in the distant future for system 3.0.
In this case it is two PRODUCTION systems running concurrently. The primary and the secondary (article calls the "backup"). Primary went down due to corruption, but the identical secondary system couldn't be switched to because the corruption also occurred there.
It sounds like their high availability system failed but they probably were able to restore from backup (offline)
>Due to temporary lack of access to Internet and malfunction of the electronic document flow system of Rosaviatsia the Federal Agency for Air Transport is switching to paper version.
>“The document flow procedure is being determined by the current records management instructions.
>“Information exchange will be carried out via AFTN channel (for urgent short message) and postal mail.
>“Please make this information available to all Civil Aviation Organizations.”
Not exactly that we need to praise them or anything, but I think sarcastic tone in some replies ("if only they did something BASIC") aren't needed.
Ideally you would want the combination of both: system doesn't usually crash/fail, but you force it to do so regularly. See https://en.wikipedia.org/wiki/Chaos_engineering
If it had uptime of 2 days but only went down for a minute at a time that would not be a problem
They were doing some mass network upgrade during the early hours maintenance window and devices weren't coming back. I don't think they had OOB access to their boxes and it requires someone going out and recovering them after hours of downtime and getting techs there.
The culprit? They downloaded the OS image using FTP and forgot to set binary mode. Various other network kit vendors I recall would do some level of validation (and anyone downloading should be checking the hash). RCA sent to customers was vague and just said it was a corrupted image / failed upgrades.
I've always been paranoid about this sort of thing when applying BIOS or firmware updates, despite those almost always having checksum validation, but I guess my caution is not unwarranted.
End-to-end and integration tests are much more helpful. But even then, they won't look at operational concerns like backup and recovery.
A fragile thing is like a wine glass. Once it breaks, it cannot be restored to its original state and especially not made better than before.
However, if you're talking about patching the system while it's still running, check out "Stop Writing Dead Programs" from Strangeloop '22: https://youtu.be/8Ab3ArE8W3s
I think a big takeaway from it is that designing systems which are failure-free is a fool's errant - no matter how hard you try, you can never get rid of 100% of the bugs.
Instead, make it failure-tolerant: sooner or later every part of the system will break, so it should be constructed in such a way that it can gracefully recover from failures, and even operate with some parts of it unavailable. Crashes are expected, so the system is designed to handle them properly.
For software, that would mean practices like Chaos Monkey. If your production system stays up while an external process is constantly killing processes and deliberately corrupting memory and files then you have good confidence of riding through unexpected failures.
You don't get what you don't pay for™
https://www.faa.gov/air_traffic/technology/swim/users_forum/...
https://www.faa.gov/sites/faa.gov/files/2021-10/NOTAM%20Mod%...
So probably Linux, though perhaps something else.
Once tuned for processing they just ran - rip out a network cable and stuff it back in again, not a hiccup.
Many of the time saw no benefit in porting forward to Slow Loris.
There are two fundamental philosophies in fault tolerant systems. One is designing fault-tolerant hardware and running non-fault-tolerant software on it. This is what mainframes do. Practically any component of a mainframe can be hotswapped without shutting down the OS.
The other is designing fault-tolerant software and running it on non-fault-tolerant ("commodity") hardware. The latter is so popular that it's pretty much the default now, but it's not the only way of doing things.
How would an IBM mainframe help you with a corrupted database file? I understand that reliable hardware makes the corruption less likely to happen for hardware reasons, but it can also be the result of a software bug, or some unexpected and not correctly checked input.
I think it is a SYSTEM
> It has a backup, which officials switched to when problems with the main system emerged, according to the source.
> Officials ultimately found a corrupt file in the main NOTAM system, the source told CNN. A corrupt file was also found in the backup system.
Perhaps the backup system is just using data from main system which currently is compatible but won't be in the near future?
Still there should be data copy somewhere, right (with corrupt data...)?
I am starting to wonder if this "backup" is an online log replica of the production system.
Failover doesn't work out in the situation where you are replicating trash. A "reboot" from an actual backup/snapshot would be required if you ate a bad log stream.
How often is the system rebooted? I haven't heard of it happening before and some quick searching didn't find any historical examples. Is it a scheduled event and no flights have departure times while this maintenance is taking place?
I've found that issues tend to manifest on production systems with infrequent restarts.
I've done testing on these types of systems in the past (carefully) and the owners will often let you test against the system that's currently the "backup".
So they probably restarted the "primary" after performing a fail-over.
https://www.newyorker.com/magazine/2018/12/10/the-friendship...
File this under "early Google's infrastructure was a low grade cosmic ray detector."
Source: 10 years airbourne geophysics, radiometric calibrations.
Addendum:
In-flight upset 154 km west of Learmonth, WA 7 October 2008 VH-QPA Airbus A330-303 [1] was a probable (but uncertain) example of cosmic ray events causing multiple spikes in one of three air data inertial reference units (ADIRUs) that also went on to cause a failure mode of the "best of three" reporting system leading to a pitch down in which [2]
> 110 of the 303 passengers and nine of the 12 crew members were injured; 12 of the occupants were seriously injured and another 39 received hospital medical treatment.
HOWEVER .. despite 250 events per second in a 42 litre volume, it took 128 million hours of unit operation to see a failure mode.
It's a lot of billiard balls going through a lot of space and a high bar for "something bad" ( just the right bit flip ) to happen.
[1] https://www.atsb.gov.au/sites/default/files/media/3532398/ao...
[2] https://www.atsb.gov.au/publications/investigation_reports/2...
( Final Report TAB )
They are not rare, just because you don't detect them (consumer hardware) means nothing.
>>study done by IBM in the 1990’s that referenced 1 cosmic ray bit flip per 256MB of memory per month
https://blog.mozilla.org/data/2022/04/13/this-week-in-glean-...
> Cosmic ray flux depends on altitude. Computers operated on top of mountains experience an order of magnitude higher rate of soft errors compared to sea level. The rate of upsets in aircraft may be more than 300 times the sea level upset rate.
https://en.wikipedia.org/wiki/Soft_error#Cosmic_rays_creatin...
Yes thats why you have max altitude for normal servers/hardware and it's also more disastrous on modern HW -> smaller transistors.
I lost the thread on that story before the investigation concluded the actual cause was autopilot error. It is, in some ways, more comforting to me to know that the issue wasn't novel atmospheric phenomena, but instead relatively-mundane cosmic radiation flipping one packet of data from the sensors to the avionics that the avionics lacked sufficient redundancy to detect or discard. As a result, the autopilot believed the plane had suddenly pitched 90 degrees and drastically corrected to escape stall.
Somebody inadvertently put a poop emoji in a text field.
OT, but didn't NOTAM use to mean Notice to Airmen? It's still that on the ICAO website: https://www.icao.int/safety/istars/pages/notams.aspx
Did it change, or is this a CNN initiative? (And how does one notify a mission?)
Given that the US often does these kinds of things first, I've often wondered if it is a US-centric way of thinking?
"An engineer 'replaced one file with another,' the official said, not realizing the mistake was being made Tuesday.
Imagine something like this happening at the NYSE, CME, et. at. Or, simply think about the last time you heard about a nationwide credit/debit card outage...
Why can't we have our national infrastructure systems running at least as reliably as the Amex network?
These systems are all information clearinghouses at the end of the day. If we have matching engines that flawlessly process millions of trades per second every day and mainframes that provide resilient source of truth, I think we could consider the same for a life safety critical system as well.
versus: how many people would get fired if NYSE went down due to "we didn't think we needed backups".
incentives matter. In political systems the main goal is to be able to point the finger at somebody else, not to ensure things run well.
Last year in Canada when Rogers had a meltdown.
E.g. https://dailyhive.com/vancouver/everything-impacted-rogers-o...
Bell goes down more rarely but also can take large swathes of nation out.
Maybe it was a corrupted HACKED filed, I'd believe that. After all, China isn't too happy with our games about Taiwan.
I have a feeling the summarized age of the hard drives in that system exceeds the age of the United States itself (which is 247 years). When fsck was introduced in 4BSD in 1980, it checked every filesystem on boot because it had no better idea. If this thing is, say, forty years old then that's exactly the right age for this...
Obviously it was due to details being entered using the now standard International foot when the software expected the now obsolete US survey foot [1] .. the discrepency across a transcontinental flight was large enough for a pilot to fall through.
[1] https://www.nist.gov/news-events/news/2023/01/new-years-eve-...
No. NOTAM is "notice to airmen."
https://www.faa.gov/documentLibrary/media/Order/7930.2S_Chg_...
"These changes include modifying the acronym NOTAM from Notice to Airmen to the more applicable term Notice to Air Missions"
This is not explanatory, and sounds like more pandering.
Now you did.
> doubt that there's any other explanation than pandering.
The airplane doesn't care about the gender or sex of the pilot. There is no pandering in that observation.
90 minutes; this certainly appears to be an advertisement about Windows OS.
Has anyone seen exactly how many NOTAM messages are generated per day, and how long of a look-back is required?
From 50k feet, it looks like something that could be replaced with a cryptographically-authenticated massively-replicated virtual data structure that's oblivious about, and robust against all sorts of failure in, the exact systems, languages, update-paths, etc used to keep it in sync or implement any one user's view.
All that pilots need to know is: "I have a full local copy, signed by the right update-authorities, as of roughly-now."
From a glance, NOTAM's uptime this year looks worse than Ethereum, but a bit better than Solana or Binance Smart Chain.
You've solved one problem and created several more problems.
Especially for a simple log-like system, with a limited number of permissioned authorities.
When I did my initial pilot training in the late 1980’s, the codes for METAR, TAF, and NOTAM had already long been in place. It was explained to me at the time that the encoding was a practice that dated back to its origins in the teletype era. I suppose the limited baud rate of these devices meant that economizing on symbol density was a good idea. I’d still much rather read these succinct formats because it’s easier to chunk it at a glance.
from there you can get to
and pilots won't stop laughing and crying when they read
https://fixingnotams.org/wp-content/uploads/2019/11/Field-Gu...
About 1.5 million per year.
So we're talking, (1.5M * 80 bytes=) 120MB a year, uncompressed. (With the tiny controlled-vocabulary, maybe 6x compression possible?)
Every pilot could have a RAM-resident local queryable copy in a their commodity handheld devices.
There is significant debate concerning the number (too many) and size (too long) of NOTAMs, but I can assure you that when you get an unexpected rerouting on a dark, bumpy, busy night, you do not want to be reading pages on plain-text prose! All experienced pilots are comfortable with the cryptic NOTAMs, and are used to scanning the abbreviations for important items :)