How to avoid a BSOD on your 2B dollar spacecraft
clarkwakeland.com
clarkwakeland.com
Re-reading the post, I see how the title, my analogies, and poor attempts at humor would give the incorrect description of what’s happening with the satellite when it enters safemode. I’ll amend the post soon.
Thanks for the feedback, I’ll be better next time.
Also do you not run these tests in an even more simulated environment where there is only the flight computer and no real hardware at all?
We normally have a Functional Test Assembly (real computer and some other hardware for testing) to run our tests against, but we only have one setup and it is consistently unreliable. This particular CLT was unable to get a clean run in the lab but it was decided that the issues were related to the lab setup rather than the actual test, so we moved forward to run on the satellite (against our team's protests).
This to me is the real crux of the issue: if we can't even trust our own testing environment, what's the point of having it at all? If the customer is so risk averse, why would we take this chance? Needless to say, I don't think we'll be running anything on the satellite without full FTA vetting anytime in the near future.
To the last part first: Good that safe mode kicked in and did the right thing, but now what? What caused it to enter safe mode in the first place?
That's why they care when it happens. If they don't know why it's entering safe mode, they can't correct the actual problems in the system.
The non-critical functions are all the things the customer actually bought the satellite for. Cool that it's still alive, but now the Space Internet / death lasers / etc. are offline.
To your point, yes you're correct. The cause of the safemode is much more interesting than the fact we entered it.
Its interesting to see that someone with a 2B budget have the same problem as someome with 5 million budget... we have an engineering model for our cubesats but its flaky
However, I did come away thinking there are other dysfunctions at play in all of this. Perhaps an excessive amount of wheel re-inventing.
I got lost recently in how the Shuttle software was managed, mostly through IBM mainframes, and z/OSs facilities for all the above. I'm curious how modern development looks in comparison.
Do you have any references for this? I also recently went down a research rabbit hole of the history of computing on Earth and in space - super interesting stuff. And the parallels are quite obvious when you look at it.
Oh yea. https://ntrs.nasa.gov/citations/20090001334
> And the parallels are quite obvious when you look at it.
The insane level of detail and strategy when writing the shuttle software is something to behold. The testing laboratory SAIL was a full scale orbiter that actually flew test missions. "Day of use I-Loads" are one of my favorite things. They couldn't change the software load, but they could move some constants around before launch, really useful for feeding wind data into the shuttle before it launched.
I am a spacecraft engineer. I don’t see anything in the linked article indicating that they are actually running Windows - the BSOD claim is tongue-in-cheek, or at least that’s how I read it. I also don’t know of anyone anywhere that runs Windows on a spacecraft, with the exception of laptops used by astronauts. Typically one runs vxWorks, or maybe QNX. Some experimental (high risk, low cost) systems run Linux. Older spacecraft don't run any OS at all, everything is running on bare metal, and that may be true for a handful of current spacecraft as well.
Windows is used in some places by ground controllers, but these days they tend to be running Linux a lot more often.
The license list is a bit long:
https://www.starlink.com/assets/pdfs/Starlink-Open-Source-Co...
In the same way that any pi project should really use i2c or spi to comms with hardware that's more purpose designed.
Software actually running on satellites isn't being distributed, so there's no license obligations there (unless it's AGPL, ha!)
I've also worked with a payload running Windows Embedded.
I'm working on a Linux distribution for space applications that addresses some of the pain points (application deployment, software update & memory-safe implementations of typical space protocols like CCSDS/PUS). If all goes to plan it will fly on a 1U CubeSat tech demonstrator this year, a cybersecurity research 1U CubeSat next year, and a 2U high performance satellite ... later. :)
citation https://old.reddit.com/r/spacex/comments/ncj4vz/we_are_the_s...
why would that be mistaken for a joke?
(Maybe this is a good place to ask, anyone have a recommendation for static testing of JS?)
People also point out stuff like `[]+[]` where Array addition operator is not overloaded to handle it at all because the correct way is to use `.concat` so I think by default duck typing comes into play and in most cases with that in JS the preferred type is string.
I do wish that there was a "cleaner" spec of JS that wasn't as backwards compatible but fixed all of the gotchas and filled in every gap but there doesn't seem to be much call for it atm.
At the time, some very good flight software engineers had been working diligently on a new UI framework that was written in the same code style and process as the rest of our flight software. However, I noticed a classic problem - we were working on the UI platform at the same time that we were trying to design and prototype the actual UI.
I made some observations:
1) We can create a prototype right now in Chrome, with its incumbent versatility.
2) The chip running the UI can actually reasonably run Chrome.
3) Web browsers are historically known for crashing, but that's partly because they have to handle every page on the whole Internet. A static system with the same browser running a single website, heavily tested, may be reliable enough for our needs.
4) We can always go back and reimplement the UI on top of the space-grade UI platform, and actually it'll be a lot easier because we will know exactly functionality we need out of that platform.
The prototype was a great success; we were able to implement a lot of interesting UI in just a week.
I left SpaceX before Crew Dragon launched, so I'm not sure what ended up launching or what the state of affairs is today. I remember hearing some feedback from testing sessions that the astronauts were pleasantly surprised when we were able to live edit a button when they commented it was too hard to reliably press it with their gloved finger.
As for reliability, to do a fair analysis you need to understand the requirements of the mission. Only then can you start thinking about faults and how to mitigate them. This isn't like Apollo where the astronauts had to physically reconfigure the spacecraft for each phase of the mission -- to an exceptionally large extent, Dragon flies itself. As a minor example of systemic fault tolerance, each display is individually controlled by its own processor. If a display fails, whether due to Chrome or cosmic radiation, an astronaut can simply use a different display.
Also, as a side note regarding "touchscreens" -- I believe some (very important) buttons did launch with Crew Dragon, but buttons and wiring are heavy, and weight is the enemy. If you're going to have a screen anyways, making it a touchscreen adds relatively trivial weight.
Updates keep us secure. /s
Please tell me you have a blog
But if it's implemented correctly, it probably won't get out of sync.
I am always amazed how HN doesn't realize many mission & life critical systems are powered by JS - especially as a front end through a browser.
I am always amazed how some people believe that adding more levels of abstraction leads to safer systems. /s
As long as the risks are acceptable (single non mission critical display out of many may rarely crash) then it's fine imo.
Still get a familiar development experience, live updates, and able to run on constrained devices.
https://news.ycombinator.com/item?id=41651715
TL;DR: the spacecraft is indeed not running Windows. It's running a custom OS written in C.
> Safemode is the satellite equivalent of a blue screen of death.
It’s about avoiding safemode, and more generally about the end-to-end QA/testing process for satellites before they’re sent up into orbit. It’s very clearly not about actual Windows BSODs, it’s just written in a tongue-in-cheek style. Those commenting about “wtf windows on a spacecraft” clearly didn’t read the article, just read the title.
FWIW I found the writing style engaging and the content interesting. I guess the title is a little click-bait-y, but not in a way that I minded much, and I probably wouldn’t have read an article titled “How to avoid safemode on a satellite.” It’s a fine line, but titles DO have to draw you in, otherwise you’ll never read the article.
Re: the article itself, I did think it was pretty wild that customers have to be informed of every incident where a satellite flips into safemode in TESTING! In real operations, sure, but in testing, that’s wild. Feels like having to report bugs caught in my local dev environment, that were never deployed to prod.
If you’re paying two billion $ for something you become very very interested in test design and test results.
Also, safe mode isn’t really the same as a BSOD. It’s a mode where the spacecraft decides something is wrong and disables a lot of functionality and focuses on pointing the solar panels at the sun and the antennas at the ground. It does not cease functioning - if that happens, you’ve probably lost your spacecraft. It is therefore VITALLY IMPORTANT that safe mode works, and a smart program manager tests the hell out of it.
We already got that it's not actually Windows and so not literally identical to bsod in every detail.
It's not the same as a common os safe mode either because it happens by itself as the last resort response to a problem, like a bsod. Not just on command.
I have worked on software where individual customers pay millions for it, but not billions, and it’s also not a physical thing that can literally crash into earth or fly off into space if something goes wrong!
Because the customers are almost certainly running their own metrics, tracking failure rates over time. An increasing rate of failures across a program is probably a sign of something going wrong at a higher level. Remember too that there is "testing" and testing. One is you playing around with the software at your workstation, the other is the more formal testing as monitored by the acceptance and standards people.
Because TFA is highly misleading
> I don’t see anything in the linked article indicating that they are actually running Windows
TFA literally begins with a picture of a Windows BSOD with a Windows error message.
Then it got bought by RIM (BlackBerry). The source code went away and the owners stopped responding to the user community and the whole thing basically became irrelevant.
https://www.eng.auburn.edu/~kchang/comp6710/readings/They%20...
Take your satellite, replace it's navigation/communication inputs with ones generated by a reasonably high fidelity real time physical simulation. Feed it's outputs back into said simulation. Ensure the satellite does the right things at the right time.
That's... Concerning. No root cause analysis? Not even an internal one?
> the US government isn’t burning taxpayer dollars on a ten figure spaceship just to have us push a Crowdstrike update on it.
That or inexperienced programmers were involved, assuming they were not scared of modifying memory addresses directly.
As for the safe-mode, if it happened maybe you could say you were randomly injecting errors in the memory during runtime and spacecraft entered safe mode as expected, would not be far off from the truth, just do not mention it was unintended :)
I can tell you a little first-hand account of where this helped a satellite formation flight mission I worked on. The communication system was working fine in terms of signal strength, but many command packets were ignored (no response). We were able to figure out that a message queue in the processing pipeline was considered full, by strategically reading certain memory locations. We then sent a memory patch to the satellite to skip some of the processing steps and this improved the communication system dramatically.
We also brought back to life the first CubeSat launched by my university, BEESAT-1, by using arbitrary memory write to patch some telemetry collection software to avoid a damaged section of the onboard flash. Pretty cool story actually.
For example, you could use an OS that is deterministic down to the last detail, and have a "digital twin" / virtual machine of the spacecraft computer here on Earth, kept in sync with all sensor and actuator activity out in space. Before issuing any command to the spacecraft, you branch the digital twin, issue the command on the branch, and make sure everything looks good.
With this method, you wouldn't even need to read memory locations on the spacecraft, you could just read them on the digital twin. Then test the memory patch on the digital twin and make very sure that it won't brick your spacecraft before you transmit it.
I can imagine that these "state dumps" would be quite big though. And there is definitely some hidden state in the hardware blocks themselves (SPI, I2C, whatever...).
Things happen, and the electronics has to take radiation hits. I'd make the software as defensive as feasible so failures have less potential to cascade.
The problem was that their parameter file had no header for the machine type. Unchecked memcpy in space is seldom good, and you are lucky if it trips a shutdown. Rotating satellites don't tend to unrotate in time, and very soon they are gone.
But what would you run? QNX? BSD?
A better way to put my argument is: could an average mom build Linux on specialized hardware in space? If the answer is "yes", then you may have a point.
I don't think the answer is yes.
Just because you can read it, doesn't mean you can or will be able to actually fix it; not because of technicality, but because of personal knowledge.
In this case, they are both black boxes.
I once spent three days trying to figure out an issue, stepping line by line through hadoop (after figuring out the issue was in hadoop and not my own code). Yay, I proved the issue was actually in Java itself. Guess what happened next? We avoided the bug. Why?
- We couldn't update Java.
- We couldn't change hadoop because we were using a packaged solution. So, we just filed a bug with them.
Had the source not been available, we would have just skipped all of that, and it would have been our vendor's problem 3 days earlier.
This happens less nowadays, because stuff like Android and ChromeOS, so why pay Microsoft when free beer OSes exist.
If you look at an LTS branch, you'll see there are hundreds of point releases. Usually a point release is created once every 5-10 days. I interpret that to mean bugs were not found until many weeks after the LTS branch was cut. Obviously, not all of them affect you, but many patches apply important subsystems which affect you.
That said, this is on the scale that I'm surprised off-the-shelf things are even being considered. I'd have thought they'd just roll something bespoke and that's the end of that.
With Windows, you need to beg and pay Microsoft for customizations and hope these changes will not cause other issues.
Plus, most space projects are on a tight and limited budgets where management would rather spend on hardware than software.
Commodity equipment running specialized distros of Linux is a growing thing however.
Linux RT merge: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
Also RTOS tend to have a lot of verification work applied so that you can be confident some facilities just won't fail.
> If the watchdog timer has not been restarted and instead times out after ~30 seconds, the satellite enters something called safemode. Safemode is when all non critical functions are automatically shut down and the satellite becomes entirely focused on generating power by pointing its solar panels towards the Sun and trying to reestablish any communication that was lost. It’s a state the vehicle goes into when something bad happens [...] the satellite equivalent of a blue screen of death.
If only Windows would be so kind!
Detailed writeup by project manager: https://forum.nasaspaceflight.com/index.php?topic=27043.0 (doesn't mention the OS)
Some indication the laptops on ISS run linux on HP ZBooks: https://training.linuxfoundation.org/solutions/corporate-sol...