The software that flies SpaceX rockets and starships
stackoverflow.blog
stackoverflow.blog
Exploring the software that flies SpaceX rockets and starships - https://news.ycombinator.com/item?id=27115372 - May 2021 (112 comments)
Also related:
SpaceX Software Team AMA - https://news.ycombinator.com/item?id=23444738 - June 2020 (6 comments)
We are the SpaceX software team, ask us anything - https://news.ycombinator.com/item?id=23432996 - June 2020 (28 comments)
Software Engineering Within SpaceX - https://news.ycombinator.com/item?id=23403800 - June 2020 (313 comments)
We Are SpaceX Software Engineers - https://news.ycombinator.com/item?id=13705762 - Feb 2017 (6 comments)
My personal view is that more of the industry, including SpaceX, should be moving toward engineering things to be more correct-by-construction, including formal modeling & verification as you mention.
As one example, I don't hear most aerospace firms talking about SeL4 for their flight computer RTOSes. For another, you might be surprised (frightened?) by how many vehicles implement their flight software as big single address space executables written in C (sometimes C++) despite running on modern superscalar processors with MMUs. Even static analysis in some of these codebases is relatively recent - they are heavily dependent on functional testing for safety.
I thought they mainly focussed on process isolation and security. A flight computer would be much more concerned about deterministic performance and real-time guarantees.
Also: https://sel4.systems/news/2021
Do you know more that would substantiate defunct-ness?
- Your comment also reinforces my point: if aerospace seeks to be more rigorous, then more in aerospace should be figuring out how to shore up development of SeL4. Instead, they continue with no vision as to what should one day replace things like VxWorks.
I’m not familiar with the standards you called out (not an aerospace engineer), but I would imagine it’s similar to what we do in the controls engineering and automation space.
I’m actually really curious about what software/methods they use to contextualize assets and what those data points are. How do They digitize the equipment data that comes from visual inspections/manual testing.
I wish it was a little more detailed in general.
Obviously they are doing some amazing, and quite high quality work, so I am quite interested in what their internal setup is like.
It’s not like regulation and quality have not fallen by the wayside or been too relaxed for other aerospace companies, especially in recent history. Quality should be meaningful and not just a rubber stamp.
It’s a fine line
I wonder how developer/engineers are expected to split their focus between design and development of their own areas of responsibilities, and doing peer review and testing of other people's areas of responsibilities. One of the dangers of this setup is always that people (and their managers) prioritize output of their own areas, resulting in neglecting testing and review of other areas. It's certainly possible to make the setup, but it takes good people (and team cohesion/morale) up and down the chain to make this work.
I definitely agree that it's a fine line.
Context switching is hard especially when understaffed and most certainly when doing different types of work. I’m still “executing” engineering projects, but I’m also performing inside sales functions, helping advance and maintain our dev systems, and now manage 2 functional groups. Unfortunately stuff does fall by the way side all the time.
Prioritizing things is key, but at the end of the day there is a capacity to what one can do. It takes good management skills to realize the capacity of the employees and when things need to change. Luckily for me I’m working with groups in all three of the areas I was working in to offload those extra responsibilities to other people so I can focus on what I’m supposed to be doing (managing).
It would be interesting to see how they operate.
Specifically, "How does my company deal with failure?"
Which is another way of saying "Are people disincentivized from telling the truth at my company?"
You can never have a process that requires truth-to-power as a regular occurrence that's successful in an environment where there are career penalties to speaking uncomfortable truths.
That's why legacy providers & government have engineered a system that accepts +50% wasted time, in exchange for not requiring truth because almost everything is checked and verified.
Good culture: We failed. I succeeded.
Bad culture: You failed. I succeeded.
I’ve been wondering about this for a year or so because I can’t seem to find one that matches my description, other than concept drawings on NIST 4D/RCS paper, but there got to be one in the world.
https://nasa.github.io/fprime/Architecture/FPrimeArchitectur...
other rockets use rad hardened computers that make decisions on their own, but for cost and time, I believe SpaceX just triplicated all necessary systems and votes out and reboots any system that gets bit flips.
It's mostly LabView I think, which is an interesting choice given all the hate it gets.
For those that don't know, LabView is a graphical (drag and drop) language for the most part where you interconnect various hardware components together. The advantage is that a lot of labs need to plug a lot of devices and sensors to run an experiment and LabView has support for a large amount of hardware. So you connect it all and can graphically view what the outputs are doing in chart form. The downside is that I don't think it's always very easy to specify exactly what you want to do, and code can pretty easily become spaghetti. There are a lot of horror stories about engineers having to maintain someone else's code.
Edit: -Flight software is C++ -Launch is mostly LabView and yes, they have LabView running in mission control
Also it's very easy to build maintainable code in LabVIEW (contrary to the popular narrative), but there's a serious lack of good learning resources, and most people don't even take Core 1 and 2 before trying to code. Kind of like how a person on the business end of things may discover VB.NET and create some horrifying macros, simply because they don't know any better.
I used LabVIEW to code the control / data logging of a prototype racecar and the hardware only broke once when we probably zapped it with a nasty ESD or reverse polarity connection somewhere. We never figured it out and didn't ask when we sent the part to be repaired.
Programming in LabVIEW is something I wouldn't wish on my worst enemy though. It's extremely hard to debug and time consuming to refactor. I always felt that I was 10-50x slower than when I would code in a standard text-based programming language. Working collaboratively also put you in a world of pain because, at least when I used it and I think it's still the case, LabVIEW doesn't support git natively so you have to use an external tool. They have a tool to merge / diff their proprietary binary files but the UX is terrible and I think even if they really tried to do something good, diffing diagrams (which is what a LabVIEW program is IMO) is going to stay hard.
From my experience the people that code in LabVIEW are not software engineer and I had discussions where I was arguing that yes the channel feature is broken but that's not a reason to put all the variables used to communicate between components in a single global variable file for example.
But once you fought against the language long enough and have something that works you can be pretty sure it will keep working forever.
Re: merging; yes absolutely, this is totally broken. NI should either release the file format for VIs, or offer conversion to something saner like XML, which you could easily build a nice git-diff tool for.
I maintain that LabVIEW is a good language with a terrible teaching problem; you can do so much if you just master the queued message handler architecture, and you can do things in a way that are extendible and maintainable for years afterwards
Take the opposite, a little quad copter, very low mass, will need very high frequencies. I used to manually fly miniature remote control helicopters before all of the automatic control stuff came in, they were extremely, inherently unstable, very tricky to control.
Not really. Only when scaling mass, margin of error, and forgetting thrust/gravity.
And even if you scaled mass, thrust, and margin of error proportionally, you'd still have an object exactly as stable as before (force X applied to the small object gives you the same acceleration as proportionally scaled force Xs applied to a larger object). Gravity accelerates all things the same, and so do proportionally stronger thrusters.
But being off by one meter can destroy a small drone just as well as a large spacecraft, so you don't have the luxury of scaling your margin of error by significant amounts.
The only reason larger objects appear more stable to humans is because we, and the environment[1], tend to apply relatively weaker forces to them, and also you don't tend to notice they're already moving at 5m/s because of their size.
[1]Larger objects will better compensate for "truly random" forces, if their size causes them to encounter more of them, canceling them out and/or they're generally small.
You can still have a small but massive flying device and would end up with the same acceleration and jerk because need more thrust to zero the gravity out.
The bigger your thing is the relatively smaller the 9.8 m/s^2 is.
Are there any processes in combustion dynamics faster than 20ms? You bet there are, and when the engine is being tested they are instrumented with high speed probes, because during development/testing you want to _understand_ what’s going. When you are flying you want to stay in control first and foremost.
Of course, it would be awesome to have everything on high speed data acquisition channels. But there is a very limited telemetry budget and everyone is fighting for it. So compromises have to be made.
There probably are high speed probes on the rocket, such as strain gages. And if an engineer thinks that having a high speed engine metric would be _critical_ to do postmortems then it will most likely be included.
So, I would say that assigning telemetry budgets it’s a holistic process where multiple teams fight/cooperate to share that budget and prioritize absolutely critical data channels.
For a SpaceX mission, there isn't a lot that can go wrong in 2 hundredths of a second that you could have magically resolved if you saw it 1 hundredth of a second earlier (i.e. at a 100hz rate).
I was thinking whether this is OK for dragon in-flight abort system. Well, probably is. Or that may use some system on Dragon itself as the rocket may be no more...
Which is why the abort system probably doesn't need to wait for that tick.
They probably use higher rates for stuff like measuring combustion dynamics during ground tests of the engines in development. If anything is wrong there, they'll probably seek to correct it via mechanical or fluid dynamics adjustments of the system itself, since active control systems could probably never react fast enough anyways.
Can anyone experienced comment if 10Hz is very different or the same as vehicle control systems such as a modern car?
Also, are these systems fixed or floating point? Fixed time intervals would help keep numerical instability of floats in check (however I'd really hope this is more formally understood for such a safety-critical system.)
Nowadays I work with satellite SW. Most of the control loops are pretty slow. The fastest ones are those controlling gyros and reaction wheels that run in 5 or 10Hz
I really love my role of spanning the gulf between these worlds. There’s so much to learn from each side to the other.
I was also reminded of James Gosling’s VJUG talk about the Liquid Robotics onboard software for autonomous ocean vehicles:
- Quadrotors control their attitude solely by adjusting the thrust generated by the four motors. Unlike rockets, or fixed-wing aircraft, or cars even, the moment there is any kind of disturbance (whether software, hardware, or environmental), the aircraft immediately starts to fall out of the sky.
- For the work I do, we typically fly around 5-7m above ground. If the motors completely quit, we have around 1.2 sec before we hit the ground at 42km/h.
- The aircraft is also capable of roll rates >50deg/sec. Even if you're high enough that you don't immediately hit the ground, motor problems can also result in you pointing upside down and pulling yourself towards the Earth instead of pushing away.
- Most of the parts have notoriously bad repeatability, plus small variations would wreck any repeatability anyway. You can't just set all of your motor throttles to 60% and expect them to have the same thrust, it's nested-upon-nested control loops to keep everything going
- The sensors are generally crap too with non-trivial amounts of bias and drift.
And the wildest part of all of this? You can get a top notch flight controller to handle all of this for you (and open source!) for $150 USD https://shop.holybro.com/pixhawk-5x_p1279.html?
No wonder John Carmack got into rockets!
From my own background knowledge I remember that the rockets actually run on Intel chips and the Dragon touchscreens are Chromium-based (less certain about that one).
I was hoping for a lot more details like that.
> ...when you plan to reuse booster rockets and shuttles on future missions.
Shuttles? I'm aware of the general term shuttle as a vehicle for going back and forth, but here I'm certain that the better term would be craft, capsule, or Dragon. Certainly not shuttle.You mean the shuttle has been replaced by something newer and hence is an obsolete word?
Doubly so when that is not a term that SpaceX uses to refer to any component or system.
People tend to solve problems with the languages they know and can find people to hire who know it. C++ is fast, and there's a good number of people to hire who know it.
Actual comments I've heard from senior developers when I tell them I write Ada in my spare time:
- "That language is still around?"
- "Other than you, I've never heard or anyone who has ever used Ada."
i must also admit, ada was hard. you had to do some mojo just to print out a value ("package int_io is new integer_io blah, blah). it had a standard of a couple of hundred pages that you kept on your desk. and you looked at it on a regular basis. and the compiler was picky (strongly typed).
companies dropped ada because they couldn't hire. C++ people were more available.
my opinion today is ada is relatively small, makes very fast code, and supports all the programming paradigms of C++ in a very sane way.
if the first ada compiler was "ada95" instead of "ada83", it would be very popular. it came out a decade before the world was ready.
> my opinion today is ada is relatively small, makes very fast code, and supports all the programming paradigms of C++ in a very sane way.
This is why I've been sticking with it. I also get pre/post conditions, a real module system (yes, I know C++20), and bounds checked arrays/strings. Easy built-in multitasking has me using concurrency a lot more often as well.
> companies dropped ada because they couldn't hire. C++ people were more available.
I work professionally in C++, and honestly think you could convert a C++ programmer into an Ada one in about a month or so, due to the conceptual similarities.
I'm confused, I knew a woman in the early 2000s that worked out of Fort Dix in New Jersey, doing ADA development for U.S. Army tanks. She told me that ADA was the standard language for military hardware. Can anyone confirm or deny this?
I don't have the whole story as to when and how this stopped being a requirement (1990s?), or if it was ever a true requirement, but in U.S. DoD aerospace today, Ada is uncommon. I'm not sure of any new systems developed over the past 20 years that are based on Ada. Perhaps some of the avionics suppliers like Honeywell or Rockwell Collins that historically used Ada still do, but I don't think most aerospace prime contractors do today.
The prevalent U.S. Government view is that the industrial entities that implement systems should have the latitude to decide what they want to use. For DoD systems, they would just be expected to have the DoD conduct reviews of their software/system design. [Sadly, such reviews vary wildly in their quality and how much companies' feet are held to any fire.]
Keep in mind the FAA requirements for civil aircraft certification are different from DoD's airworthiness assessments. And then launch vehicles and spacecraft are assessed/certified on a different basis than aircraft.
To give a reference with other languages:
- This is a bit less than C# (whose standard language spec ends at page 462 in ECMA-364 2nd ed.).
- The python spec for 3.9.9 is 151 pages. Okay, but let's compare equivalent things in the C++ spec and Python spec: awaiting a coroutine.
* In Python (the whole "await" spec):
Suspend the execution of coroutine on an awaitable object. Can only be used inside a coroutine function.
await_expr ::= "await" primary
New in version 3.5.
* In C++ (5% of the "co_await" spec): The co_await expression is used to suspend evaluation of a coroutine (9.5.4) while awaiting completion of the computation represented by the operand expression.
await-expression :
co_await cast-expression
But, for C++ this is only the first sentence (which is not written in an abbreviated style but as a proper sentence). Then a whole page explains the semantics in detail, and gives a complete example of how to use it. The python spec is just insufficient in comparison.Most of the bad things about MongoDB are actually bad things about the NoSQL fad, and that fad was huge around here a few years ago.
It's way easier to understand than relational databases, can be simply browsed and retrieved as json (usefull when your only language is JS) and had good adapters for the cookie cutter frameworks used at the time. And you don't need to bother with a schema or joins.
There are use-cases for NoSQL but it's not a once size fits all solution.
> Dragon is a fully autonomous vehicle, so it is capable of completing the trip to and from the ISS without any interaction from the crew. But the displays and button panel do provide the crew with capabilities should they need to take action due an unexpected scenario or emergency. As for the tablets and displays, the tablets themselves act as a sort of backup, and include copies of important data such as procedures. The displays themselves are designed to be fully redundant, so if a single display failed, the other 2 could fully take its place. Even if we were to experience a failure of all 3 displays, the crew has a button panel that can be used to initiate emergency responses, and ground commanding is also possible. > -Jarrett
https://old.reddit.com/r/spacex/comments/ncj4vz/we_are_the_s...