Mars Helicopter Lands Safely After Serious In-Flight Anomaly
spectrum.ieee.org
spectrum.ieee.org
> and while the flight uncovered a timing vulnerability that will now have to be addressed, it also confirmed the robustness of the system in multiple ways.
This is so important. Imagine if it would have crashed and damaged communications, that would have been the end of it. This extra step in engineering effort really payed off and gave Ingenuity a second life, and NASA a really good opportunity to learn important lessons and do further debugging.
Also, the article at mars.nasa.gov is a delight to read, with all the information one would like to know being in there. Not to mention that there is even a video for us to see.
> Flight Six ended with Ingenuity safely on the ground because a number of subsystems – the rotor system, the actuators, and the power system – responded to increased demands to keep the helicopter flying. In a very real sense, Ingenuity muscled through the situation, and while the flight uncovered a timing vulnerability that will now have to be addressed, it also confirmed the robustness of the system in multiple ways.
The code being robust and self-correcting, rather than the increased demands coincidentally being within tolerance levels, would have been more interesting or laudable.
But then again, they've been working with semi-autonomous systems about as long as anybody.
> Then, once the vehicle estimates that the legs are within 1 meter of the ground, the algorithms stop using the navigation camera and altimeter for estimation, relying on the IMU in the same way as on takeoff. As with takeoff, this avoids dust obscuration, but it also serves another purpose -- by relying only on the IMU, we expect to have a very smooth and continuous estimate of our vertical velocity, which is important in order to avoid detecting touchdown prematurely.
https://mars.nasa.gov/technology/helicopter/status/298/what-...
I think it's interesting and laudable that the hardware coped. But yeah, on the face of it this sounds very much like brittle software. For 46 seconds the craft's performance suffered from a single frame being dropped. It sounds like either the calculations were smeared over way too much time, or (more likely) the code made too many assumptions that relied on every piece of the pipeline properly behaving, not only in the moment but also over the entire history of the flight.
I'm speaking from extreme ignorance obviously. But this reminds me of a million code reviews I've done where I've asked developers to make fewer assumptions about the state of the system. Often the response is something like "but how could that ever happen?" And my response is always "i have no idea, but shit happens."
I would love to see a postmortem that discussed the specifics of what went wrong in the software, and whether they can attribute the lack of robustness to system design flaws.
It was not a single lost frame which caused this issue. The issue was that the glitch then corrupted the timestamp accuracy of the frames that followed. I guess the dropped frame was just a symptom, maybe an initial memory corruption or something like that.
But it was not the fact that the frame dropped and that this missing piece of information then had the negative effects.
This is the good thing about the overall outcome of this incident: They now have a chance to learn from this.
It would be interesting to know from NASA if, had Ingenuity crashed and lost all COMMS, they would still know what was the cause of the issue. I mean the wrong timestamps of the images, not what caused them.
It would also be interesting to know what caused them, if it was a bug in the software, or a particle flipping a bit in the timestamping-code.
I understnad the missing video frame is not just physical build issue. This is also a video of video pipeline and speed estimation implementation.
I imagine it is sort of like the organisms early in evolution that simply believed their sensory system when it had in fact been hijacked by something that wanted to eat them. There aren't many sensory systems with those issues these days because anything that doesn't have an internal model that it can use for comparison winds up being a snack or crashing on an alien planet.
From the minimal descriptions I've seen so far, it seems like they were using producer/consumer work queues for the image processing without further synchronization.
Now, I get that this is probably not a real-time system.
A real-time system can still have inputs that can not be guaranteed, and internal processes can have an upper and lower bound for their execution time. So a result would be late if either:
1) $T_{input} + t_{process_min} > T_{deadline}$
2) $T_{input} + t_{process_actual} > T_{deadline}$
(where $T$ is a timestamp given by some time source and $t$ is process duration). Condition 1 can be determined on arrival of the data, and the input can be discarded immediately. Condition 2 cannot be determined beforehand, but if the processing isn't finished at $T_{deadline}$, the process/thread could be killed without waiting for it to complete.Of course, this requires that an accurate deadline can be determined for each input packet. The textbook use case is for rendering live video streams, where stream latency is more important than rendering each frame accurately. This flight control system is a similar use case, since the utility of the camera feed for determining location or drift rapidly declines as the picture ages.
But if the timestamp-keeping isn't accurate as it seems in this case, it really doesn't matter if the system was real-time or not.
I believe the CPU is the same that has been used in a few smartphones.
So it’s not a real-time system.
https://9to5google.com/2021/04/19/nasa-ingenuity-qualcomm-sn...
ref. https://news.ycombinator.com/item?id=26907669 for more details for example.
A RTOS is typically a much slower and much smaller OS. FreeRTOS e.g consists of 3 small source files, and the control loop might be from 1KHz to 10KHz (for extremely high-dynamic systems). Compare that to the snapdragon loop of 2GHz. Factor 1e6.
https in real-time is obviously a joke. There exist proper networking protocols for hard realtime, not randomized and spiky Ethernet based protocols. Telcos use such proper realtime, slicing protocols.
If it really were just producer/consumer workflows one frame would have been lost, the program would be comparing images two frames apart and seen it moving twice as fast as it was supposed to, this would have caused a minor excursion but then the next frame would be right and it would very quickly have settled back down to proper flight.
The problem is that it wasn't simply comparing frames, but it had an internal clock the frames were being compared against.
Drive by question: do you have recommendations on good literature for real time systems?
That in the distance/time code they didn’t take timestamps of each successful frame vs just assuming it’s a 30hz camera so surely each frame will be 30hz!
It seems to me, as someone who does similar embedded work, that if I would account for a missing frame failure, they surely would. I am a professional in embedded timing systems - I am guessing this isn’t the full story or it’s been paraphrased for normal people and something important was lost.
EDIT: Better info below. The photo was lost but it’s timestamp was not. They apparently didn’t tightly couple the image and it’s metadata together (hashes... what are they for!?) so there was some sort of mess up there. Knew there was more to the story.
A struct or pair of {Image, Timestamp} would be the simplest & most robust solution.
Yes, you can make collisions arbitrarily unlikely, and there are so many situations where they come in useful, but I can imagine an engineering culture where they are just never an easy tool to reach for.
IDK, I would call more possibilities than grains of sand in the universe a little more than “arbitrarily unlikely”.
I would imagine you would want to demonstrate the P(collision|domain+operating lifetime) << acceptable level or risk, and after you’d done that, take the uniqueness of hashes as a axiom for the proof system?
That’s just speculation on my part, since I haven’t done much with formal proofs of working code
Now you’d need to hope the inertial guidance is good enough at those altitudes. The article says it doesn’t use image based calculations for landing due to the possibility of kicking up dust.
It seems that the timestamping is detached from frameshooting. That is, as if there were a burst happening, then the burst's frames get assigned the sequential timestamps. Thus if a burst had a glitched/skipped frame, a part of the series and the subsequent bursts become mis-attributed.
I understand that integrating the timestamping to be atomic to frameshooting is likely to slowdown the burst. So there has to be some expected internal timeline to validate the timeline that results from bursting.
Can you describe this further, or link to further reading?
0. https://scholar.google.com/scholar?q=Nicholas+Strausfeld
1. https://www.mn.uio.no/cees/english/services/van-valen/evolut...
I learned only way too late in the project that most GigeVision cameras emit timestamps that can be synched using NTP and/or PTP. The camera specs/manual said nothing about this. Only found it by accident looking through the GeniCam XML API.
Moral of the story: ensure your cameras emit timestamps on device. Don't rely on system time.
Google/Amazon do have a similar problem with eventually consistent databases, but low power (NB-IoT) meshes are harder, as synced wakeup's are critical. Don't rely on device time.
Looks to be ~$3 on Digikey: https://www.digikey.com/en/products/detail/omnivision-techno...
The video from Ingenuity itself is taken by a black-and-white camera and has no color information.
Edit: appears so: https://astronomynow.com/2021/04/11/nasa-delays-mars-helicop...
Actually, I remember reading somewhere that more and more they use FPGA chips, allowing to even "change out the CPU" if needed.
http://www.esa.int/Enabling_Support/Space_Engineering_Techno...
Lets call it an "OPA" - Other Planet Update
Or perhaps "USTLLTFMMA" (updating software thats literally like thirty fucking million miles away)
https://en.wikipedia.org/wiki/Aether_theories#Quantum_vacuum
Sacrifice redundance, to maintain rollback capability, and flash one bank. Power on, test, and if test fails, restore to previous configuration.
(not sure what distro they're running)
"Tim Canham, the Mars Helicopter Operations Lead, shares Linux’s origins at JPL and how it ended up running on multiple boxes on Mars.
Plus the challenges Linux still faces before its ready for mission-critical space exploration."
edit: just saw someone already mentioned the podcast in this thread
Can I get a tl;dr on this
My favorite is we got details. Legitimate details that give us a pretty clear picture where it went wrong and how it can be corrected. No calling it "software glitch" and stopping there. Those kinds of articles frustrate me to no end.
As a very small portion of the HN crowd, believe we are quite pleased.
The communications from NASA's Mars teams has been incredibly thorough and transparent. That they share all of the raw images soon after they are received from Mars gives an incredible feeling of inclusion.
"For the majority of time airborne, the downward-looking navcams takes 30 pictures a second of the Martian surface and immediately feeds them into the helicopter’s navigation system. Each time an image arrives, the navigation system’s algorithm performs a series of actions: First, it examines the timestamp that it receives together with the image in order to determine when the image was taken. Then, the algorithm makes a prediction about what the camera should have been seeing at that particular point in time, in terms of surface features that it can recognize from previous images taken moments before (typically due to color variations and protuberances like rocks and sand ripples). Finally, the algorithm looks at where those features actually appear in the image. The navigation algorithm uses the difference between the predicted and actual locations of these features to correct its estimates of position, velocity, and attitude.
Approximately 54 seconds into the flight, a glitch occurred in the pipeline of images being delivered by the navigation camera. This glitch caused a single image to be lost, but more importantly, it resulted in all later navigation images being delivered with inaccurate timestamps. From this point on, each time the navigation algorithm performed a correction based on a navigation image, it was operating on the basis of incorrect information about when the image was taken. The resulting inconsistencies significantly degraded the information used to fly the helicopter, leading to estimates being constantly “corrected” to account for phantom errors. Large oscillations ensued."
> Flight Six ended with Ingenuity safely on the ground because a number of subsystems – the rotor system, the actuators, and the power system – responded to increased demands to keep the helicopter flying. In a very real sense, Ingenuity muscled through the situation, and while the flight uncovered a timing vulnerability that will now have to be addressed, it also confirmed the robustness of the system in multiple ways.
Description:
Tim Canham, Mars Helicopter Operations Lead at NASA’s JPL joins us again to share technical details you've never heard about the Ingenuity Linux Copter on Mars. And the challenges they had to work around to achieve their five successful flights.
I assume it was just a failure to record the frame, but it'd be quite something if it was a drop frame from 29.97 video.
54 seconds of 29.97 frames: 1618.38 frames
54 seconds of 30.00 frames: 1620.00 frames
Toss in a bit of rounding errors and/or tolerances and it's definitely suspiciously close to where a 1 frame error appears between those two.I'd suggest that the evidence against it would be that while this may be the longest flight on Mars, it seems certain that they'd have taken longer flights on Earth than this. But still, the math nearly checks out.
I bought a cheap drone on a common importing website and the thing likes to get lost even though it has GPS and is in the perfect clear. It flew into my truck, uncommanded, at full speed. On Earth. With GPS. So for this thing on another planet to do what it just did? Wow.
https://mars.nasa.gov/technology/helicopter/status/305/survi...
OK, so from the NASA article, it seems that they use an inertial measurement unit [0] (think an accelerometer/rotational sensor similiar to what your smartphone uses) that reads out heading and acceleration 500 times per second and sums it up to get the the helicopter's position. This is "dead reckoning" [1]) which unfortunately suffers from accumulating errors. To work around these accumulating errors, they sync their prediction to camera data thirty times per second by predicting how features visible on the surface should have moved in this time. So, sensor fusion [2], in a sense. The images are delivered with timestamps (I would have been surprised if they were not, to be honest) but one of the images went missing and for some reason, its timestamp was not. (Variable not cleared? Are they maybe using parallel queues to match images to timestamps (would be very weird, imo but I am no helicopter scientist?)). Either way, the following timestamps were now attributed to wrong images. I really wonder how that happened and hope they expand on that later. As the system predicts how far a feature in an image should have moved between frames, things were now considerably off (example: from the IMU prediction, a feature should have moved 1 unit per second but with a missing frame and the previous timestamp, it now looks as if it moved 2 units in a second) and the state gets updated with this wrong information.
Here is an interesting bit about safety margins from this article by the way:
Despite encountering this anomaly, Ingenuity was able to maintain flight and land safely [...]. One reason it was able to do so is the considerable effort that has gone into ensuring that the helicopter’s flight control system has ample “stability margin”: We designed Ingenuity to tolerate significant errors without becoming unstable, including errors in timing. This built-in margin was not fully needed in Ingenuity’s previous flights, because the vehicle’s behavior was in-family with our expectations, but this margin came to the rescue in Flight Six.
[0] https://en.wikipedia.org/wiki/Inertial_measurement_unit
Even in absolute terms, once you are 110 seconds into the flight, we are talking about an offset of 1s/(30*110s), or 0.03% — if control algorithm can get confused by an error of this magnitude, it's not a very good one.
But honestly, I am sure the algorithm is much better than that, and I believe the explanation is "watered down" for the general populace, and a lot more is in play that's not being shared.
Unless they box that in with limits you could easily get to a point where it starts flying like there’s a chimp on the yoke and either crash or saturate the navigation system and have it fall back/fail entirely.
They should have just used GPS!
/s
I wonder, why not have a real time clock capture time stamps for different inputs for same frame, like one when written to memory, one when the trigger for the frame is done, etc and have a comparative analysis. Discard the process when there is a discrepancy, and re-start the captures from beginning, and go into a default hover state when such a fault is found.
You have to make some assumptions. For example, that pure functions will always return the same result with the same arguments. In fact, in real life, it is not always true, especially in space where a cosmic ray can flip a bit. You may try do be defensive to account for an unreliable hardware, but you have to draw the line at some point.
Yes, is should have been tested, but it doesn't surprise me that it wasn't even for a company with as good a track record as JPL.
#include <stdio.h>
int f(unsigned n_steps, int from_val, int to_val){ int step_size = (from_val - to_val) / n_steps; return step_size; }
int main(){ printf("f(3, 0, 30) = %d\n", f(3, 0, 30) ); printf("f(3, 30, 0) = %d\n", f(3, 30, 0) ); }
Yeah, they are, but your example is not good enough, it is like an example from a textbook on traps of C.
Corner cases hard when you are trying to do something new, because if you did it before, you'd know the most of them. Or if someone else did it before and wrote a textbook. =)
The android camera API for example will tag every frame with the timestamps and exposure settings it was captured with. Unit and integration tests can simulate dropped frames and verify the algorithm still runs as it should.
I would guess they tried to go for direct camera hardware access and ended up writing their own logic to schedule frames, and it contained a bug.
I'm also reasonably certain that the drone isn't running Android, and that not doing so is the correct design choice.
Considering Qualcomm never releases their kernel sources, I am really worried that NASA might have been forced to use a modified Android system.
Most likely they just used the specific kernel version with closed binary blobs from Qualcomm for a typical embedded Linux, not the full Android system.
I don't think anyone is seriously suggesting the thing should run android.
But if you are writing something yourself then looking at what existing libraries already (and more importantly their test suites!) do is a good idea.
Maybe they did and maybe they didn't in this case, hard to know with the data available.
this isn't magic. Image processing isn't a wholly unknown field when under the influence of Martian gravity.
It's a staggering achievement, and a testament to humankind's ability to wrangle technology, but -- to be frank here -- visual odometry is technology from 1960s era spy planes, not beyond-human voodoo.
Sometimes errors are stupid ones, even if the stupidity is causing problems a zillion miles away on distant unknown frontiers being first-explored by human kind.
tl;dr : This visual odometry problem isn't a space problem, it's a computer science problem -- one that has been encountered by numerous people even here on earth.
The repercussions of such a simple non-space-problem might very well turn into space problems when the craft crashes, however -- and that's a damn shame given that this particular set of problems has been encountered-and-fixed by numerous earthlings over and over again.
super tl;dr : as an only half-informed-participant in this discussion I bet this problem is NIH-syndrome-related, a problem NASA has encountered a lot in the past, a problem that's entirely managerial and very hard to eradicate.
We have plenty of faux-mars environments to practice these techniques on. NASA even has their own.
many of the toy indoor-use drones use such techniques for attitude/altitude-hold features when under a roof.
In the context of a space helicopter, I’d guess that deep understanding and ownership of each software component is super important when you’re likely to have to support a debug it from millions of miles away. At this scale, and in this context, doesn’t “rolling your own” critical components, in some cases, make sense?