Computing glitch may have doomed Mars lander
nature.com
nature.com
>>...The lander even switched on its suite of instruments, ready to record Mars’s weather and electrical field...
It's a shame. But, reading that, I can't help imagining the lander in the role of the Blue Whale from Hitchhiker's Guide to the Galaxy:
"I wonder what this big flat thing rushing toward me is? I think I'll call it 'The Ground'"
I don't think GP was judging ESA. I think they were merely pointing out how impressive NASA's success rate has been lately.
You also have to pre-program the entire flight and find out half an hour later if your parachutes and retro-rockets deployed, or you made a relatively small crater.
It went riding on another orbiter, Mars Express, but at least reached the surface intact.
NASA's budget is larger, but they're doing more missions (and spending a lot on SLS and James Webb). I suspect blaming the ExoMars crash on budget is inaccurate.
https://en.wikipedia.org/wiki/Mars_Polar_Lander#Landing_atte...
I wonder what causes the confusion.
But it's not the case, isn't it?
Maybe they used an US-developed microcontroller library and screwed up the unit setting.
Much simpler to test fire the engines in a lab somewhere to make sure they run, and test the software in a simulation of the landing. It'll probably come out in the next year or so that there was either just a flat error in the code, or that the simulation was incorrect in some way.
If I were ever to somehow become absolute tyrant over the entire Earth, my first edict would make open and notorious use of non-SI units a capital crime, then I'd go live in my imperial hermit shack on a remote mountainside (paradoxically, one with a >1 Gb/s network connection) until my minions finished mopping up the last of the blood.
I think the first way should be fast enough with today's processors (but IIRC the processors on the Rovers are slower).
That said, _if_ there were any US companies involved then perhaps there might have been complications. Many years ago I worked on guidance software for tunnelling machines and we had a US client who insisted that we reprogrammed the entire system for them to work in "centifeet", i.e. 100th of a foot or approx 3mm.
I well remember the chaos and arguments caused by adopting such a ridiculous chimera of a unit, and was struck by how resistant some people are to adopting something as universal and transferrable as SI units.
SI is an improvement but it is not quite perfect. As long as we have a choice of imperfect systems, we have to pick which is most convenient.
I somewhat agree with that - I had the exact same conversation with my kids very recently.
Re. physical object: well, whatever physical reference artefact is used, to all intents and purposes the physical stuff is water, with the basically sensible equivalence 1L = 1kg = 1000cc. As you point out, it's not perfect because of the somewhat arbitrary scales (kg, cc), and it gets worse when you bring temperature into it (I believe it's supposed to be measured at 4 degrees C, but I'm not totally sure of that).
You're right in that it is not perfect, but at least it is a fair attempt at coherence.
Correct; it's because at that temperature water is at its densest: https://en.wikipedia.org/wiki/Water#States
Redefining the Kilogram with the DIY Watt Balance: https://www.youtube.com/watch?v=ewQkE8t0xgQ
But what kind of "on the ground" test condition could have made it to do that ?
The fact that it was "designed to decelerate the craft for 30 seconds until it was metres off the ground" but instead it ran only 3, raises many questions. Not to say that there are zillions of ways to actually test that you are on the ground or not..
I don't know about that. I've worked on systems where there is some hardware with thruster duration that software is writing. So if software crashes, the thrusters stop. (But I have never worked on a lander). If software is running the control loop, things would get very bad if the engines kept running at their previous output levels for 27 seconds. Thrusters are never aligned perfectly, who knows what kind of rate the lander had, there may be wind pushing it, etc. If a thruster blindly kept running, the lander may have spun up, flipped over, etc.
But the point that I wanted to make is that a software crash wouldn't explain why the lander behaved as if it was on the ground. So it's more probable that it reached the "on the ground" condition.
If that condition was adequately programmed, the chances that some sensor glitch could have triggered it are quite low. So the probability is higher that it was either badly programmed or maybe someone flipped a bit somewhere ..
Given that the parachute was deployed at the correct time and its deployment was governed by air speed (to be deployed around 1mach), the gyroscopes were likely working correctly. The radar detection wasn't available until after deployment of the parachute (more specifically, after ejecting the heat shield, which occurred at the same time), so if the machine's decision was based on faulty input, the doppler radar is likely to blame.
It is also possible that the sensors were working correctly, but the inputs were processed incorrectly. I'd expect that an issue like that could have been caught while testing the lander on earth...
edit: it's also unlikely the engines would have kept firing after computer malfunction: there were multiple (9 iirc) thrusters all around the craft and they were fired independently using pulse-width modulation. This approach was intended to both slow down the descent and to stabilize its course/orientation.
Unless they used a ground proximity radar which sent back false pings due to overheating caused by the too early discarded heat shield?
I'll wait for the final crash analysis and report, but I trust that the lessons learned here will significantly improve the chances of future missions, if at least by adding some more redundancies into the mix.
Somehow people seem to assume software is easier, but then there are at least as many critical failures blamed on software as on hardware.
Maybe there are different curves: software is easy to get to 90% working but hard to get from there to 99.9%. Hardware is hard to get to 90% working but easy to get from there to 99.9%.
So the question reminds me of someone asking, "What's harder, physics or philosophy?"
> led the craft to believe it was...
Don't anthropomorphize computers, they hate that.
As an armchair space explorer, I would have imagined such a crucial part of the mission had serious redundancy.
Redundancy adds either weight if you're computing things on two systems at once, or time if you're recomputing using different code. I'd guess that working out the balance between risk of failure, weight and time is a very complicated problem.
It could be something straightforward that presented an unexpected input and made it behave in an equally unexpected fashion. I'd imagine they're currently replaying the sensor data to an instance on earth and seeing how it responds.
The anthropomorphized version is more concise for someone familiar with and more understandable for someone unfamiliar with the inner details of computer systems.
I actually don't think anthropomorphization is helpful in discussions with the lay public but I'm not that serious about it.
Of course, this is a bit different for software we ship to other planets :P
The root cause is that hardware has clear purpose and is limited by the real world.
An aircraft from 2015 or from 2000 does the same thing: Get you from Point A to Point B, same world, same constraints, same goal.
A software that does something today. It doesn't give any clue about what it will be or do in 2-5-10 years.
IMO: Software projects are way more ambitious, more complex and evolving faster than hardware. What is expected from software is NOT expected (nor achievable) in hardware at all, that's one of the first thing I noticed.
---
Nonetheless, hardware can run software onboard. So you get hardware and software dev projects, both running along. That get some quirks of both world. IMO, the hardware [and real world] limitations will quickly limit the software.
That's hard to say in general. With current practices, I'd say software is harder to get right, but that is mainly due to a failure to spec correctly. Hardware is engineered with very tight controls: dimensions, shapes, and properties are all produced according to a specification that doesn't just specify its nominal values, but also the acceptable tolerance for any deviations.
No software that I know of is engineered to the same level of precision as hardware. The lack of specification-guided tooling (and validation for 3rd-party components) is what puts software development not even in the same league as hardware development.
Are you counting NASA-produced software? I'm not familiar with ESA's software pipeline but I imagine it's similar.
[1] http://digitalassets.lib.berkeley.edu/techreports/ucb/text/E...
"Hardware eventually fails. Software eventually works." -- Michael Hartung
Perhaps the hardware altimeter was sending errors or the gyroscope was out ?
The comet lander Philae bounced and was lost because the hardware harpoons failed to secure it.
Neither. Rather, it's the total lack of design accountability procedures in place for software. When a mechanical or electrical engineer messes up, there is a legal framework in place to hold the involved parties responsible for their actions. Depending on how bad the negligence was, this can include jail time. This motivates those designing the hardware aspect of a product to check and recheck their work many times over, and as a result you rarely read about hardware failures with this level of frequency.
In contrast, when a software system fails, everyone seems to just throw their hands up in the air and say how this is par for the course. There is currently no legal PE certification available for software, so we have seen disproportionately many catastrophic software failures.
Because of this, my guess is that the camera technology that we can safely get to Mars is significantly worse than what we expect in our day-to-day lives.
These projects started out 10-15 years ago. So upgrading them mid-stream is almost impossible.
They're stuck with the hardware available when they chose that part. Re-designing for new hardware means changing everything.
But as I'm developing software for my Universitie's cubesat, it's incredibly interesting to read all that goes into Single Event Upsets and other hardware failures from being outside of Earth's protective atmosphere.
For silicon devices, the VERY hard rad environment of inter-planetary space is pretty much instadeath. So you are forced to be 'dumb'. You can shield it, but that adds weight and heat issues. Redundancy is out, it just doubles/triples the weight, heat, and interfacing issues. You can use rad-hardened devices (best idea) but then you have to fab that silicon anew, and this is really expensive. Then you have to test all these things before you put them up there. So you need to put your device in a vacuum-rad chamber. Governments are not so hot on just having these rad sources laying about and letting a piece of silicon 'cook' next to a power station that is in use. So you do accelerated testing, aka, super radiation. I've been in those chambers before, it's scary. After the magic glow machines are turned off, you have to evac all the air as the O2 in the room has been turned into O3 and you can't breathe in there. The roaches in the corners are literally cooked and brittle. It's nuts. Unfortunately for your testing, this environment is more like the sun's corona than the Earth-Mars transit, so you have to take the testing with a grain of salt. Also your vac chamber is not nearly as good as real space. So everything is oxidized and hat messes up your analysis. You do the best you can, but we don't get many sample returns from outside the Earth's magnetic field to test against. NASA et al do a GREAT job, but they are dealing with very extreme jobs and we just don't have a lot of data about what we are dealing with. We can't do a lot of in situ tests.
http://blog.lenovo.com/en/blog/thinkpad-laptop-nasa-youtube-...
As pointed out elsewhere in this thread, the consumer grade devices cannot withstand the extremes of space (not just temp extremes, also high energy rays/particles we're shielded from here).
This [0] is an interesting read on some of the considerations for the Curiosity rover cameras.
[0] https://www.dpreview.com/articles/0353350380/curiosity-inter...
(I realise it's not your main point, but broadcasting 60fps video across the solar system sounds like it might need a LOT of power, no matter of sensitive the receiver is.)
- radiation: electrons, protons, heavy ions (MOSFETs are very vulnerable)
- cold temperatures (up to -235 degree celsius)
- high temperatures (up to +250 celsius, and there's no air-cooling in space)
The performance of electronics vary greatly depending on radiation amount and type, and temperature. You can absolutely not throw consumer electronics (GoPro) in deep space and expect it to function.
And since the fabs that manufacture this kind of stuff are working at a relatively large technology node, it's non-trivial to design something like an i5 for space. Why? Well, Intel and friends have no interest setting up an entire fab process that will be utilized once in a blue moon. Even if you could use a more advanced node, the extremely high circuit area overhead required to realize the requirements I listed above will remain a limiting factor.
I would be amazed (and saddened) if someone had started a new, large-scale, serious, rigorous, engineering project in the past couple of decades (i.e. since the 90s) and used imperial units (I am sure there are lots of small-scale/personal/etc stuff done in imperial though - I dont count interplanetary space flight as small-scale!)
Imperial is pretty much dead in the UK (or maybe just London?) apart from roads and conversational/casual usage where its often easier to say the imperial equivalent than the metric (e.g. "pint of beer" is easier than saying "568 millilitres of beer", "about a foot" is easier than saying "about 30 centimetres" just because of fewer syllables if nothing else).
Also IIRC technically most engineering schools in the US also use metric these days, it doesn't mean stuff doesn't get screwed up, imperial on it's own is also annoying since you have both decimal (thous) and fractional units which I always found frustrating on it's own.
A verified DSL written in Haskell on the other hand ...
Oh, and s/wastely/vastly/ ;)