NASA can't figure out what's causing computer issues on the Hubble telescope
npr.org
npr.org
There doesn't seem to be any nuance or respect that they're trying to repair an orbiting telescope that was launched 30 years ago and designed 40 years ago - and that people are patiently trying to sort through a fully autonomous system 400 miles above the surface of the Earth with a very large set of failure options.
For me - huge props to NASA and other organizations that do this kind of work and keep these systems running for decades. I need to reboot Windows every 2-3 days
> Overworked Engineer misses semicolon. All night review session finds it, data gets loaded!
> Management insists Jira stories be routed to new Epic. Team Lead spends hours learning Jira API before giving up and doing it "the hard way"
That's a positive spin on clickbait.
Extra points if you can throw "Einstein" in the headline.
I didn't get that impression.
This does intrigue me - I like browsing hackernews, but I often get the impression that some people (not the PC specifically) here are either ridiculously anal about English or genuinely do not parse sentences the way I do.
People with autistic traits sometimes parse sentences in an overly literal or precise manner. I rarely do this anymore, but when I was a child and teenager I did it more often.
When I'm speaking of "autistic traits", I'm not speaking just of people diagnosed with autism/ASD (who are of course represented here), but also people with broad autism phenotype (BAP), the subclinical manifestation of ASD. BAP is when you have more of the symptoms of ASD than the average person does, but not enough to justify an actual diagnosis of ASD. BAP is quite common in software engineers, and STEM professionals more generally, so I think there are likely a lot of people on this site with BAP (albeit most of them have probably never heard of it.) The people you are talking about quite possibly do have some degree of BAP, and this behaviour is quite possibly a manifestation of their BAP.
That's what the article says. The headline is clickbait.
In terms of downvotes, what really irks me and what I often see is people posting factually correct information, but still being sent into faded oblivion because some sect of the community's worldview doesn't agree with the facts.
If someone uses factually correct information to make a comment thread worse, I can see how downvotes could be justified.
Since comment scores was removed this is the only way to signal this to others besides the original commenter.
That said it should not be overused. If it annoys someone I guess they should downvote it but I don't think there is a need to reflexively downvote every time someone adds a friendly meta comment.
(And if people start gaming it for karma farming I guess it should be downvoted relentlessly until that stops :-)
So no, it's probably not as useful as the comment it was referring too, but it was useful (to some of us) as it pertains to the community as a whole.
Well, in opposing it I especially read the faded comments and upvote any of those that are not completely abhorrent.
Take that, ycombinator.
This seems to happen a lot more frequently here than anywhere else.
I'm not really sure what that says, other than people still read comments that are faded. Also that people shouldn't worry about self-censoring.
I don't have a problem with downvotes or the karma needed to do it, but
I do sometimes wish it were possible to reply to a dead comment, especially if you vouch for it and it's still dead.
Sometimes they're worth defending, or is relevant in a non-obvious way, and sometimes the comment itself is discussion worthy, as it relates to the topic, even if it's wrong or seems trollish.
I'm being sincere here since I also appreciate book recommendations and I get probably half my book recommendations from HN.
https://www.youtube.com/watch?v=2KSahAoOLdU&list=PL-_93BVApb...
My favorite part is when they needed the version of the software that was used for the moon landing but they only had the source code for a previous version (scanned from giant binder) and the hash value of the version of the landing. By a series of educated guesses, by reading memos and by analysis of the source code they modified the old code the exact way so it gave them the correct hash, confirming that they correctly and exactly recreated the original code.
It's being a while and I go from memory, I might have some details wrong. See this video for this story. https://www.youtube.com/watch?v=-JTa1RQxU04
Apollo 13 was the story of a 'successful failure', while Apollo 14 shows how hard work and creative thinking can turn failure into success.
I saw this but it's a short documentary and may not be what you meant: https://www.amazon.com/Apollo-14-Complete-Downlink-Edition/d...
Jim Lovell was undeniably a badass but I watched that movie and thought the heroes were the ones reading telemetry off a computer screen and using their slide rule to figure out what to do. I hope Hidden Figures does that for another generation.
I noticed on the last season of the Expanse that Luna headquarters was named after him and bothered to look him up. Dude's 93 and still kicking!
It's amazing how well the astronaut medical screening worked. Unless they get killed in the line of duty, these guys are all living incredibly long.
On the other hand watching a stream involving something on Mars, let alone voyager, would be pretty boring!
Send: ls
Ok let's take a 20 minute break.
For LEO satellites, that usually means you have 2 blocks of 3 13 minute passes, when the groundstation is in The Netherlands. For a Svalbard groundstation, you get a lot more, but still 13 minute or less passess.
Imagine dictating shell commands, over the phone, to a salesperson who has no idea what half the characters are that you're asking him to type, and the only output signal I could come up with was ejecting the CD tray, which was just visible from the ground...
(Note that the goal wasn't usually to fix things on the spot, it was more to triage things like whether we needed to have a replacement projector on hand, which was a big deal.)
I don't know enough about ssh and terminals to know if it's possible to type "12345" and see "12345" echoed back to me but really what the remote session sees is "1245".
Nowadays, links and systems get easier to work with, and you can sometimes have a literal TTY open to the system, like Reactor Hello World has ( https://reaktorspace.com/reaktor-hello-world/ ). However, this is over S-band, which is a 2Mbit/s link, so overhead for a stable TTY (or ethernet connection) is a lot less than using UHF/VHF.
Ground processing was a totally separate program and contract from mission control, though, so my team only ever received data from the satellites. We never sent data to them.
You can see how NASA people react to tough situations by watching the videos of mission control during the Challenger and Columbia disasters. No shouting. No arguments. Just cool professionalism and restrained emotions.
They have a job to do, and they do it well even under stress. "Steely-eyed missile men/women" indeed.
There are plenty of NASA engineers and leaders who lose their cool. I’m only saying that so people don’t overly lionize them in a way that prevents them from pursuing a similar job because they feel they are somehow cut from a different cloth.
“The O-rings were never tested in extreme cold.”[1]
There wasn’t data which led to discussions about uncertainty, but that shouldn’t be conflated with irrefutable evidence of failure.
The obviousness of it (like many engineering failures) was only apparent in hindsight.
“Evidence, in retrospect, points to a long period of time, especially based on post-flight inspections when the joint design weakness was ‘sending a message’ and the true potential of this message was not perceived and reacted to.”[2]
“Not perceived” isn’t compatible with “irrefutable evidence that it would fail”.
[1] https://www.space.com/31732-space-shuttle-challenger-disaste...
[2]https://www.govinfo.gov/content/pkg/GPO-CRPT-99hrpt1016/pdf/...
Satisfying, but not exactly must watch tv.
What's in your head could be though. That's my pet theory on the movie Hackers, what we're seeing on the computer screens isn't what's actually there, it's the characters' mental constructs visualized.
Think of this as the computer equivalent of that scene.
Guy 1: Why don't I just poke him with the pointy bit
Great scene though, makes me want to watch the whole movie.
It includes powers like becoming weightless, killing with a single movement, flying, etc.
You have a computer that you can only interact with over a radio link, and need to make it start working again with only what you know about how the system is built and a limited set of remote commands. Sounds like something I'd get obsessed with solving.
Confirming that not being able to ping/connect to it during the failed attempts was absolutely not exciting :)
Sigh.
pg down
Sigh
...
Repeat for hours.
And considering there are no life or death stakes, it still wouldn't be as exciting as Apollo 13.
This guy took about 30 hours of video of him porting an 80s version of unix to the ESP8266. Warts and all -- live!
I've started to watch it and it's fascinating!
https://www.youtube.com/watch?v=cDHcGY7EzUM&t=62s
You could have a whole channel with different teams debugging satellite technology and if you're bored, it would probably be quite interesting. The bigger problem is most likely concerns about IP and secret protocols and so on.
"And now Bob will log into TeleSat123 via SSH." <We see bob type in root / password123>
"Oups, uh..gosh we'll just go to a commercial break!"
Did that work... no. Well, what about... THIS... still no. 3 hours later... clear the cache?!? Aww crap
> Most of Hubble's components have redundant back-ups, so once scientists figure out the specific component that's causing the computer problem, they can remotely switch over to its back-up part.
Wait, then why don't they just switch over each component in turn? The "divide and conquer" debugging strategy.
My guess would be that they want to try that method only if this debugging doesn't work. Imagine that there's an electrical issue in item 1 that fries item 2. If you switch over to item 2b, then you fry item 2b too!
This is exactly what happened with the Soviet Salyut 7 station. They tripped an over-current protection, didn't fix the root issue, and remotely turned the circuit back on. A series of electrical shorts then rendered the entire station without power, resulting in the need for one of the most daring station rescue stories of all time:
https://arstechnica.com/science/2014/09/the-little-known-sov...
Are shifts this long still common practice in US or RU space programs?
The best case scenario of a bunch of engineers flailing about on a bridge turning knobs is that you luck into a fix but don't know how you got there. But you're more likely to make things worse.
Business requirements and requests change all the time. 90% of our work is done in response to that. The other 10% is fixing up technical problems due to increased scaling or bugs found, and then basically never do we upgrade a system to keep up with security updates or change to a more modern tech.
Let's say the CPU is the actual issue, but the problem manifests itself in the memory module. You swap over to the backup memory module, and suddenly the problem vanishes!
Two months later, the problem manifests again. Identical presentation. This time, there is no backup to switch over to test.
You fly a Very Expensive Mission to the telescope only to find out the CPU was the issue, and if you had figured that out originally you wouldn't be up here with four memory modules.
So you swap in the backup storage module and all your problems go away. Until it happens again and corrupts _that_ too.
This is fascinating to me. Do you have any pointers to information/research/projects focused on hard real-time garbage collection? A Lisp with hard real-time garbage collection (even if Herculean to implement) would be fantastic.
http://web.archive.org/web/20020331165324/http://home.pipeli...
Almost C style, and I'm sure just as error prone, but it seems like it could work.
Lisp with ref counting (assuming acyclic references) instead of GC could also be interesting to try to hack together. I have a feeling closures would be a particularly quick way to get reference cycles, so some concept of weak references may be necessary.
Some of those elements will be part of the major service windows, and have expected operational and standby lifetimes.
So if a component with two elements has a service window of 10 years, and each element contributes to meeting that service window, then you've bumped your major service window from 10 years to a significant factor less than that.
e.g. the expected use profile might be: use element 1 for 6 years or 60% of service, switch to element 2 for 4 years, replace both during 10 year maintenance window. Interrupting that by bringing element 2 up reduces that window and contingency plans if the service window cannot be met.
I don't know, and I'm just talking out my you-know-what.
It's better to understand the problem than to just start changing stuff hoping you find the right thing even systematically. There's not a huge rush to fix this since it's the payload computer and the telescope is still being maintained by other systems. A lot bad could happen, if the switching system is flakey you could maybe get stuck in a bad system, or if there's a number of faults you might damage one of the backups. Without the shuttle there's not a plan to service it any more so why take the risk rushing through to the most simplistic debugging method?
Also a repl in space only makes sense in earth orbit, but not farther away, with 8-20min waiting time for a packet roundtrip to Mars. Those machines really need proper and faster decision making (AI, think lots of `if` statements and proper modeling) on board, eg to perform landing or docking maneuvers. Or to detect and workaround radiation damage in its own circuits.
- First they had a DF-224 flight computer and a - Science Instrument Control and Data Handling (SI C&DH)
Initially DF-224 between missions got installed a coprocessor: https://asd.gsfc.nasa.gov/archive/hubble/a_pdf/news/facts/Co...
During another servicing mission they replaced it with something called the Advanced Computer with Intel 80486: https://asd.gsfc.nasa.gov/archive/hubble/a_pdf/news/facts/FS...
It looks like its about 50,000 lines of code in the C and Assembly programming languages.
https://www.nasa.gov/pdf/327688main_09_SM4_Media_Guide_rev1....
Fig 5-10 is the Data Management Subsystem
https://asd.gsfc.nasa.gov/archive/sm3a/downloads/sm3a_media_...
"Welcome to the Hubble Space Telescope Help Desk"
"The operations team is investigating whether the Standard Interface (STINT) hardware, which bridges communications between the computer’s Central Processing Module (CPM) and other components, or the CPM itself is responsible for the issue. The team is currently designing tests that will be run in the next few days to attempt to further isolate the problem and identify a potential solution."
So "can't figure out" sounds more like "haven't yet figured out", but they have remaining ideas to play through.
[1] https://www.nasa.gov/feature/goddard/2021/operations-underwa...
Of course they do! I wonder if they ever had to put another part out of service. I also wonder whether the twin of the part could also suffer the same failure at the same time without being used.
Hubble started out with 6 gyroscopes, in 1996 they replaced four of them, by 1999 four had failed so they replaced all six, by 2009 three had failed again, so they replaced all six. Now they are again down to three, and one of the remaining ones has a defect that required some workarounds. The last three gyros are at least a new design that should last a bit longer.
No part of that effort has actually repaired anything in space.
Funny enough, nobody posted the link to the article that says "70% of bugs are memory issues" (or something like that) yet.
https://news.hitb.org/content/microsoft-70-percent-all-secur...
This isn't a security issue, NASA isn't Microsoft, and physically degraded memory isn't the same as a memory safety programming bug.
I'll certainly bet that article is super popular with the rust crowd though.
(And yes you're right none of this has to do with this stuff, for sure.)
https://baltazaar.wordpress.com/2009/07/20/a-story-about-lis...
https://www.nasa.gov/feature/goddard/hubble-memorable-moment...
And as someone who has been invited into "war rooms" by managers who do the "you're smart so of course you can help these other smart people stuck on a hard problem" there's a real skill to being able to read the room and know when to back the fuck out of it or just shut the fuck up -- which most intellectuals don't have. Sit and listen for awhile and take it in. Then maybe take your best idea and ask a very toned down question. If the person who seems to be leading the troubleshooting instructs you on why that's wrong, throws in 3 neighboring ideas that also don't work, with 5 reasons you haven't considered for why that's the entirely wrong path, then just nod in agreement and be quiet and see what you can learn from the domain experts.
Peppering that team with a dozen outside "experts" is going to be useless because they'll just start getting really defensive after awhile, and even if someone winds up throwing out the right solution they'll probably reflexively reject it.
OTOH if that team ASKS for someone who has expertise the team lacks and needs, then go assemble a team skilled at the use of cellphones and the internet to hunt that person down and drag them into the conversation.
Despite the article trying to phrase it as if they have no idea what to do, they know there computer incredibly well, it's a matter of going through the steps and isolating the problem.
Most likely school/college/university teams knowledgeable in the subject matter would be the "individuals".
NASA hasn't yet figured out what's causing computer issues on the Hubble telescope
Edit: and Hubble was built from a surplus skeleton of one, through a government transfer.
https://en.wikipedia.org/wiki/USA-245
USA-314, launched this year, is allegedly a KH derivative.
The question is not "is the US military still doing telescope-based surveillance?" it's "is the US military doing it with Hubble-class stuff, or are they doing it with newer stuff?"
Given the age and amount of hands-on maintenance required to keep the Hubble working, it's almost certain they're doing it with newer stuff.
So it's time to do what every gamer does when the rig fails. Switch parts and see who's the culprit. And yes, I do understand the next quote: <"The rule of thumb is when something is working you don't change it," Hertz said. "We'd like to change as few things as possible when we bring Hubble back into service.">
But between having nothing anymore, since Hubble had its last maintenance in 2009 (per quote: "The last time astronauts visited Hubble was in 2009 for its fifth and final servicing mission.") and have something now that definitely would fail later, I'd choose the latter.
Hubble is different. It's not like it's a ship that you can board. So you need two things: Ability to attach yourself to Hubble, and ability to leave Dragon to perform a spacewalk. It's not clear whether you can just have everyone in the Dragon suit up and open the hatch. And even then, you still need to attach yourself to Hubble somehow. I think you can via the port... but then you can't leave. Unless you go out the other door? Can you open that from the inside and get out with a space suit?
My rambling isn't meant to be an actual answer. It's more to show that it's wayyyy harder than "Let's just send up some people to Hubble!".
AIUI, The ISS can be serviced from the ISS if the appropriate supplies and personnel are sent up, but it doesn't have the delta-V to zip around other orbits servicing other satellites, so it is okay without the the shuttle for itself, but doesn't substitute for it for other things needing orbital service.
[0] Except Soyuz I guess their orbital module would allow you to keep the descent module pressurized but it's still way outside the design so there's no telling if the module would remain operational.
Solar probe sent to comet, rebooted years later to try to put it back into it's solar mission.
They've just mostly described my career.
I am going to go out on a limb here and post my diagnostic: There is some global counter that overflow as the system was not rebooted for a while...NASA...take it from here :-)
But in all honesty, how would an internet across the solar system work?
There's open efforts to work on a protocol that would work with the extreme latency and packet loss. They're really quite fascinating.
Look into "Delay-tolerant networking" for more details.
Realistically, each ~150 ms sphere would have its own cloud infrastructure. Those systems would then bridge with one another. So idk AWS on Earth and DogeNet on Mars.
I would love for a distributed model as much as the next guy, but it's unlikely to happen for the same reasons that it isn't happening today.
The idea is sound, but the zone needs to be a bit bigger in reality. I think the moon is close enough to earth to be in the same zone (assuming antennas on "both sides", and special routing protocols to deal with day/night cycles)
(all 20 years away ofcourse)
I do not understand why they can't just swap out the computer for a better and more modern one. Am I missing something here?
The space vehicles we used for this purpose have been retired and we have no replacements.
When dealing with high-latency, high-radiation environments, more modern isn't necessarily better: denser ICs mean greater susceptibility to radiation (and consequently more expensive hardening). They also can't exactly fly up there and swap out random bits on short notice -- I'm not sure if the US even has a the current capability to perform physical maintenance on the Hubble.
> Hertz said that because Hubble was designed to be serviced by the space shuttle and the space shuttle fleet has since been retired, there are no future plans to service the outer space observatory.
It's arguably an interesting question whether the government would consider using commercial relay satellites instead of just the Quasar constellation, though. The data stream doesn't need to be decrypted on the satellite, just forwarded. Obviously, you can't prevent radio from being intercepted, so throwing in a hop you don't own doesn't actually add any risk. You're totally reliant on the strength of your encryption either way.
Getting downvoted so I guess I need to clarify I was being sarcastic. I work in IT and we say this a lot.
The most common sentiment I see when people attempt cheap jokes is that Reddit is a more appropriate forum.
(It's UFO season! Everything is aliens again.)
I'd like to learn how to use Rust to work around memory corruption resulting from irradiating the hell out of the RAM in my PC at home.
I'm talking enough radiation to flip more bits than ECC is capable of detecting & fixing.
I want to do all of my programming remotely… right next to the the Elephant's Foot.
583 closed, 156 open. That much to "memory safety" in rust.