Vulnerable transistors threaten to upend Europa Clipper mission
science.org
science.org
This is a classic thing with Industry, they qualify a process that is working and shows good performance, but this process needs to be changed for reason XYZ, often because it is maybe a bit too expensive or doesn't align with the rest of their processes. The small change in the process wasn't that small and takes a little while to be identified because by the time you catch it you might be further down the line and this would be caught by a QA process and not a QC process, that might have deemed at that point not necessary because you had no reason to fault the part.
The second part is that some things are rated and verified but not tested extensively, since you might have prototype you might misdiagnose a failure of a component for a behaviour of your prototype, when in fact you had a deeper problem, but timelines with the added fact that so far you didn't think about that problem because it shouldn't have been a problem can catch you really off guard. This is usually where people testing the same thing in an exotic environment can ring alarm bells for others and that often happens at conferences...
People often under estimate how much you can get bitten in the back by such little details that become huge details.
Depending on the electronics and where the MOSFETS are, I would be them I would probably trash the electronics, take the spare that they had, validate components that get in and rebuild a control box and re-integrate it, provided that this is doable. It's expensive but provided that you have no choice that gives you a backup system that you can test code on before pushing it on the actual probe and might help for problem solving by being able to do measurements and test on the actual setup... Provided that they have the time and resources. Otherwise I wouldn't YOLO it given the fact that it might just straight up not work at the moment you need it the most and a little delay is better than nothing and they can spend the time re-checking part of the design that might also be weaker...
But heh, who am I but a random guy on the internet...
They found it was cured by lemon juice, but they didn't understand the details. Over years, they switched to lime juice (less vitamin C), put it in copper pipes (leaches vitamin C). But ships were faster so there was more fresh food available, masking the problem. Then scurvy starts mysteriously popping up again 100 years after it was first "cured."
Hard to keep track of the effects of all the details in the face of various co-dependent things changing simultaneously. Recipe for surprises.
This is a more-consequential example of the things you can learn by chatting with others; it is an extreme example of, "Hey, are you guys using components from Widget Inc.? Their datasheets are good, but sometimes we get a bad batch."
Those little things can save you a ton of time. In this case, it may have prevented mission-failure.
Part of the blame falls to NASA, too. If the outcome is your responsibility, then open-loop trust of a vendor for a known failure-mode may not be acceptable. Integration rad-hard testing may be requisite.
In the spacecraft environment, qualifying components is very difficult -- there's a good chance that NASA has these MOSFETs on an approved list because they've worked well before and have had few (or known) faults. They're probably not on that list anymore.
My guess is the parts failed TID at the more stringent levels, and Infineon didn't follow up with NASA or their contractor because they assumed that NASA was okay with the lower rad tolerance levels typical of space. Usually that would be the case, but Europa Clipper is special because it's going to an extremely harsh radiation environment.
The big question for me is: did the Europa Clipper program order a lower TID and try to upscreen, or did they order the high TID part? If it's the former, it's on NASA. If it's the latter, that's extremely concerning because Infineon should know that nobody orders expensive high TID parts for funsies, and they should have followed up with all customers as soon as they confirmed there was an issue. Just assuming NASA over-specified a part is absurd. The rad hard electronics market is small, everyone knows each other. Trust is king.
Finally, I'm not sure if it's the part in question, but it looks like Infineon discontinued their 1 MRad MOSFETs in 2020, citing low order volumes: https://irf.com/product-info/hi-rel/alerts/fv5-d-21-0004.pdf. In the light of this reporting, I have to wonder if there was more to it than that?
It's more likely that Infineon's folks talking to NASA were equally clueless about this change.
Is there any actual evidence they didn't reach out to every single buyer of the electronics?
The article goes out of its way to say Infineon did not contact NASA. But even in your description, they would not have, they would have contacted NASA's contractor working on the electronics.
I still go back to "if there was actual evidence that Infineon did not notify who it was supposed to, the article probably would have cited it". There isn't, so they instead cast aspersions.
Instead they make a bunch of hay about a statement from Infineon that seems totally innocuous - they didn't notify people they didn't know about. Shocker.
Look, i actually hate Infineon - i've been forced to try to make their wifi and bluetooth modules work properly before ;-)
But this kind of lazy-at-best journalism doesn't help anyone.
Personally, I despise status meetings. 99% of them are worthless fluff.
But, every now and then, you get something like this.
I think that highly usable dashboards are a good way to deal with this.
It's possible that AI could be a big help, here.
I experienced this kind of situation, where only by chance conversation was a crisis averted, very much at my last FT. So much that I'm working on a startup for fostering serendipitous communication for remote teams, like private notes from coworkers left on stackoverflow questions (or anything on the web)
Have you even been in one of those meetings that just won't finish despite everything being done? Making it written doesn't solve the problem. Instead, it makes it worse.
Computers are good at this though.
Now the only question is how you can automate the spec comparison such that issues with the spec and the parts used can be automatically compared.
And that starts with a computer readable spec that is updated by the manufacturer.
AI it is, then...
Surely they can be replaced! Humans (in cleanrooms) can put their hands inches away from them. Here's a picture, admittedly from 2022, before it was sealed:
https://europa.nasa.gov/resources/342/electronics-vault/
It's not the same as saying they've failed and are in orbit around Jupiter and so can't be replaced. They simply need to open and re-seal the vault or build and seal a new electronics vault before October.
In the meantime NASA built a whole spacecraft around it and the craft went through numerous tests. If they'd have to undo and redo all of that, they may even miss the 2025 backup launch date. I'm sure NASA is very eager to seek other ways to deal with this problem.
Not exactly responsible disclosure! NASA buys rad-hard transistors, and Infineon "didn't know what they'd be used for"?
But it's a reasonable idea to notify all potential large consumers that are likely to have bought your specialty product; these are not numerous, and the impact may be large (as in this case).
Did NASA assume these were rated higher then they were? Did Infineon make a mistake in documentation, or did they straight up not test them or test them incorrectly.
Unless contractually specified otherwise, it's generally up to the buyer to check the delivered goods for defects and report those without undue delay*. If this is not done, the goods are deemed to have been accepted.
Sure you can contractually specify that the product has to meet certain specs and pay extra for the seller performing QA, but the default often is "you're buying whatever comes out of our factory, check the goods yourself on delivery". The reason things are done this way in the business world is that it is generally cheaper to accept certain failure rates than to perform testing at every step of the supply chain and add a whole lot of bureaucracy and complications because of returns.
Whether custom contracts existed in this case is unknown, but it is likely that Infineon notifying customers was already a courtesy. They could've just said nothing.
* Under German law, which likely applies here since that's where Infineon sells from.
If the MOSFETs don't meet the specs on Infineon's data sheet, including rad hardness, then Infineon would be in breach of contract.
Is my reasoning correct?
German law differentiates between "open deficiencies" and "hidden deficiencies". If you neglected to properly check for an open one, that's on you. You now have no warranty under the law. In case of a hidden one, which will likely only show during large-scale production and can't really be detected beforehand, you have to immediately report it once you discover it, and it is your responsibility to document & prove that you did so without delay.
Under this system it's up to the buyer to decide how much reliability they need. They can forego testing and save money because it's not important to test every single screw when building a garden shed, or they can rigorously test every single thing because they're building a spacecraft.
* It is enough to prove that you did perform checks. If you got unlucky and the random samples just happened to be good, you are still protected. But if you didn't check at all or not sufficiently, you're screwed.
I don't really know this field, but might they have switched away from silicon-on-sapphire?
There are so many types of radiation that I do not think it unreasonable that they only notified customers who used these devices in particular environments. Most military use would be near radio transmitters (radars) or nuclear reactors (navy). Neither use case are an exact match for the radiation environment of Jupiter orbit.
If not, does that mean that maybe NASA is using them outside of their designed spec?
(See also: https://nodis3.gsfc.nasa.gov/displayDir.cfm?Internal_ID=N_PR...)
For keeping track of power use and interfaces specifically it turns out doing it all with SysML diagrams wasn't so great. Aside from all the pointless futzing around with boxes and arrows the model eventually became so huge the authoring software could barely handle just opening it up. So it must have been shortly after these slides when all the power use tracking was shifted to a custom tool with a more tabular user interface that we were already using for tracking electrical interfaces (slide 15) with version control in git.
There was also a semantic object-level diff we got for "free" by virtue of building on top of the Eclipse Modeling Framework. It was integrated into the Eclipse git UI and could help resolve merge conflicts without having to touch the XML directly, but merge conflicts were still annoying to deal with so generally engineers coordinated with each other to not touch the same part of the model at the same time.
Normally for review though I think users tended to compare reports generated from the model rather than trying to diff the source model files directly. There was a sort of automated build process that took care of that once you pushed your branch to Github.
People are reading this as Infineon didn't know that the parts were going into a probe when it's far more likely they meant they didn't know how the transistors are being used in that probe, which might have a large effect on whether or not the problem will affect them.
(Note: I'm not a physicist and have no idea what I'm talking about in this domain)
The problem being that high-energy cosmic rays are unlikely to interact with the lightly built spacecraft, going right through it. But if you add a thin layer of a good radiation shielding material, then there is substantially increased chance that they will interact with that material, and produce a very large spray of secondary particles. And those secondary particles will also be going fast enough that when they hit more shielding material, they will also result in more particles.
Then some of those secondary particles will be neutrons, which will easily penetrate the thin shielding (lead half thickness for 4MeV neutrons is 68mm), and irradiate the surroundings.
This has been very clearly demonstrated on the ISS, any metal tool has substantially higher radiation levels around it.
A thin lead sheet would be a rounding error next to that.
This is an oversimplification that's rather wrong, but: a decrease in altitude of just 300 meters, at airliner levels, puts an additional atmospheric mass equal to ~1 cm of lead (Pb) above your head.
https://www.nasa.gov/wp-content/uploads/2009/07/284275main_r...
Different sources of radiation interact with electrons or nuclei (1:1 with number of atoms) or nucleons (individual protons/neutrons, 1:1 with the mass). For instance, neutrons bounce off nuclei in nuclear reactors, and the lighter they are, the more energy the bounce can siphon off from the neutron. So having more, lighter (low-Z) nuclei (hydrogen in water and carbon in graphite are commonly used) provides better slowing of the neutrons vs. heavier (high-Z) elements, like lead.
Smashing ions (alpha, protons, heavy ions) into materials can also cause a https://en.wikipedia.org/wiki/Particle_shower
Lead "is effective at stopping gamma rays and x-rays" [1]. Jupiter's radiation comes from "trapped particles [that] are about ten times more energetic than the ones from the equivalent radiation belts of Earth" and "several orders of magnitude more abundant" [2]. When those encounter lead they cause bremsstrahlung radiation [3], a sort of subatomic shrapnel that can be more dangerous than the original radiation.
Lead is also heavy, which means not only increasing the mass of the spacecraft, but its balance and thus propulsion profile. That might mean upgrading and moving thrusters and propellant tanks--in effect, a complete redesign.
(It's a good question that doesn't deserve to be downvoted.)
[1] https://en.wikipedia.org/wiki/Lead_shielding
[2] https://www.spenvis.oma.be/help/background/planetary/traprad...
Were it designed today we'd probably dope it with titanium [1][2].
[1] https://www.tandfonline.com/doi/full/10.1080/10420150.2023.2...
[2] https://www.sciencedirect.com/science/article/abs/pii/S01491...
Hopefully SpaceX is able to resolve its Falcon second stage problems before Clipper is scheduled to launch.
* There were some discussions about adding a Thiokol Star 37 or Star 48 apogee kick motor to the Falcon Heavy stack for Clipper but for various reasons this didn’t happen.
> Falcon Heavy rocket, having three launches under its belt, has proven more powerful than originally anticipated. Previously, it was thought that launching Europa Clipper on a Falcon Heavy would require a “kick” stage — essentially a small booster attached to the top of the rocket. The Falcon Heavy’s impressive performance has made that unnecessary. Moreover, mission designers at Jet Propulsion Laboratory have found a path to Jupiter called a MEGA trajectory: after launch on a Falcon Heavy, Europa Clipper would fly to Mars for a gravity assist, and then return to Earth for another, and then on to the Jovian system. (The mission previously believed that the rocket would necessitate a Venus gravity assist, which would require special thermal protection for the spacecraft.)
> The window for a MEGA launch opens in 2024 and would take only three years longer than an SLS flight. A Falcon Heavy expendable launch is about $150 million. A single SLS launch is now estimated to cost $2 billion.
Source: https://www.supercluster.com/editorial/europa-clipper-inches...
You’re still changing the spacecraft’s balance. Imagine moving one of an airliner’s engines a foot to the left. It can be done. But it’s a big change.
Now consider that “modern jet airliners have…useful load fractions, on the order of 45–55%,” while orbital rockets’ payload fractions are “between 1% and 5%” [1]. Deep space craft are another order of magnitude more sensitive.
Adding a little shielding here and there is the aeronautical equivalent of hanging a bag of bar bells off the tips of one of the wings.
[1] https://en.m.wikipedia.org/wiki/Payload_fraction Note: useful load != payload fraction, but within orders of magnitudes they’re comparable
(I wonder if Starship is useful for this type of problem: if you could adapt the orbital-refueling method to serve as radiation shielding, and put an electronics vault in the middle of the propellant tank? Could you adapt Starship into a spacecraft bus in this way?)
2. Even a thin sheet of lead may be too heavy.
> In June 2022, project scientist Robert Pappalardo revealed that mission planners for Europa Clipper were considering disposing of the probe by crashing it into the surface of Ganymede for Europan protection purposes, in case an extended mission was not approved early. He noted that an impact would help the ESA's JUICE mission collect more information about Ganymede's surface chemistry
What about Ganymedian protection, eh? The Ganymedians, should they exist, are going to be furious.
Also it is weird that the outcome of the EJSM divorce (originally there was going to be a joint NASA-ESA mission to the Jupiter icy moons; Europa Clipper and JUICE are a result of the breakup), is that America explores Europa and Europe explores Ganymede, as the other way around would be less confusing.
Either the parts were in spec or they weren't. Which is it?
When the requirements for a part are specified, it is based on assumptions that may or may not hold true.
For example, if an issue tends to be all or nothing, then testing a small percentage of a lot should reasonably be expected to catch an issue. So you might specify that 1% of these transistors be tested and so long as that 1% passes the rest are considered good. If let's say there's a process change and lots become more variable, the confidence with which you can say the others are good based on that 1% testing goes down, but you are still testing to the same standard that you were before, which is what the specification calls for.
The issue gets even more thorny when issues are conditional. For example a part might meet the voltage specification, the temperature specification, and the radiation specification individually, but when you put that same part simultaneously in a high voltage, low temperature, and high radiation environment it doesn't perform as well. Or perhaps one component used downstream of a particular other component has an effect. Perhaps the most basic example is oversized but in tolerance shaft meets undersized but in tolerance hole.
I am not disagreeing, but at some point humanity should switch to actual probabilistic device models instead of vague datasheets. Imagine every datasheet has a .sample() function and you get a randomized SPICE model as if it came from the manufacturing line, you can draw a 100 and plot the properties. Want to measure the dynamic range of some ADC? A-weighted or not? Instead of specifying values highlighting figures of merit, each such figure of merit corresponds with an explicit SPICE circuit that measures that figure of merit on one and the same generator for random SPICE models with that specific part designation.
If a brand tries to fool its customers by insinuating a desirably high or low value by changing the test method, its immediately clear. A user may specify his own test circuit, or reuse the test circuit from the interactive datasheet from a competitor etc in order to make apples to apples comparisons for different DUT's.
More seriously, I'd be interested to learn what failed in the quality assurance process, as NASA & its suppliers have a legendary reputation in these topics. RCA will be enlightening.
They might not have "known", but come on, you're selling radiation-hardened chips to NASA. You can sure make an educated guess that they might be used for a probe.
I'm guessing there's a clause missing in the contract that says Infineon must disclose all known problems to NASA regardless of how the chips will be used.
Regardless, there are some people at NASA to whom 'Infineon' is now a curse word.
The article doesn't say or even imply that NASA has any contract with Infineon. It seems much more likely they are buying the chips through one of their approved distributors.
Without something saying that NASA bought directly from infineon:
1. It's not obvious how they would know who they sold to.
2. It's not obvious how they could get the information out beyond how they usually do it - issuing erratum notices.
Honestly, it feels like the article goes out of its way to try to imply Infineon should have notified NASA, but gives no data to suggest it had any idea at all what was going on.
If they had data that infineon and NASA had a contract, they would have put it in the article and used much stronger language. All these contracts would be public and are easy to find.
The fact that they don't have anything in the article about this suggests the contracts don't exist, and as usual, they are just using implication instead.
The question here is whether Infineon had a contract with NASA or otherwise should have known these were sold to NASA.
Again there is nothing cited in the article that says "yes".
If you've got data that says yes, awesome, what is it?
[1] https://www.eenewseurope.com/en/nasa-tests-infineon-power-mo...
[2] https://hardwarebee.com/electronic-breaking-news/nasa-tests-...
[3] https://www.eenewseurope.com/en/ir-hirel-rad-hard-components...
But do people ever actually "invoice NASA" for components. It was probably one of 100 different sub contractors building the actual circuits to NASA specifications, i.e. it was lower in the chain rather than NASA itself.
(Doesnt excuse the non-disclosure to those subcontractors)
Yes, absolutely they do. I'm not a part of this mission, but I'm currently working on another NASA spacecraft mission. I don't know the percentages off hand, but a substantial portion of our spacecraft is built in house with parts purchased directly by NASA from the manufacturer.
Regardless, there are lines of communication to subcontractors. The mere fact that they found out about this at a conference is significant evidence that Infineon didn't notify who they should have.
They may also use this as a spur to wargame failure mitigation strategies, so they'll be ready, if they do go belly-up.
IIRC some did not even reach Mars while others failed soon after orbit insertion due to various sub-systems failing.
While mars is an elevated radiation environment when compared with earth, the Jovian radiation belts are on a whole other level, particles up to 1-2000 MeV are fairly common. To put that into context, a medical radiation beam therapy deals with 2-300 MeV on the absolute highest end. To get into the 1-2000 MeV range you generally are talking about energies found in the low end of particle accelerators. Ingenuity mostly had to worry about Total Lifetime Dose (TLD), one example of a TLD issue is dopant migration induced by high-energy heavy ion collisions which can change the on voltage of a transistor. At high energies you can have single events with enough energy to cause fatal latch-ups. For instance modern rad-hard FPGAs start encountering major issues around 60-70 MeV.
Furthermore, these parts are power MOSFETs which control power for whole subsystems so their reliability is critical to the operation of the spacecraft. In addition, the biggest issue here is not just that there were issues that were addressed and fixed, it's that Infineon didn't issue an errata to the datasheet or inform NASA of the issue. As a result there are now transistors littered throughout the spacecraft which don't meet the radiation needs. This is going to require reworking the boards, re-validation of the subsystem, and re-integration of the subsystem into the spacecraft. This all comes at a non-trivial impact to budget and timelines which is to say nothing about what this does to the launch window the project was trying to hit for gravity assist / proximity.
I hope you find this informative! :)
EDITS: Spelling and an "is"
https://en.wikipedia.org/wiki/Ingenuity_%28helicopter%29?wpr... (in case anyone wanted to refresh their memory)
> Engineers at NASA’s Jet Propulsion Laboratory (JPL), which leads the development of Clipper, discovered the problem in May after talking with colleagues about a classified satellite at a conference.
...I immediately think "that's what they want us to think"
if classification of projects means anything, the true intent of the passing of dis information is something dat we can only guess at.
fnord.
This is criminally incompetent on the part of Infineon. WTF, NASA could use those transistors for a fancy inteliggent toilet FWIW, it doesn't matter, NASA doesn't have to tell you how they are going to use those parts. They bought parts based on a fucking SPECIFICATION, and if the parts you sold them don't meet the specs, you communicate immediatelly with the customer offering a replacement for free.
Really, someone should be jailed for that.