HNHacker News
TopNewBestAskShowJobs

thewonderidiot

334 karma · joined July 21, 2014

submissionscomments
thewonderidiot··on We found an undocumented bug in the Apollo 11 guidance computer code
Mike Stewart here! I led the restoration of the AGC documented on CuriousMarc's channel and co-administrate VirtualAGC. There is a lot to unpack here.

First: this is indeed a real bug in the AGC software. However, it did not go unnoticed for the whole program. It was discovered during level 3 testing of SATANCHE, and late development branch of the Command Module software COMANCHE. It was assigned anomaly number L-1D-02, and was fixed between Apollo 14 and 15. There are two known surviving copies of the L-1D-02 anomaly report:

* https://www.ibiblio.org/apollo/Documents/contents_of_luminar...

* https://www.ibiblio.org/apollo/Documents/contents_of_luminar...

The fix described in the article is partially complete, but as noted in the anomaly report there's a little bit more to it. Rather than just adding the two instructions to zero LGYRO, they restructured the code a bit and also cause it to wake up pending jobs. You can compare the relevant sections of the Apollo 14 and Apollo 15 LM software here:

* Apollo 14: https://github.com/virtualagc/virtualagc/blob/master/Luminar...

* Apollo 15: https://github.com/virtualagc/virtualagc/blob/master/Luminar...

The bug would not manifest silently in the way described in the article. For starters, LGYRO is also zeroed in STARTSB2, which is executed via GOPROG2 on any major program change: https://github.com/virtualagc/virtualagc/blob/master/Luminar...

This means that changing from any program to any other program would immediately resolve the issue. This is almost certainly a large part of why it took them so long to notice. Hitting BADEND while actively pulse-torquing is quite rare, and avoided by normal procedure. The scenario presented in the article can't happen since the act of starting P52 will zero LGYRO.

Moreover, in the very specific scenarios in which the bug can be triggered and remain, it results in multiple jobs stacking up attempting to torque the gyros. Eventually the computer runs out of space for new jobs -- similar to what happened on 11 -- and a 31202 (the Apollo 12+ equivalent of 1202) is triggered.

Since the issue was found before the flight of Apollo 14, a further description of how it might occur and what the recovery procedure should be was added to the Apollo 14 Program Notes: https://www.ibiblio.org/apollo/Documents/LUM159_text.pdf#pag...

Some other notes:

> Ken Shirriff has analysed it down to individual gates

I've done the bulk of the gate-level analysis. :)

> the Virtual AGC project runs the software in emulation, having confirmed the recovered source byte-for-byte against the original core rope dumps.

We've only been able to do that in very specific circumstances and only for subsections of assorted programs, but never for a full program. Most AGC software either comes from a program listing, from a core rope dump, or from reconstruction using changelogs and known memory bank checksums. We've disassembled all of the rope dumps into source files that assemble back into the same binary, but the comments and labels will be different from what was in the original listing. And to be extra clear: I've never had the opportunity to dump a module containing Apollo 11 software for either vehicle. Our sole source for both programs is a pair of printouts in the MIT Museum's collection.

> Margaret Hamilton (as “rope mother” for LUMINARY) approved the final flight programs before they were woven into core rope memory.

Jim Kernan was the rope mother for Luminary at least up through Apollo 11. Margaret was the rope mother for Comanche, the CM software, and was later promoted to lead the software division. Their positions at the time of 11 can be seen on this org chart: https://www.ibiblio.org/apollo/Documents/ApolloOrg-1969-02.p...

> Their priority scheduling saved the Apollo 11 landing when the 1202 alarms fired during descent, shedding low-priority tasks under load exactly as designed.

This is a huge topic on its own, but the AGC software was not designed to shed low-priority jobs. Ironically, the lowest priority job during the landing was the landing guidance itself, with high-priority jobs being reserved for things that needed quick response like antenna movements or display updates. If the computer were to shed the lowest-priority jobs, it would shed the landing guidance. This memo contains a list of all jobs active during the landing and their priorities: https://www.ibiblio.org/apollo/Documents/CherryApollo11Exege...

> For example, the ICD for the rendezvous radar specified that two 800 Hz power supplies would be frequency-locked but said nothing about phase synchronisation. The resulting phase drift made the antenna appear to dither, generating roughly 6,400 spurious interrupts per second per angle and consuming roughly 13% of the computer’s capacity during Apollo 11’s descent. This was the underlying cause of the 1202 alarms.

The frequency-lock prevents phase drift, so the phase is essentially fixed once the power supplies are up. Ironically, however, the bigger issue is that one reference was 28V while the other was 15V. Initial testing on actual Apollo hardware suggests that at least for Apollo 11, this voltage difference was the key contributor rather than the phase difference: https://www.youtube.com/watch?v=dT33c70EIYk

thewonderidiot··on Virtual Apollo Guidance Computer
For the communications project specifically (and also a couple of other projects we have going on), I combined the FPGA AGC and the Monitor into a single design that runs on a Digilent Cmod A7-35T: https://github.com/thewonderidiot/cmod_agc

It's a lot cheaper and more accessible in this form, and much easier to integrate into projects that need AGC stand-ins.

Aside from the FPGA design though, yeah, all of the AGC software and the assembler is in the VirtualAGC repository.

thewonderidiot··on Apollo 11 implementation of Trigonometric functions (1969)
The Command Module system test code Sundial includes a slightly earlier version of these same routines: https://github.com/thewonderidiot/sundiale/blob/master/sundi...

This is a work-in-progress disassembly of the core rope modules we dumped at the MIT Museum, so apologies for the less friendly formatting!

thewonderidiot··on Apollo 11 Guidance Computer source code for the command and lunar modules
No, it's not. The 1 here is using interpretive index register X1 to index onto GAINBRAK. The star on the DMP means that multiplication is indexed. GAINBRAK is one of the pad-loaded descent targeting parameters; the index selects the appropriate number based on the phase of the descent.
thewonderidiot··on Apollo 11 Guidance Computer source code for the command and lunar modules
People frequently seem to think this is about the line number it's on (666), but that doesn't have anything to do with it. That line number is a totally modern construction; the original source code was on punch cards. That particular comment was punched onto card number 0562 in the LUNAR LANDING GUIDANCE EQUATIONS log section. The original developers only referred to code by punch card number, and page number in the listing. So it really is just a coincidence, and the "NUMERO MYSTERIOSO" is referring to whatever is in GAINBRAK,1.
thewonderidiot··on Apollo 11 Guidance Computer source code for the command and lunar modules
That call to STOPRATE was there to zero out attitude rate commands at the moment the astronaut switches into the semimanual final descent program P66. It was removed in the final few revisions before the first, unflown release of the Apollo 14 software (Luminary 163 [1]) because it was preventing attitude control of the spacecraft when Rate-Of-Descent commands were skipped [2]. Skipping ROD commands wasn't normal, but was something that was added as part of the effort to make the computer more cleanly handle large unexpected additional load, like happened in Apollo 11.

[1] https://github.com/virtualagc/virtualagc/blob/master/Luminar...

[2] http://www.ibiblio.org/apollo/Documents/LUM156_text.pdf -- look for PCN 1037

thewonderidiot··on Apollo 11 Guidance Computer source code for the command and lunar modules
These files were originally put on GitHub in the VirtualAGC repository [1] in 2015. The code was publicly released and digitized into these source files in 2009, through cooperation with the MIT Museum.

[1] https://github.com/virtualagc/virtualagc

thewonderidiot··on Apollo 11 Guidance Computer source code for the command and lunar modules
Wherever we added modern comments, they are denoted by a double hashtag (##) at the beginning of the line. Everything else, minus formatting concerns, is directly from the original listing. You can see the scans we transcribed this from here:

https://archive.org/details/Comanche55J2k60

https://archive.org/details/Luminary99001J2k60

thewonderidiot··on Weird-Looking Freak Saves Apollo 14 (1971)
Don is still around, and is still awesome! His memory for events around that time is surprisingly sharp, for it being 50 years ago.

He's been an invaluable resource for us at the VirtualAGC project [1]. He provided over half (!) of the AGC programs we've made available, including LM software for Apollo 5, 9, 10, 11, 12, 13, and 15-17, plus several other ground development programs. He also was the source of a majority of the schematics that we used in the AGC restoration [2] -- I can say with 100% confidence that it wouldn't have been possible to get that computer working again without his help.

He's got a book out now, Sunburst and Luminary [3], that is very good and goes into a lot of technical detail on the LM software, without getting too much into the weeds. I highly recommend it if you're interested in that sort of thing!

[1] http://www.ibiblio.org/apollo/

[2] https://www.youtube.com/watch?v=2KSahAoOLdU&list=PL-_93BVApb...

[3] https://www.sunburstandluminary.com/SLhome.html

thewonderidiot··on A computer built from NOR gates: inside the Apollo Guidance Computer
The Interpreter was written by Charley Muntz, not Margaret. Margaret wrote a large chunk of the DSKY display routines and a lot of the alarm/abort logic.
thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
Nah, they put a couple of precautions in to make sure you couldn't put the computer into standby accidentally. First, software has to set a bit to enable standby mode to be entered. For the astronauts, this was keying in VERB 37 ENTER 06 ENTER, which started P06, the pre-standby program. They then had to press and hold the button for a period of time between 1.28 and 2.56 seconds (exactly how long it takes depends on where the clock divider chain was at when they first press in the button). To bring it out of standby, you have to push the button in for a similar duration.
thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
Many years now of research and simulation of the system (I led the restoration of the computer mentioned in the article). There's not a single place where you can read everything, unfortunately, aside from the comment above. We're planning on making a video on it in the future. But I can cite sources:

CDU theory of operation (starting PDF page 15): http://www.ibiblio.org/apollo/Documents/HSI-208435-003.pdf

CDU coarse module schematic: https://archive.org/stream/apertureCardBox462NARASW_images#p...

Grumman memo (from 1968!) describing the problem, and mentioning it is due to the reference switching to a 15V 800Hz source: https://www.ibiblio.org/apollo/Documents/Memo-GAEC_LMO_541_1...

Excerpt from the LM-8 Systems Handbook showing the reference voltage RR switch wiring: https://i.imgur.com/fMsQ7RI.png

Don Eyles describes the software side best in his book Sunburst and Luminary (which I highly recommend) but he also talks about it in some detail on his website: https://doneyles.com/LM/Tales.html

thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
Solder is used extensively in satellites nowadays. The catch, though, is that you must use leaded solder, because the lead dramatically decreases the occurrence of tin whiskers. RoHS has been rather annoying in the aerospace industry, because anything RoHS compliant can't safely fly.
thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
Did it use wet tantalums, though? If the alternative is "solid tantalums", then it didn't -- SCD 1006755, the only type of tantalum used in the computer, specifies them as "Capacitor Fixed, Electrolytic, Solid Tantalum". Our computer has KEMET KGxxJyyKPS parts (xx=capacitance, yy=rating), but they also at least evaluated Sprague 150D, 151D, and 350D to fill the role (and utilized at least one of them).
thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
Sadly most of what you read by googling these problems is misinformation. It was actually an incredibly sinister systems integration bug, that wasn't well described even at the time.

It wasn't a switch misconfiguration; the Apollo 11 astronauts were flying to the checklist, and did as they had simulated. The Rendezvous Radar switch has three settings -- LGC, AUTO TRACK, and SLEW. In LGC mode, the AGC controls the positioning of the antenna; in AUTO TRACK, the radar automatically tracks the CSM based on return strength; and in SLEW it is automatically positioned.

The trouble came from how the trunnion and shaft angles of the antenna were measured. They used "resolvers", which are sort of like variable transfomers. Resolvers look like motors, and attached to the shaft there are two windings positioned 90 degrees apart from each other. An AC "reference voltage" is applied to an outer winding in the case, and that voltage couples onto the two inner windings with a magnitude proportional to the angle on the shaft. One winding (the "sine" winding) produces an output equal to Vrefsin(theta) and the other ("cosine") winding produces an output equal to Vrefcos(theta), where Vref is the reference voltage and theta is the angle of the shaft. The voltage and phase of both windings can be used to determine exactly what the theta was that produced them.

The circuitry to do this is a bit involved though and lived outside the computer, in a device called the CDU, or Coupling Data Unit. The CDU constantly maintained its own idea of what the angle ("psi") in a digital register. It translated the incoming sine and cosine voltages into a digital representation by mechanizing the equation +-sin(theta-psi) = +-sin(theta)cos(psi) -+ cos(theta)sin(psi). It did so by using the bits of its digital register containing psi to switch on and off resistor dividers that effected cos(psi) and sin(psi) onto the incoming signals, which were then added together with a summing amplifier. The goal of the CDU is to zero this sum; to accomplish this, it "counts" the angle register up or down to reduce the magnitude of the sum. As it counts, switches are changed, which switch out resistors in the circuit, which in turn change cos(psi) and sin(psi) in the above equation. And also, with every other increment, a pulse is transmitted to the AGC to indicate that the angle has changed slightly.

The problem comes in because in addition to the above, the CDU also, for many angles, added to the sum some fraction of the reference voltage directly. This is fine when the switch is in the LGC position; the resolvers are supplied with the same 28V, 800Hz reference voltage that is used inside the CDU. However, when the switch is put in either of the other two positions, the reference voltage for the RR resolvers is switched to an unrelated 15V rail. Critically, this 15V reference has no defined phase relationship with the CDU's 28V 800Hz reference. The phasing is locked in by the exact millisecond at which you power up your subsystems.

So when the switch is changed, the sine and cosine outputs from the resolver are suddenly derived from the 15V reference -- they are much lower before and at a random phase. The CDU doesn't know that this has happened, and still tries to perform the summing as before. However, for many theta/phase relationships, it becomes impossible for the CDU to actually null the sum. In these cases, the CDU becomes "manic", and starts seeking back and forth, frantically changing switches to try to figure out what the angle is, but never succeeding.

This causes a huge flurry of +1 and -1 pulses to the AGC. In order to minimize circuitry, the AGC implemented what was called "unprogrammed" or "cycle-stealing" instructions. The computer only contains a single adder, and adding or subtracting 1 from the current angle requires use of that adder and a memory cycle. Rather than generating a full interrupt, which would require many memory cycles and instructions to handle, the computer simply transparently inserts a single-cycle instruction in between two "programmed" instructions that performs the addition or subtraction. This is totally transparent to software, normally. But with a manic CDU that is incessantly seeking on both RR angles, the AGC receives something close to 12,800 pulses per second, which translates into something around 15% of its total computational time. The landing software had only been designed with a margin of 10% or so.

The 1202s were also a lot less benign than is often reported. They occurred because of the fixed two-second guidance cycle in the landing software. That is, once every two seconds, a job called the SERVICER would start. SERVICER had many tasks during the landing. In order: navigation, guidance, commanding throttle, commanding attitude, and updating displays. With an excessive load as caused by the CDU, new SERVICERs were starting before old ones could finish. Eventually there would be two many old SERVICERs hanging around, and when the time came to start a new one, there would be no slots for new jobs available. When this happened, the EXECUTIVE (job scheduler) would issue a 1201 or 1202 alarm and cause a soft restart of the computer. Every job and task was flushed, and the computer started up fresh, resuming from its last checkpoint. It was essentially a full-on crash and restart, rather than a graceful cancellation of a few jobs. And unlike is often said, the computer wasn't dropping low-priority things; it was failing to complete the most critical job of the landing, the SERVICER.

Luckily, the load was light enough that of the SERVICER's duties, the old SERVICER was usually in the final display updating code when it got preempted by a new SERVICER. This caused times in the descent when the display stopped updating entirely, but the flight proceeded mostly as usual. However, with slightly more load, it was fully possible that the SERVICER could have been preempted in the attitude control portion of the code, or worse yet, the throttle control portion. Since each SERVICER shared the same memory location as the last one (since there was only ever supposed to be one running at a time), this could lead to violent attitude or throttle excursions, which would have certainly called for an abort. Luckily, this didn't happen -- and the flight controllers didn't abort the mission not because 1202s were always safe, but because they didn't understand just how bad it could be, were the load just a tiny bit higher.

thewonderidiot··on Apollo Guidance Computer switching power supply works after 50 years
To the contrary! In our AGC, all of the capacitors not tucked into cordwood holes sport a KEMET logo and part number: https://photos.app.goo.gl/oGsrgXBpNMccmkWV7

I think the Computer History Museum's might have Sprague caps in it, though.

thewonderidiot··on Restoring an Apollo Guidance Computer, Part IV
Yeah, it was absolutely incredible to see that clock being generated for the first time. That was the first module we powered up, since it's the easiest -- just give it 4V and 14V power inputs, and it should start clocking. But it was also a scary one, because that's one of the three potted modules in the computer, so trying to fix it would be difficult. And the design review book for the computer has nothing good to say about the durability of one of the components in it.
thewonderidiot··on Restoring an Apollo Guidance Computer, Part IV
Yes, the current plan is to do a bit of an exhibition/tour with it. Details are still murky, since we still have a good bit of work to do to get it fully working.
thewonderidiot··on Restoring an Apollo Guidance Computer, Part IV
Yeah, definitely. We took a ton of pictures and video. I'm not sure when they will be posted, since I'm sure there's a lot of editing to do.
thewonderidiot··on Restoring an Apollo Guidance Computer, Part IV
Hey all, Mike here! I'll be more than happy to answer any questions about the project, the AGC we're restoring, or anything else!
thewonderidiot··on Tindallgrams
I haven't used or prepared rope memory, but I have done quite a bit of research and planning to potentially construct some in the future.

Using it wasn't terribly exciting; the rope memory for a program was broken up into six rope modules that could be installed into and removed from the back of the AGC pretty easily with a screwdriver. Dedicated rope modules were only really used for flight and for completed test programs. For the most part during development they made use of "core rope simulators", that simulated the electrical properties of a core rope memory, but read data from a traditional coincident current ferrite core stack. This let them much more easily and quickly test programs out on hardware, without going through the pain of shipping a release out to the factories to manufacture.

The long and expensive assembly process caused last-minute changes before mission to be, in some cases, a bit hacky. They would do their best to localize changes to a single module, if possible, so that they would only have to re-manufacture one of them instead of all six. The Apollo 11 LM thus flew with 3 modules of Luminary 97, 2 from Luminary 99, and 1 from Luminary 99 Rev. 1. And this page from the Apollo 5 software Sunburst 120 shows how messy that could get: https://archive.org/stream/yulsystemforagcr00nasa#page/n485/...

Rope memory led to one of the most interesting "binary" output formats from an assembler that I've ever come across. The assembler (https://archive.org/details/yulsystemsourcec00hugh), upon successfully assembling a program, could punch a paper tape for manufacturing. The tape didn't contain the words of the program, directly. Rather, it contained commands for a special machine that was designed to help with the construction of the rope modules. The machine would position a small loop in front of the next core a wire was supposed to be threaded through. One of the two operators would pass a needle and wire through the loop and through the core to the operator on the other side, and the loop would then be repositioned in front of the next core the wire was to pass through. This video describes the process in more detail and shows it being done: https://www.youtube.com/watch?v=YIBhPsyYCiM

Preparing it nowadays is a bit challenging. Contrary to most descriptions you'll find online, rope memory is not "just" simple transformer coupling between the drive lines and the sense lines. The operation is a lot closer to the concept of coincident current ferrite core memory. Just like regular ferrite core memory, rope memory relies on the switching action of cores with "square" hysteresis loops. When a particular word is being read out of memory, a single core is addressed (each core stores the data for twelve 16-bit words). This core is "set", or driven to one magnetic polarization, by a set current, with all other cores being held by a series of inhibit lines. After a short time, the core is then "reset" to its original polarization by flowing current through the other direction. The changing magnetic field of the core couples into the sense lines that go through the core (and not onto those that don't), which are run into a traditional sense amplifier circuit.

Anyways, the cores used were (as far as I've been able to gather) metal tape cores, wound with 1/8mil thick 4-79 molybdenum permalloy tape. Similar stuff is still made today, but it's not super easy to come across. Each of the six modules contains 512 of them, so you'd need 3072 cores total if you wanted to weave one of the full manned flight programs.

thewonderidiot··on Some Funny Things Happened on the Way to the Moon (2002)
Indeed! He has a fun anecdote on his website [1] about the throttle control routines, which he was responsible for. His account of it is worth reading, but long story short: The ICD for the descent engine specified that its throttle response time was 0.3 seconds, and so that's what the simulator hardware implemented. Don discovered, through testing, that he only needed to code in 0.2 seconds to get a clean throttle control, and so he opted to not give it any more compensation than needed.

At some point, though, the throttle response time for the descent engine was improved from 0.3 seconds to 0.075 seconds, but the ICD was not updated. So at least Apollo 10-13* all flew with 0.2 seconds of compensation. As it turns outs, the throttle was just barely stable with this much, and if Don had implemented the specified 0.3 second compensation it wouldn't have worked.

Apollo 15-17 (and likely Apollo 14) all had this issue corrected. You can see the difference in the constant THROTLAG comparing the Apollo 13 [2] and Apollo 15-17 [3] lunar module source.

* We haven't yet found a copy of the Apollo 9 or Apollo 14 flight software, so I can't definitively say that either had the error. We do have Apollo 5 source, but THROTLAG, the constant in question, does not have the same name there, and I'm not sure what the equivalent value is.

[1] http://www.doneyles.com/LM/Tales.html (towards the bottom)

[2] https://github.com/virtualagc/virtualagc/blob/master/Luminar...

[3] https://github.com/virtualagc/virtualagc/blob/master/Luminar...

thewonderidiot··on Some Funny Things Happened on the Way to the Moon (2002)
Yep, that's right -- Don was the guy who came up with that procedure.
thewonderidiot··on Some Funny Things Happened on the Way to the Moon (2002)
If I recall correctly, over the course of the program some 300-400 people wrote code for the AGC, so it's not surprising that not everybody gets mentioned. Don Eyles (who saved Apollo 14) has an MIT org chart from the time of Apollo 11 on his website: http://www.doneyles.com/LM/ORG/index.html That only includes people at MIT that worked on it -- there were also many at Raytheon, AC Electronics, and various other contractors (like Adams Associates) that we don't know the names of.

The code that kept Apollo 11 going was Hal Laning's, and he's mentioned a lot in the article.

thewonderidiot··on Source code for the Apollo 11 Guidance Computer
They do! They're in EXECUTIVE.agc (search for OCT 1201 and OCT 1202).

Those error codes mean "No VAC Areas" and "No Core Sets", respectively. Core sets were the basic task control blocks, including each task's entry point, priority, flags, some memory for temporary variables, and a few other things. VAC (Vector Accumulator) areas were a bit more interesting. Most of the real guidance code was not actually written in AGC assembly because it was so primitive and limiting. So, they created the "Interpreter" (INTERPRETER.agc) that's essentially a little virtual machine, that had its own assembly language (complete with mathematical and vector operations). Interpreted tasks needed more than the 7 words of temporary variables provided by the core sets, so they also allocated these VAC areas for more storage.

Those error codes showed up on Apollo 11 because of a weird electrical power phasing bug, essentially causing the radar to generate thousands of "interrupts" (actually cycle stealing operations) every second. With all of this additional work, the AGC didn't have enough time to finish its low priority tasks. And since those tasks hadn't exited by the time they were expected to, when the executive attempted to kick off new tasks, it found that no core sets or VAC areas were available, and sounded those program alarms.