The parsing broke speech into phonemes--actually a string of candidate phonemes, each candidate having an assigned probability. It made a lot of mistakes--it couldn't generally distinguish between "a" and "the" in rapid speech, and the semantic phase didn't help disambiguate those. It worked better for female voices because they have an extra formant. It didn't work well if the speaker was intoxicated--we learned this from some anomalous results that one researcher dug into and discovered that there was a "knee" in the data--it turned out that our late night speaker Bill (a giant bearded guy who wore overalls that he ordered specially from Iowa IIRC and was known as wabblezabble) had taken a break, during which he drank a considerable amount of beer, on the hypothesis/excuse that it would make his speech more, er, fluid. It had the opposite effect--the automated recognition was consistently better before the break than after.
Coming up with the answer required doing a nondeterministic parallel search of the candidate phonemes through a DAG of phrases--the problem was contained because the DAG was highly restricted to the subject matter, in this case facts about Navy ships. This was a pilot and the dream was to have a much more massive semantic net of the English language. We had linguists and a resident lexicographist (he distinguished this from a lexicographer, though the dictionary says they are synonyms--but lexicographists know better than dictionaries created by lexicographers, heh heh) working with us. The parsing code that dealt with the audio signal was written in FORTRAN and assembler, IIRC, but all the language stuff was written in a local version of LISP. Jeff Barnett, on our team, was the author of SDC's LISP2, but I'm not sure that's what we were using. He was working on developing a more performant algolish LISP called CRISP when I left. Jeff had written the parallel search algorithm, which had a "knob", as he called it, which was a floating point value that controlled the depth first/breadth first balance--any possible balance could be achieved by dialing the "knob". This was needed because it took too long to do an exhaustive search--it bailed with an answer as soon as it found one that passed some threshold. Anyway, it required recording onto tape, digitizing it and feeding it to the minicomputer, running many passes, feeding the results into the LISP program running on a mainframe, waiting an indeterminate time to make a match against a highly restricted vocabulary--more of a grind than futuristic. I remember when programs like Dragon Speech showed up ... way advanced over what we had, but still needing to be trained on a specific speaker. Now we have realtime language translation in our pockets. The other day I accidentally turned it on and my friend at the other end of the line asked who was speaking Spanish ... everything I said was being repeated in Spanish.
BTW, when I left SDC because I wanted a break from work, they offered me a spot with their new development called EFTS, but I was pretty set on leaving. EFTS--Electronic Funds Transfer System--is the backbone of all of today's digital money transfers ... ATMs, ACH, etc. I really missed the boat on that one.
P.S. In trying to remember why Bill (aka Billy) also had the nickname wabblezabble, I managed to remember his last name, which yielded his initials WAB (at UCLA initials were used as login names). I found this lovely obit which very much fits the guy I knew: https://www.legacy.com/us/obituaries/latimes/name/william-br...