Based on the recording, the information I could find, and imagining how I'd try to do the same thing using the technology of the era, I assume the melody is based on single-cycle samples of a piano and Max Matthews playing the violin. The vocals sound like formant synthesis like the Votrax SC-01 or TI LPC series, although of course those chips didn't exist until 15+ years after the work at IBM. But I'm very curious about the details. Did the team develop a general-purpose sequencer for the melody and/or speech, or were all of the notes, slides, etc. hardcoded? Did the computer actually output all 3+ parts together, or were they separate elements mixed after the fact? I assume the output was not realtime, but it would be a neat surprise if they achieved that in the 60s. Was it all handled digitally in the computer, or was the computer controlling some add-on hardware, maybe with analogue filters? Etc.