Oh I doubt that will be necessary. I bet you could find a useful "feature vector" by repeating the process on about 10-20 new voices - then extract a set of words (let's pessimistically say 100) that activate these features the most - then have a person retrain with 5 examples of each of these 100 words (an hour, at the most).
I honestly wouldn't be surprised if you could find a really feature rich sentence (say 5 words long) that you could use to crack pretty much all voice-activated biometric password systems.