so if i understand this correctly — you want the speech recognition model to identify a vocabulary of specific terms that it wasn't trained on. instead of fine-tuning with training data that includes the new vocabulary, you input the full vocabulary at test time as a list of words and the model is able to generate transcripts that include words from the vocabulary.
seems like it could be very useful but it really comes down to the specifics.
you can prompt whisper with context — how does this compare?
how large of a vocabulary can it work with? if it's a few dozen words it's only gonna help for niche use cases. if it can handle 100s-1000s with good performance that could completely replace fine-tuning for many uses