The problem with implementing this idea is which technology would actually have to be built, and in which technology the real value would lie. Once the technology is established, scanning videos, tv, radio etc for any source of spoken audio and building up a database of indexable dialogs would almost be the easy part.
Building a speech reconizer is not only difficult, it has also been attempted many times before and unless a speech recognition guru could bring something new to the world, the best we could do is what is already available - so probably best to use existing technology, which often is not cheap to get a license. This is also true with voice print technologies.
The key to getting this up and running lies in finding or building a really good speech recognizer and voice print generator/varifier...
Maybe this is something Y Combinator would be interested in funding? I am based in Europe (Spain at the moment) and I think it would be really hard to convince people to fund this type of technology over here.
If anybody is up for the challenge, I'd love to be involved!