Thanks for the feedback. We use captions for our prototype in order to extract all the data which is just one way to do it. They are many ways to extract content from videos and we're working on some interesting ones!
About youpronounce.it (watch out for trademark complaints), they focus on a single use case? pronunciation? Which is fine.
We just built a quick prototype for validation. We're working on scan.video to make video content more accessible for people on the Internet. What you see today is %0.1 of what we'd like to build. So, stay tuned! definitely..
(We added a link for 1st time users to make things easier)
Thanks again for all the feedback!