The Wakaru app does allow ebooks reading with a dictionary, but I always found it clunky and I don't think it allows easily generating cards from a book, the dictionary was less useful than yours looks, and it doesn't learn your vocab over time.
For audiovisual context, I'm really thinking of:
- images to match vocab items (potentially auto-sourced from a Creative Commons licensed image search engine)
- images from articles to match sentences from articles
- sound clips to match text taken from audio captions
- video clips OR audio clips + screenshots to match sentences from video subtitles
Regarding image rights, I would say that there is a huge amount of stuff in the public domain or liberally licensed (e.g. on Mediawiki). But you could also allow the user to add their own images, where they take responsibility for making sure they have the rights.