In contrast, image data on the intent for image generation models is very highly annotated in most cases.
Another potential source of data is voice acting script of animations. I always thought the storyboards of films/animations can be great annotated training data but it seems there are no open datasets, probably because of copyright issues.
It also does not account for where stresses, emphasis, pauses, etc. are placed to enhance the delivery of a given text.
How do you get sentiment analysis to properly annotate an audiobook that has a dramatic reading, or something akin to the narration of the Game of Thrones or Harry Potter books where the narrators switch characters, accents, manarisms to portray the written content?