Looks awesome!
Is there anything you can share about the architecture or pipeline you used for it? A high-level overview would be enough.
I’m guessing you’re doing video-to-image, image-to-text, and then text-to-docs, right? Since not all of the models you mentioned are multimodal.