Meta's new Video Understanding Multimodal Model used Qwen model for trainingarxiv.org7 points·BUFU··1 commentOpen articleSaveView on HN