Multimodal Chain-Of-Thought Reasoning In Language Models
arxiv.org
arxiv.org
Referenced snippet from the abstract:
With Multimodal-CoT, our model under 1 billion parameters outperforms the previous state-of-the-art LLM (GPT-3.5) by 16% (75.17%->91.68%) on the ScienceQA benchmark and even surpasses human performance.
But it doesn't look all that easy to stand up.