Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
microsoft.com
microsoft.com
It seems like a lot of the most exciting research is happening here - making unbelievable progress with such small parameter sizes.
Am so much more excited about tiny models gaining real intelligence. Just today I have been running Qwen3.5 0.8B model on images and am pleasantly surprised by how good it is compared to even 4B and 8B models from a few months ago.