at least for vision? Are there similar comparisons for language? It seems like vision is often an afterthought when it comes to bootstrapping these newer inference engines
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.