yeah there has definitely been work done in that space: it’s called multi-modal models
not sure if this is the latest work but here’s some results from Google’s AI Blog
https://ai.googleblog.com/2017/06/multimodel-multi-task-mach...
not sure if this is the latest work but here’s some results from Google’s AI Blog
https://ai.googleblog.com/2017/06/multimodel-multi-task-mach...
Like imagine the vision part making a phonecall to the natural language part to ask it for help with something.