Offtopic sort of, but does anyone know if folks are working on combining vision and natural language in one model? I think that could wield some interesting results.
The results are quite interesting:
https://www.reddit.com/r/Art/comments/p866wv/deep_dive_meai_...
not sure if this is the latest work but here’s some results from Google’s AI Blog
https://ai.googleblog.com/2017/06/multimodel-multi-task-mach...
Like imagine the vision part making a phonecall to the natural language part to ask it for help with something.