I can understand that this works for cats or cars, but how does this work for images that are not in any training set yet? E.g. highly specialized images like x-ray pictures, or astronomy pictures?
Also for the medical domain, I think vision-text segmentation models like SEEM (https://github.com/UX-Decoder/Segment-Everything-Everywhere-...) are really cool. You could for example ask “Where is the tumor located on that image?” and then the tumor is highlighted in the picture.