The problem seems fairly tacklable . Learning what is on a display screen is relatively easier than most computer vision problem spaces. There are many repetitive patterns in typical application UX.
For example let say there is a label for Save Icon that is an image (a Floppy Disk in most apps) and not alt tagged. By visually reading the image of the screen the model should not have to much difficulty in tagging it that as Save button ?
Most consumer / biz app UX do follow many standard conventions if only out of convenience and lack of imagination, so building a learning algorithm around these components should be possible ?