> We then used the location of the images as the output data and the screenshot of the webpage as the input data. And now we have exactly what we need — a source image and coordinates of where all the sub-images are to train this AI model."
I don't quite understand this part. How does this lead to a model that can generate code from a UI?