It already uses a harness (lightweight viewer) to generate screenshots with. The agent already sometimes compares it with the aerial layer, but I didn't gave it much thought. Currently, the most difficult things to get right are tunnels and bridges. Comparing it to real location imagery would probably help. I'll give it a try. Thanks for the suggestion!