Loved it! But I think with advent of llm, there should be a way to somehow feed the current view and generate a realistic 3D view, based on the real data from the same location to be as accurate as possible.
It already uses a harness (lightweight viewer) to generate screenshots with. The agent already sometimes compares it with the aerial layer, but I didn't gave it much thought. Currently, the most difficult things to get right are tunnels and bridges. Comparing it to real location imagery would probably help. I'll give it a try. Thanks for the suggestion!