How to find a street in 2 minutes [video]
youtube.com
youtube.com
A short & catchy variation of "Overpass turbo - a super powerful mining tool for OSM" would be a more appropriate title [0][1]. It exists since 2013 but got a massive boost from the Pokémon Go craze in 2016 [2].
[0]https://wiki.openstreetmap.org/wiki/Overpass_turbo
[1]https://wiki.openstreetmap.org/wiki/Overpass_API
[2]https://www.itechpost.com/articles/34612/20160930/pokemon-go...
i.e. "banks in my area" a la google maps here: https://i.imgur.com/pSRn1xo.png
In my experience it seems like prompting for older technologies works much better on large LLMs, i guess it makes sense considering that there is probably more crawlable documentation out there
EDIT: someone in this thread linked a tool which seems to do this: https://github.com/rowheat02/osm-gpt
There was also a great talk about Overpass Turbo and related OSINT tools on the 2021 CCC Congress, also showing examples outside the US.
The talk is available in english as well: https://media.ccc.de/v/rc3-2021-r3s-112-osint-ich-wei-wo-dei...
Openstreetmap data is used in Overpass Turbo -> ask ChatGPT NLP into Overpass Turbo scripted API -> copy paste script into Overpass Turbo -> select Location data from query -> copy/paste coords into Google Street View images -> Matching the original image by eye.
This part of the his search could be collapsed into a single tool. Or, if that is maybe too much of a specific use case, imagine moreso a 'workbench' tool that would put all this in a single place. Every time he switches chrome tabs, or copy/pastes... that could be taking place in a single environment, like a GPT4-plugin for example.
We are of course very good at object recognition from still images. We are reasonably good at inferring depth from a single still image (monocular depth estimation [0-2]).
Combining these, we should be reasonably good at determining relative positioning of identified objects in 3D, and thus computing their overhead 2D positions. Imagine discretizing overhead x/y positions, essentially yielding an image where each pixel corresponds to a 1 meter square looking down, whose color is the inferred identity of what occupies it.
For example, in the linked album cover, we could infer that a street occupies a set of pixels with coordinates S = {(x1, y1),...(xN, yN)}; a building occupies pixel set B1 = {(xb1, yb1),...(xbM, ybM)}; another building pixel set B2; etc.
Now, we are also reasonably good at object recognition (and MDE) for overhead imagery [3]. This would let us build a giant set of overhead object identities inferred from the entirety of satellite imagery. The tricky part would be efficiently indexing this, such that the overhead coordinate sets inferred from the ground-level image could be quickly (and fuzzily) queried in the satellite data. A brute-force approach would essentially convolve the overhead object coordinates inferred from the ground-level image over every region of the satellite imagery, at every possible angular orientation (and likely several different scales, since MDE is often imperfect at inferring absolute depth), but this would be impractically slow.
I would bet that intelligence agencies have had this capability for some time now.
[0] https://www.sciencedirect.com/science/article/pii/S095070512...
[1] https://www.nature.com/articles/s41598-022-20909-x
>The model can['t] pinpoint exactly where a street-level photo was taken; it can instead reliably figure out the country, and make a good guess, within 15 miles of the correct location, a lot of the time
I could imagine that a more qualitative model like this one could be used to restrict the search space of a more quantitative model akin to the one I (very roughly) proposed.
Also, picture has to contain all these clues of course, conveniently, including a couple of street numbers.
Knowing exactly what they look like would help, but the point is more that you're supposed to look them up and see if they match to rule out regions or confirm suspicions. You've got the wealth of human knowledge at your fingertips, use it - if knowing what a Kentucky numberplate looks like might be useful, you can just look that up in a matter of seconds.
> Also, picture has to contain all these clues of course, conveniently, including a couple of street numbers.
Given how many times this individual has narrowed down an exact location from a photo of a straight road with nothing but trees on either side, I think it's fair to suggest that he was picking an easy target for the purposes of tutorialisation.
It might take more than a couple of minutes to narrow down a photo with less obvious clues than this one, but it's fairly difficult to take a photograph outdoors that doesn't give away enough for someone with sufficient knowledge to identify exactly where it was taken.
The ability to have ChatGTP generate the query is actually pretty huge.
I wonder how long until there are AI tools where you can just say things like "give me skate parks within 100 meters of dog parks in Boulder CO", and it will identify the tool to use and generate the query and execute the query and just give you the answers.
It’s amazing plus it’s probably used in several tools/apps/services you use regularly.
And best of all, if you have a look at the map around someplace you’re familiar with there are probably several improvements and additions you can make this very day using their online map editor.
He made a funny video about taking the “test” on the CIA’s website which was an online game for kids (a really weird thing to exist when I stop and think about it).
https://www.worldstandards.eu/electricity/plugs-and-sockets/
PIGEON is going that way I guess: https://arxiv.org/abs/2307.05845
It's also not as if most other video sites provide an open api that doesn't need to be scraped
Years of use, rarely isn't perfect condition.
I really think those tools should be banned, they're such an immense security problem.
I know those can be useful, and I understand that information should be free, but they can still do quite a good amount of harm.
Lesson: ALWAYS blur anything that could identify a street or anything when you post a picture online.
Legitimate users no longer have access to it, Bad actors continue to use the software on their own hardware anyways.
I think your point is valid to consider but it uncharitably overlooks the positives of the other side's proposition and reduces it down to all or nothing.
It’s a set of basic techniques to find where a photo was taken. I have a hard time envisioning an actual new issue here.
Obviously, people shouldn’t post pictures of things they don’t want other people to see or know (like where you live if you don’t want people to know that) but that was already true before people could easily mine Openstreet map data.
Those things could be used by robbers or any other kind of people who browse photos online.
Now, of course it depends on people handling their data straight, but since the internet is mostly used as a place where people share anything to the public, I still believe it's a bad combination of problems waiting to happen.
I think I’m lost here because what you are describing is exactly the meaning of mining openstreet map data. That’s the GIS database they use.
Austria for example has very good maps in OSM, South Korea on the other hand has much lower quality maps (probably fewer people there using OSM)
https://jsis.washington.edu/news/south-korean-data-localizat...
> The purpose of the Korean Land Survey Act was to prevent anyone who is threat to national security from stealing the country’s maps during the post-war period. ROK government enacted the law in 1961 and it concerned the establishment and management of spatial data — it has subsequently been amended in 2009. The Article 16 of the act states strict regulation on taking maps, photos, the results of a survey, or any land surveillance data abroad, because of the likelihood that it could harm South Korean national security interests.
There's certainly regions of the world where doing this would be much more challenging (e.g., Central Africa, China, rural India), but the stuff he covered in the video is going to be extremely helpful in the vast majority of cases.
Exactly the same thing that Rainbolt did in this video: cut down the amount of work you have to do on each street from checking dozens or potentially hundreds of photos/angles to just 1-3.
Of course if you've only narrowed the streets down to 20,000 candidates instead of 20 of them, that doesn't get you straight to the answer, but it's still a massive proportional improvement.
But the lesson being presented here is to use data that's available to you in the photograph. Maybe you don't have any street numbers (or any particularly useful ones) visible, but you can see the sun at the end of the road and therefore know that it's running directly East-West. That filters out tons of roads. Maybe you can see that houses are only on one side and a river is on the other, you can use that as well. In the video he mentions similar constraints with regards to local parks as being other options for this kind of search narrowing.
The point of that portion of the video isn't "hope the house number is really weird lmao", but to extract any geographical information out of the image and then query that with open mapping databases. House numbers are just one of the most common ones, and while they're typically not quite as powerful as was shown in this video, they're very often going to help dramatically.