Unfortunately trying to construct street level from this technique would be difficult as people would have to manually geolocate to an extent, which would generally grow the older the photos are. But I think sites like Mapillary and OpenStreetCam don't have a required freshness of photos, though I've found by accident Mapillary has a minimum at 1 epoch time.
http://www.freep.com/story/entertainment/movies/2017/02/13/d...
Looks like the NYC one has more coverage though.
During WWII the UK government (maybe the US? I'm unable to find a reference but the search terms are pretty generic) requested pictures and postcards from people who'd taken recent trips into Europe. They wanted them for intelligence reasons - I think it was the UK.
That's the closest thing I can think of, where there's been a request for massive photos. They wanted people to include information about locations and dates, as I recall.
Someone here might be a historian and know more details than I do.
But, it seems like it'd be a large pile of data and work. Unless they knew when and where they came from, I'm not sure how much value they'd have - other than aesthetics.
You're right, its a huge pile of work. Without the geotag and temporal EXIF data, you're relying on machine vision/structure from motion (think https://mapillary.com) to rebuild 3D scenes solely from image data.
https://help.mapillary.com/hc/en-us/articles/115001770329-Ma...
"Mapillary uses a technology called Structure from Motion (SfM) to create and reconstruct places in 3D. By matching points between different images, SfM is able to locate the point in a three-dimensional space and therefore determine its location on the map. The more images there are available for a specific point, the more accurately it can be reconstructed. In addition, SfM creates the kind of smooth transitions between images that you can see in the Mapillary viewer—again, provided that images have been taken within close proximity with enough overlap between them.
We also run semantic segmentation on the images. This means that the computer tries to understand what is in the image and assigns a category tag to each pixel. That enables us to detect different areas in the images (such as buildings, pedestrians, cars etc.). Semantic segmentation together with 3D reconstruction enables us to extract 3D positions of objects such as traffic signs and display them on the map. You can get more detailed information on what is currently available on our product page."
Disclaimer: No relation, just dig their tech (they're using machine vision to build world models for self-driving cars, using crowd sourced imagery).
But... If it had some initial training done by hand, and by people including what dates and locations they know, maybe it could work backwards? It'd still end up being a HUGE amount of data and the compute power for that would probably be obscene.
Which, of course, means I think it's an excellent idea and that someone should do this. I can't even begin to imagine the costs. This might just be the best idea ever, or the worst.
It does tie into a strange idea that I've been tossing about for the past decade.
A cluster of used cell phones. They're easy to power, probably free for the taking, and might make for big computing while just using something that'd normally be thrown away.
My original idea was SETI or one of the unfolding@home type projects, but something like this might make for a good project.
Alas, I'm way too lazy (and not qualified) to do this. Laziness is one of the perks of retirement.
I'd say a few dozen are of the places they were living.
The ones that show recognizable places are pretty interesting to look at though.