Researchers use AI to turn sound recordings into street images
news.utexas.edu
news.utexas.edu
So it doesn't really, as the title claims, turn recordings into images (it already has the images) and the distorted fake images it creates are only "accurate" in that they broadly slot into the right category in terms of urban/rural setting, amount of greenery and amount of sky shown.
It sounds like the matching is the useful part and the "generative" part is just a huge disadvantage. The paper doesn't seem to say if the LLM is any better than other types of models at the matching part.
The generated images are only vaguely similar in detail to the originals, but the fact that they can estimate the macro structure from audio alone is surprising. I wonder if there's some kind of leakage between the training and test data, e.g. sampling frames from the same videos, because the idea you could get time of day right (dusk in a city) just from audio seems improbable.
EDIT: also minor correction, it's not an LLM it's a diffusion model. EDIT2: my mistake, there is an LLM too!
It is still decent way to start I think, but it needs to get more varied data after that and use different geographical locations for eval and test.
[edit] The bottom right image is even more suspect. There's a vertical green sign in the same place on the right side of the image, but also some curious red striping in the distance in both images. One could argue 'street signs are green' but the red striping seems pretty unique, and not something where one would just guess the right color.
Time of day seems almost easy. Are there animals noises? Those won’t sound the same all day. And traffic too. Even things like the sound of wind may generally be different in the morning vs night.
This is not to suggest the researchers are not leaking data, or that the examples were cherry picked, it seems probable they are doing one or the other. But it is to say, if you were trained on a particular intersection, and heard a sample from it, you could probably train a model to predict time of day reasonably well.
Here's how the results were scored:
"Computer evaluations compared the relative proportions of greenery, building and sky between source and generated images, whereas human judges were asked to correctly match one of three generated images to an audio sample."
So this is very impressive and a cool piece of research, but unsurprisingly not recreating the space "accurately" if you assume that means anything more than "has the right amount of sky and buildings and greenery".
This was established mathematically, answering an old 1966 question from famous mathematician Mark Kac: "You can't hear the shape of a drum" -- there isn't a unique answer even when allowed to use arbitrary test sounds.
Wikipedia: https://en.wikipedia.org/wiki/Hearing_the_shape_of_a_drum
Article in American Scientist 1996 Jan-Feb: https://www2.math.upenn.edu/~kazdan/425S11/Drum-Gordon-Webb....
Proof of concept: echolocation.
Just hearing someone hit a drum wouldn’t give you the shape of it.
https://news.utexas.edu/wp-content/uploads/2024/11/AI-street... https://news.utexas.edu/wp-content/uploads/2024/11/AI-street...
I looked up if someone is already doing this, and found this tool: https://www.youtube.com/watch?v=_xgZeJlL3RY
I would not rely on this tool for any meaningful data collection.
Beyond that, you are correct that the 3D shapes themselves cannot be derived perfectly accurately (see my other post)
There are going to be real useful tools - but we need to play for another century before we have that aha moment. Probably :-)
What I’m saying is that if you were to replace ‘AI’ with “ask humans to draw an image based on these sounds,” you’ll probably get somewhat similar results.
Which is still interesting either way.
In this particular case, it is not.
tldr:
you can view the image directly at https://news.utexas.edu/wp-content/uploads/2024/11/AI-street...
Still not overly useful.