I suspect there will be content lost at the bottom and the top of the image depending on the frequency response of the microphone/speaker.
I suspect there will be content lost at the bottom and the top of the image depending on the frequency response of the microphone/speaker.
This uses a linear frequency scale (which is just the nature of Fourier transforms), whereas our ears are sensitive on a log frequency scale. In other words, the information that's most important to our hearing, which is what a mic & speaker will preserve the best, is in the bottom 10% of the image.
A cheap speaker & mic will probably lose a lot of content above about 10 kHz - which is the entire top half of the image. Even though this wouldn't be that huge a difference to our ears, it would sure look bad in the image.
As for background noise, the difference would probably look like the difference here: http://www.sweetwater.com/insync/media/2010/09/RXAdv-e-xlarg... (that's a screenshot of audio restoration software that removes noise, so it's technically doing the opposite process as best it can, but the difference would be similar).
The software can act as a transmitter or receiver. In trasmitter mode, you can provide it a static image or animated gif, which it will convert into audio which plays continuously. In receiver mode, pixivisor listens via the mic or line-in (depending on hardware platform and whats attached) and reconstructs the image from the audio. You can then manipulate the audio however you want.
This demo uses a korg monotron's low pass filter and LFOs to mangle an animated gif of a cat: https://www.youtube.com/watch?t=63&v=g2W1W4fwEkg
Its really interesting to me to see how the audio modulation is represented in the receiver's output.
Edit: depending on the time alignment / phase response of the speakers, you might see the low frequency parts of the image get distorted to the right.
sox -c 1 -r 48000 -b 32 -e float -t raw out.raw -n spectrogramscience is fun