MP3 for Image Compression (2006)
keyj.emphy.de
keyj.emphy.de
What I did instead was to run images through an audio editing tool, which lets you apply echoes or do mindboggling things like change the volume of the image. The script can be found on github[2].
I haven't even tried the wahwah effect, you've definitely piqued my interest.
I had a look through the SoX effects list[1] for the wahwah, but can't seem to find it. The closest thing I found was the 'flanger' effect[2].
Also like how clearly you can see the overdrive introducing noise into the previously pure blue sky at the top
In any case, JPEG murders pretty much everything in thag representation.
Heck, I've been 17 when I did that. I certainly didn't really know much about the theory behind or why it almost certainly would fail. But I guess by now I wouldn't even attempt such craziness, so in a way it's probably a good thing to not know things too well, sometimes.
sox mo.wav -e unsigned -b 8 -c 1 -r 48k mo.raw
bytes=`stat -f %z mo.raw`
width=`echo sqrt\($bytes\) | bc`
square_bytes=`echo $width \* $width | bc`
dd if=mo.raw of=mo_square.raw bs=$square_bytes count=1
gm convert -depth 8 -size ${width}x${width} gray:mo_square.raw -quality 50 mo_square.jpg
gm convert mo_square.jpg gray:mo_square_jpg.raw
sox -e unsigned -b 8 -c 1 -r 48k -t raw mo_square_jpg.raw mo_jpg.wavJust tried converting the sound to 8 bit, compressing it as a monochrome image, then back.
You can't really hear much difference at higher quality settings, and as you turn it down it sound like there is more and more digital-ish noise. I had to turn the compression way up before the voice became almost illegible, but you can still hear that it's someone talking.
Seriously, when it's unpleasant to listen to. As i said in the other reply, with music the quality is pretty bad on almost any setting, voice just happen to be an exception.
Even of high quality, there is a background sheesh kind of noise, similar to what you get if you play a 3D printed record.
On higher compression it's turning into something rather crumbly, with a recognizable tune but horrible sound quality.
So, yeah. Not quite usable.
I guess I'm behind the times, not relating to that reference...
So what happens when you encode visual information into sound in such a way as to have a sound-to-environment signal and you just keep that active and running real time. We already know that this mechanism exists in nature and can even be developed in humans, since echo-location is a thing.
Edit: It occurs to me that actually doing sensory experiments like this might really mess with someones head... maybe not a good idea to play around with it casually.
Edit 2: If we were going to swap out one of our senses for another, which should we get rid of and which should we gain?
[0] http://gizmodo.com/these-synthesia-glasses-help-blind-people...
[1] https://www.ted.com/talks/wanda_diaz_merced_how_a_blind_astr...
https://play.google.com/store/apps/details?id=vOICe.vOICe&hl...
I think the 2.00 bits/ pixel result looks quite more "analog" with a film grain effect to me.
https://hackaday.com/2017/04/16/mangling-images-with-audio-e...
Thanks for that.
There are also some results of experiments with the Opus codec posted in the comments.
https://www.tablix.org/~avian/blog/archives/2006/01/lossy_co...
I think this alone would be an improvement for the mp3 encoder.
Also, today I stopped halfway through importing a FLAC of repetitive electronic music into Audacity, and I was surprised to discover that there was almost no repetition (it would play one bar then skip ahead over the identical parts)
Edit: Whoops, totally wrong!
Turns out I was loading a half-torrented file and forgot that torrents don't download in order!
https://en.wikipedia.org/wiki/JPEG#/media/File:JPEG_ZigZag.s...
It's a shame the author didn't do the same transformation, because it would de-correlate a lot of the error noise. You can see in the highest compression settings that the "MP3" image compression is smearing everything horizontally. If it used a zigzag transformation, it would be a more smeared both horizontally and vertically, but probably less visually bad.
If you're asking why they use the zigzag (within a block) instead of a Hilbert curve, IIRC (quite fuzzy on this, so take it w/ a gran of salt and verify) the reason is that it allows for better spatial encoding (imagine having a ripple in one corner and going out - that's essentially what you want to encode w/ your DCT). Using a Hilbert curve would preserve locality, but I don't think it would line up with the spatial distribution of frequencies in an image.
No. The zigzag ordering is applied to the frequency components, not to the image pixels.
1. Divide image into blocks (as in JPEG),
2. Perform two-dimensional FFTs (as in JPEG),
3. Scan frequency components in zig-zag order (as in JPEG),
4. Run all of the steps of MP3 compression aside from the initial "split audio into blocks" and "perform FFTs" stages.
That would pretty much just give you a less efficient version of JPEG; both JPEG and MP3 take advantage of knowing how much each frequency component "matters" (i.e., how precisely it's necessary to encode the value to avoid artifacts noticed by humans), so using the MP3 quantization logic on frequency amplitudes from images would result in wasting bits by encoding certain amplitudes more precisely than is useful.
https://www.youtube.com/watch?v=Tr4zb-HHZs4
and
https://www.youtube.com/watch?v=Ee5evlN8Bbs
and if you go down the rabbit hole, you end up here:
https://github.com/skratchdot/audio-generation-loss/tree/mas...
So, mp3s add a bunch of silence to the beginning of the file, and ogg files start to "chirp". I never got around to putting this info in a consumable, easy to understand format though. The videos in these folders just continuously re-encode a source file w/ a given lossy format.
Its no surprise it works, however you wouldn't necessarily get "as good" compression as you would from an optimized DCT coder (JPEG etc) based on the data duplication (2x for the overlapping blocks).
See https://en.wikipedia.org/wiki/Modified_discrete_cosine_trans... https://en.wikipedia.org/wiki/Discrete_cosine_transform
I'd like to store music files in my phone's camera roll, and easily upload them to a website where some Javascript could decode and play them.
That's something your Turbo Pascal code never attempted.
[1]: I remember when I was taking an information theory lesson at the university and told the (video/jpeg compression guru) professor that I compressed an AVI file with my dump arithmetic coding implementation, he was shocked, turns out the file had some large crappy header from the editor
ZIP consists of sections of both compressed and uncompressed data, so some compressibility is to be expected.
Needless to say, a perceptual model designed for audio is a pretty bad choice for a long string of grayscale pixels, interpreted as sampled audio. It looks like a lot of high-frequency content was discarded, resulting in horizontal blurring.