Using AI to compress audio files for quick and easy sharing
ai.facebook.com
ai.facebook.com
It did a great job of reducing FLAC size from 100mb to less than 1mb using stereo 24kbps preset but audio quality suffers a lot in some places. Maybe training it with 48kbps or 64kbps would make it a feasible alternative for storing music without much quality loss.
In comparison, lame insane preset (320kbps) produces 12mb mp3 with almost indistinguishable quality from flac.
For those who want to listen to sample: https://a.pomf.cat/pcjynr.wav (first flac then encodec)
The sound of the AI encoder is distinct and probably not suited for music at that bitrate, but would probably serve fine for Facebook videos and podcasts. I'd be really interested in seeing a model optimized to compress human speech...
https://www.mentalfloss.com/article/19727/how-toms-diner-tun...
EDIT: I downloaded the model from the code (https://github.com/facebookresearch/encodec/blob/3837f0db2a3...). It looks to have 15M parameters in total (this is a CPU-only model, by the way). For inference, I assume we can get away with 16-bit floats, so this clocks in at about 30 MiB. Not particularly favorable compared with LAME on my system (164 KiB), but not at all unreasonable to use - these days apps take 30 MiB dependencies all the time.
The tiny music clip in the sample, encoded at 6 kbps, is obviously not any kind of evidence for "CD quality" one way or the other. (The clip itself, if you download it from the page, is re-encoded with 64 kbps AAC.) No way to know how it would stack up against 96 kbps Opus on a stack of CDs with blind testing, I don't think.
Using a generative-adversarial network as a compressor (with a 'perceptual' discriminator, how ever that works) probably means that all introduced artefacts will sound entirely plausible. I.e., it's possible that you didn't mishear the word, but that it was mis-compressed to a perceptually very similar word.
Xerox copier flaw changes numbers in scanned docs [2013]
I have more fun trying to break stable diffusion than getting interesting pictures anyhow.
"CD quality" is 44.1khz over 16bit.
> We achieve an approximate 10x compression rate compared with MP3 at 64 kbps, without a loss of quality.
but then later:
> The key to lossy compression is to identify changes that will not be perceivable by humans
On your second point: they probably meant that there is no loss in perceived quality, of course it's a lossy algorithm.
6kbps sound demo is really impressive! Initially I was turned off by aliasing artifacts (hiss) but to be fair 6kbps is a really really low bitrate.
Between this effort and the recently announced Google-led multi-channel/immersive audio codec initiative, I am pretty excited about the future of audio streaming and distribution.
... but only if you are a bat.
Link?
The flute in particular suffers quite a lot and the triangle disappears completely!
Reading this extremely charitably, I think the contrast is supposed to be to speech which is normally uses much smaller sample rates. E.g. AMR-WB samples at 16 kHz. Speex, back in the day, supported up to 32 kHz.
Full quality digitally distributed audio is often sampled at 48 kHz (e.g. Opus - the default audio codec on YouTube and many other sources), so I think "CD quality" is just supposed to emphasize that it's full-band rather than wide band or narrow band.
Most research in that field operates at 16kHz. Enough for most speech applications, not enough for CD quality. In general, CD quality in this field means a high enough sample rate that you can play all audible frequencies.
I understand all of this reads like pointless pedantry, but a) it feels good indulging in it b) certain word combinations (such as "CD quality") have meanings backed by broad consensus, please use them responsibly so readers do not have to overexert their loose interpretation muscle.
If they mean wide band, say wide band, noone will bat an eye.
The entropy coding seems like a lightweight version of AudioLM.
On the other hand, I think AI-informed video compression will be absolutely transformational; it could enable practical transmission and reproduction of high definition holography (reproduction of an entire wavefront, allowing real 3D scenes viewable from multiple angles) that currently requires a small data-center to process.
1.6kbit/sec for speech.
None of which really seems to make any difference in a world of near-limitless bandwith. It just isn't the constraint it once was.
And it doesn't matter what level you operate on, cache and latency is always relevant. Whether it's registers, L1, L2, L3, same-core NUMA RAM, cross-core, SSD, disk controller cache, disk, same-location distribution server, cross-location distribution server, tape archive backup, etc, going up a level of cache is always a lot slower regardless of bandwidth if you're doing a small read.
Does to me - I have more than a terabyte of field-recorded FLACs that I'd like to put up on the net for people to listen to or download.
Compressing them to q1 OGG (and I'm not sure that's good enough quality for some of them) only gets that down to about 20-25% of the size (eg 200408_0403.flac at 540M goes down to 128M) - even if I host them on S3 or B2, it's still going to be costing me a not inconsiderable sum if people actively listen or download.
If this can get them down to 5-10% with usable quality, that makes life a lot easier (but obviously would depend on browser support, etc.)
Currently contemplating doing this using Hugo since there's nothing really workable out there.
[1] I've got a lot of recordings of the local square and there's frequently screaming children, screaming alcoholics, the occasional person having a mental crisis, dogs barking violently, etc.
Or suppose your are trying to stream a realistic 6dof VR view of a motorcycle ride through a city.
Or just want to quickly load a realistic scene as a user goes through a portal, without waiting 20 seconds for GBs of texture and geometry to load.
Or you just want a Mozilla Hubs scene to load quickly.
Tell that to Youtube, Netflix and Facebook - These companies don't have near-limitless bandwith by default, they have to spend many millions to get their networks serving customers.
Or people like me, who can quite easily max out my mobile data-plan with Youtube and high-quality audio.
Until truly unlimited mobile plans without throttling are available everywhere for cheap, that simply isn't true.
Near-limitless bandwidth just means that the world generates more ephemeral trash, like another Youtube video or TikTok bandwagoning on the topic of the day. These will be viewed for maybe a few weeks, and then they'll never come up on anyone's feed again, they'll just lie dormant on a server somewhere.
A laptop can easily compress h. 265 or avi format without requiring intensive hard wire like AI.
Hell, even the slightly outdated combo of LAME for personal audio library and AAC for multi-channel audio in movies is good enough.
> Nb: complaints require an account in good standing at as this is your third strike your account is now suspended.
Near enough is good enough for most people paying-for/selling the tech.
I am permanently pleased to learn one size does not fit all even in gargantuan saas companies.
I'd like to see what happens when you modify the stream. If the representation is really that compact, making small changes should change the output considerably. Treat it like a synthesizer, really. That should tell us something about its usefulness in identifying audio.
Also see OpenAI’s jukebox.
Diffusion currently wont really help with (B) or even (A)... but there is a lot of new experiments going on with (3)... where the decompression using stable-diffusion is providing very interesting results using far less compute... check out JuicyJukebox notebooks (link to come... cant access github colab from my work computer).