TSAC: Low Bitrate Audio Compression
bellard.org
bellard.org
This is definitely better than some of the others out there. I threw together some comparisons here at 7kb/s for mp3/opus/aac: https://non.io/TSAC-Comparisons
Happy to add other comparisons if others want any.
Overall, it's FAR better at these lower bit rates, but that doesn't mean it's necessarily good. One issue I see off the bat is that volume is fairly inconsistent in the output for TSAC, which makes stereo in particular quite hard to listen to with the volume "flickering" in each channel independently.
Also I don't seem to be able to access your page, so there might be error.
Finally, when doing opus comparison it's good now to denote if it is using Lace or NoLace decoder post processing filters that became available in opus 1.5 (note, this feature need to be enabled at compile time, and defying decode a new API call needs to be made to force higher complexity decoder) . See https://opus-codec.org/demo/opus-1.5/
Interesting, do you have javascript turned off? Can you access this page? https://html.non.io/TSAC-Comparisons/
Also awesome to see comparison to EnCodec , which I think is one of the better ones available : https://ai.honu.io/papers/encodec/samples.html
Also, can you confirm if Opus decode is classical or with Lace or NoLace post processing filters that are available in Opus 1.5?
This is the codec that TSAC extended, so it could be a nice comparison to see. I'd also echo Vocos (from sibling comment), it operates on the same Encodec tokens but generally has better reconstruction quality.
echo -ne "\x90\x90" | dd if=/dev/stdin of=tsac bs=1 seek=23914 conv=notrunc
you can corrupt the compressed files with very interesting results: https://meow.social/@mimir/112238998609778334
the fast mode (you don't have to patch the binary for this one, it seems to not do the CRC check?) and the normal (non-fast) mode sound different, but both quite interesting
- Can't use it in telephony (obvious application for low bitrates); phone handsets and headsets don't have the power to do it in real time.
- Very small files of good quality would be useful in tiny embedded systems that have low flash space: but what systems of that type have the processing power for decoding? Very low storage more or less goes hand in hand with weak processing.
The quality is astonishing for the bit rate, though.
I like that they share their work, it can lead to something some day.
Maybe add this URL to the calendar on today's date in 5 years an go back and reply with the answer :-D
(Later on I was so surprised MP4 didn't replace MP3!)
Edit: Ok I just saw the thing is over 200 MB, might not be feasible for a while.
It would probably have to be optimized a bit further though, both in terms of computing as well as size. Goal would probably be real time encoding on an iPhone SE without breaking too much of a sweat, and a encoder/decoder perhaps less than 200MB?
I am curious how well this does work with a full orchestra music — that’s where encoders usually croak. Give me a sample of the Star Wars theme.
As the music was "avant gardé", it almost fitted the genre.
Or maybe you do want really good quality in order to fingerprint the voices. Vocoder artifacts can give parties plausible deniability (that's not my voice).
They're putting neural accelerators in everything these days, I wouldn't be surprised if they got it to where it could work on a phone, in which case you could do voice over Meshtastic.
I think this codec could be optimized to run relatively efficiently on the various AI accelerator chips modern phones have, which is “kind-of” doing it in hardware.
It's not a use for "everybody", but it might reduce the costs for those people who need this (or make new things viable).
Check out the citations of Encodec (Facebook's open sourced audio codec) for more examples: https://scholar.google.com/scholar?cites=1126914113099467682...
Not now but in another 5 years they will start to and in 10 years all new ones will probably have the power for this. I find it really exciting although it will consume more battery to run that; but if that is less than the radio antenna requires then it might make sense.
I often wonder, how much of the data currently in circulation will be lost at some point? HDD/SSD last a couple of years. Most of the data in the cloud will be copied over, but some will be lost. If you extrapolate to a 1000, 1 000 000 years, how much will remain? Will something survive the civilization collapse? I guess most people don't care, but some will ...
One way to make data mediums last longer is to make them lower density, and for that such super-low bitrate could be useful.
One approach would be to have a layered strategy - simple (but inefficient) encoding for an initial set of data, accompanied by a bootstrap for the next level which would unlock access to a much larger collection of efficiently stored data.
Until quite recently my phone was the faster computer I owned.
What phone cannot decode these?
I wonder how it compares to https://en.wikipedia.org/wiki/Codec2 and related codecs, which go even lower for bitrate.
In a certain sense, maybe they are. Or more accurately, small fragments of samples, and just how to mix them together, is what is transmitted. It reminds me of pre-generated dictionaries with classic LZ compression. If an algorithm is going to work on mostly English text, then it might make sense to include an English dictionary with the algorithm. Brotli does this [Wikipedia]:
> Unlike most general-purpose compression algorithms, Brotli uses a predefined dictionary, roughly 120 KiB in size, in addition to the dynamically populated ("sliding window") dictionary. The predefined dictionary contains over 13000 common words, phrases and other substrings derived from a large corpus of text and HTML documents
$ tar tvzf ~/Downloads/tsac-2024-04-08.tar.gz
drwxrwxr-x bellard/bellard 0 2024-04-08 14:47 tsac-2024-04-08/
-rw-rw-r-- bellard/bellard 3040 2024-04-08 14:47 tsac-2024-04-08/readme.txt
-rwxrwxr-x bellard/bellard 3979504 2024-04-08 14:47 tsac-2024-04-08/libnc_cuda.so
-rwxrwxr-x bellard/bellard 565336 2024-04-08 14:47 tsac-2024-04-08/libnc.so
-rw-rw-r-- bellard/bellard 49639706 2024-04-08 14:47 tsac-2024-04-08/tsac_stereo_q8.bin
-rw-rw-r-- bellard/bellard 85407494 2024-04-08 14:47 tsac-2024-04-08/dac_stereo_q8.bin
-rw-rw-r-- bellard/bellard 49633561 2024-04-08 14:47 tsac-2024-04-08/tsac_mono_q8.bin
-rw-rw-r-- bellard/bellard 31 2024-04-08 14:47 tsac-2024-04-08/Changelog
-rwxrwxr-x bellard/bellard 287536 2024-04-08 14:47 tsac-2024-04-08/tsac
-rw-rw-r-- bellard/bellard 85143422 2024-04-08 14:47 tsac-2024-04-08/dac_mono_q8.bin
So yeah, it's not exactly a compact stand-alone implementation, but on the other hand it does advanced GPU stuff so I guess nobody expected it to ... or perhaps I did, just a little, based on the author's reputation. :)Compression is getting so heavy that soon it isn't possible to perform it on normal hardware. AV1 already proved that, the future audio/video codecs will be even heavier.
Decompression is also getting heavier. Poor mobile devices.
I'm starting to appreciate well written algorithms which don't require massive computing power. JPEG XL is a good example. It has the same compression ratio as AVIF, but requires less processing power.
Makes me happy to see DAC getting built on! Thanks!
Following the docs, `./tsac c myfile.mp3 myfile.tsac` generates a tsac file that's unplayable with mpv. Trying ffmpeg to convert to mp3 didn't work: `ffmpeg -i myfile.tsac compressed.mp3` ("myfile.tsac: Invalid data found when processing input"). Using a wav input file has the same result.
I can use `./tsac d myfile.tsac output.wav` (I don't really want to decompress anything, but worth a try) but then after compressing `output.wav` with `ffmpeg -i output.wav output.mp3`, output.mp3 is the same size as if I hadn't used tsac (of course). If I use ffmpeg with a low bitrate like `-b:a 16k`, I get the usual low-quality gargle rather than the tsac output.
Which is totally fair given their applications, but I always wonder how much improvement they bring in high bitrate scenario. For example, are there codecs that have much better (perceptible) quality than Apple AAC 256kbps (or achieving similar quality at, say, 160kbps?) How much better are AV1 at 10Mbps compared to H265/264 (the improvement of H265 compared to H264 in "transparent" encoding was pretty disappointing IMHO).
Opus achieves ABX transparency at around 128kbps (as in, the threshold where the vast majority of users taking a fidelity test are unable to tell the difference between the opus-encoded and lossless version).
> NOTE:Opus doesn't support 44.1kHz sample rates, so encodes to 48kHz sample rate. As this causes browser playback issues, it has been resampled back to 44.1kHz. This may affect the sound quality, so this test should be taken with caution.
Is very surprising to me, in two ways.
Firstly I knew 44100 is a relic due to historical reasons, but it's still a quite widely used sample rate in audio world. I have no idea Opus does not support it.
Secondly, it seems to imply browser can't playback 48kHz audio properly. I didn't dig the details, but this sounds weird. Just like 44100, 48k is a very common sample rate, I can't imagine browser would have trouble with it (or any arbitrary sample rate, to be honest).
https://github.com/xiph/opus/issues/43
The browser audio limitation is presumably a workaround to some bug or performance limitation that was relevant at some point in history (the site was created in 2014).
How is this possible? Does it use floating point and concurrency?
Cross-platform floating point determinism is seriously difficult. The Rapier physics engine could do it [0] at the expense of disabling simd and multithreading. It also works only on platforms that strictly comply to IEEE 754-2008 which I think that GPUs usually don't qualify (regarding subnormal numbers etc). Another thing that may have issues is fused multiply-add which may give higher precision than doing multiplication and addition separately (I think some platforms don't have FMA in hardware)
For example, it seems that TSAC currently runs on CPUs and nvidia GPUs. Could porting to AMD GPUs affect determinism?
If someone has audio encoding, playback, and/or DSPs experience email me to be invited to our our Discord server so we can take another crack at it! :)
Using a ~300 MB model, on a 1 TB hard drive, at 8 Kb/s, we can store... ~30 years of music.
https://arxiv.org/abs/2206.07307 https://arxiv.org/abs/2210.13827
Separately:
>The Transformer model is evaluated in a deterministic and reproducible way. Hence the result does not depend on the exact GPU or CPU model nor on the number of configured threads.
That's neat. So even though it's "AI-based" its output is guaranteed to be the same for a given input?
https://ieeexplore.ieee.org/document/7075313
All the patents around that are long-dead so good time to do an updated version I guess.
If you wanted to do something similar but with way lower bitrates (e.g. 300bps), then look at the NRV codec:
https://www.researchgate.net/publication/224209493_300_bps_n...
Are we almost converting music to MIDI at this point?
As I understand it the model is learning the landscape of sound combinations that are interesting to humans and as such there will be no combination of raw bytes in the recorded file that will result in white noise (for example) being heard because this is never trained for.
What if it was though?