QOA, the Quite OK Audio Format
phoboslab.org
phoboslab.org
Minor quibble but, at least for music, most GC games used a native 4-bit ADPCM ("DSP", decoded by the DSP, or the ADP/DTK format, like CD-XA ADPCM and also handled in hardware), and the most common cross-console audio middleware (CRI) also generally used 4-bit ADX ADPCM. In this old list, everything marked ADP, ADX, DSP is using 4-bit (and others are usually just different containers for DSP): http://web.archive.org/web/20080420105759/https://hcs64.com/...
PS1 and PS2 were similar. I've usually only encountered 8-bit DPCM on PC, though the 3DO had a version in hardware.
The old table-driven style of ADPCM might be the poor quality the author has in mind, all of these consoles get much better quality at 4 bits, using the same kind of linear predictor (usually with 2 samples of history) with scales per frame as QOA.
Edit: I hadn't read carefully enough, QOA is doing something more complicated by updating the weights rather than using a fixed set of filters chosen per frame (ADX has only one per sample rate, GC DSP uses 8, XA/PS uses 4 or 5). That seems a little overcomplicated, but maybe needed at 3 bits per sample?
- Original, 4039kb: https://phoboslab.org/files/qoa-samples/adpcm_comp/orig.wav
- MS ADPCM, 1022kb: https://phoboslab.org/files/qoa-samples/adpcm_comp/ms_adpcm....
- QOA, 812kb: https://phoboslab.org/files/qoa-samples/adpcm_comp/qoa.wav
IMA ADPCM is in the same ballpark as MS ADPCM. ADX (not listed) comes close to QOA, but at 1136kb is also larger than those traditional ADPCM flavors.
MS ADPCM sounds horrible, a lot of unpleasant high-pitched noise. QOA is much better but still audibly degraded from the source.
Not sure I would ship content at this quality.
> What makes QOA work is 1) a reasonably good predictor and 2) storing the scalefactor for a bunch of samples explicitly instead of guessing the right one from context, like ADPCM does.
suggests a comparison with DVI/IMA ADPCM [1] or derivatives, which go out of their way not to have too much state or use multiplies, and they also tend to be used at really low bitrates, so they have a somewhat deserved bad reputation. "Guessing" the scalefactor is overstating it, the scale does dynamically adjust but that's all considered from the encoder.
The 2-history-samples style ADPCMs in the BRR [2] family usually have very short frames (at most dozens of samples) and headers specifying scale and predictor/filter index.
I know a lot more about what's used on consoles than about the math of audio encoding, though, so I can't say whether QOA is making the wrong tradeoffs.
[1] https://wiki.multimedia.cx/index.php/IMA_ADPCM [2] https://en.wikipedia.org/wiki/Bit_Rate_Reduction
[1] I ship an Opus encoder and decoder to the browser to support https://www.jefftk.com/p/bucket-brigade-singing and it's just 310KB for compression and 470KB for decompression, which are small enough that I don't even minify them.
(Normally this would be in the commit description but it looks like we forgot to include that https://github.com/jeffkaufman/bucket-brigade/commit/d96b6f3...)
I haven't done any formal benchmarks, but with a simple `time` on the command line QOA encodes 10x faster and decodes 7x faster than Opus.
QOA should be quite suitable for SIMD optimizations, which would improve performance even more. Still on my todo list.
Citation, please.
Opus (at that time it was called CELT) does have more resource requirements in terms of memory than an MP3 decoder. However, I have run the decoder on things as small as a 33MHz ARM7 and still had lots of CPU left over. An MP3 decoder had no hope on that system.
> Citation, please
The parent did provide some data. Admittedly very simple.
Couls you try a different simple benchmark that shows the opposite?
That would be interesting.
TOA is citing about 300kilobits per second which is roughly 30kilobytes per second which is too much data for a 33MHz ARM7 to be able to process let alone do anything to it.
The reason for "The Triangle of Neglect" is that your chips are either under 100Mhz (often significantly as you are on bare metal) and this is too much data or above 1000MHz (you are running Linux) and nobody cares.
ADPCM was more useful back when chips didn't have hardware multipliers.
Similarly, BSWAP / MOVBE might be cheap or free on x86_64 but IIUC RISC-V doesn't guarantee a dedicated instruction for that (and RISC-V is little-endian). "Does [endianness] really matter?" It might, for embedded devices a few years from now. I can't really say without real hardware to get real CPU profiles. But that question is entirely avoidable by just picking little-endian.
Though of course, it's no shock that something simple could be fast. There are also lossless codecs which are much faster to decode than opus.
The games also used DirectAudio or some other horrible mess that relied on registering codecs to the OS, so one of the most common tech support complaints was that the game didn't play any voices (regular SFX were uncompressed WAV), because the user only used media players that bypassed Windows' central codec registry.
Re-encoding everything to WMA (guaranteed to be a registered codec) was suggested several times to management, but since they all used Windows Media Player with properly registered codecs, they never really got the severity of the problem, and didn't care that the fix would cost them nothing (the community had already done all the re-encoding work). I imagine it got changed later, when the games were ported to other OSes, but I had left by that point. (And my NDA expired a long time ago, too.)
So, yes, something like QOA that you can just drop into any C(++) codebase, as games tend to be, without licensing problems or run-time overhead, would have been really, really useful 20 years ago. The file size differences might have been a headache in the CD games era, but by the time DVDs became standard it wouldn't have been a problem.
The web platform is incredibly backwards compatible; I'm having trouble thinking of any cases where they've dropped a format (which doesn't mean there aren't any!)
On the other hand, while the browsers generally ship with many codecs [1] the APIs for interacting with them are pretty terrible.
[1] https://developer.mozilla.org/en-US/docs/Web/Media/Formats/A...
Flash?
Flash used to run on every standard browser.
Now it doesn't exist.
Beyond API shifts, there are constant regressions and incompatibilities across browser implementations, especially in rich media APIs like image, audio, and video.
From a software point of view, more complex methods are also more prone to implementation error. Dominic's post is not opposed to Opus or the other formats, but argues (to me, convincingly) there is a valid spot on the map of audio formats for very simple format that still has more decent audio quality than other very simple formats. The result has compact file sizes, rapid decode times and still sounds well.
EDIT: The post is a good educational read (as any good software developer IMHO should aspire to creating simple code), initially and/or eventually. His other posts (e.g. the one on his Pagenode CMS) also show this simplicity-seeking mind set.
Opus decoding consumes ~10MHz of modern CPU core per one 128kbit stereo full-band track, or ~30MHz of Armv8. This translates to Raspberry pi3 at 1.2GHz decoding ~30 stereo tracks in parallel on single core.
Actually, FLAC is under 400-lines-simple.
> This is my independent implementation FLAC, optimizing the code for clarity to a human reader. The decoder is implemented in about 300 lines of source code, and the encoder in ~200 lines.
The problem with a QOI that would have 16-bit is that lossless becomes more expensive, exact match in the color table is more rare too and not worth it anymore.
You will start to need more prediction modes + offsets.
Using QOIX for elevation maps in PBR currently.
Not only does mp3 go to 320kbps, 192kbps was considered the 'standard' for ripping audio CDs, with the V0 variable bit rate being another common choice. if this only does 277, what am I getting aside from. another audio codec?
The ATRAC codec used in Minidisc is similar in bitrate to the QOA although it is transform-based. I had a music technology prof ask me "How do you stand listening to something compressed?" I also have a monster CD changer
https://www.crutchfield.com/S-sTSOm8D5jfj/p_158CDPX355/Sony-...
which I am filling up with 5.1 DTS disks that play on my home theater. I told my son that I find it hard to listen to 2-channel minidiscs next to really good 5.1 recordings with good bass management, but this weekend I did some heads-up listening testing between 2-channel CDs and minidiscs I made from the CDs and I could not tell the difference on my stereo. I am somewhat picky, I think "128 kbps MP3 suck" and can prove it. More careful A/B testing through headphones might reveal more weaknesses in ATRAC and any codec has some files that will stress it, but I am impressed with the quality of ATRAC and also with the low complexity. I have a portable player that plays for hours that runs off a single AA battery, mechanical parts and all.
QOA would have likely been the perfect alternative! We’d have better sound quality with no downsides, I think.
But yes, not quite fast enough for mp3.
A few years ago, a guy made his own video compression format, but I can't find it anymore, is it still available online?
Anyway, I'm grateful.
https://bengarney.com/2016/06/25/video-conference-part-1-the...
[1] https://github.com/phoboslab/qoa/blob/master/qoa.h#L324-L336
int quantized = (slice >> 57) & 0x7;
You wouldn't need this shift with a little endian format. If you store samples in order and read them as big endian, the first sample will be in high bits and the last sample in low bits, which is unnatural for shift and mask, that's what I call complication.
Really cool work!
https://github.com/xiph/flac/tree/master/src/libFLAC
https://github.com/SerenityOS/serenity/blob/master/Userland/...
It does not mean that the code is bad or inefficient, but the style and architecture is not the simplest to read.
Pretty much in the same way that published scientific papers are quite unpleasant to read.
AIFF and WAV say what?
On the other hand, AIFF does have a compressed variant ("AIFF-C") spec'd in 1991 [1], which I suspect the author wasn't aware of.
[1] https://www.mmsp.ece.mcgill.ca/Documents/AudioFormats/AIFF/A...
Most DAWs store audio internally as uncompressed 32-bit floats. That gives them plenty of headroom (which is important for gain staging).