Quite OK Audio Format
phoboslab.org
phoboslab.org
Haxe comes to mind but this seems a lot simpler.
QOA, the Quite OK Audio Format - https://news.ycombinator.com/item?id=34625573 - Feb 2023 (78 comments)
Quite Ok Audio Format Benchmark Results and File Format Specification - https://news.ycombinator.com/item?id=35729036 - April 2023 (0 comments)
QOA Benchmark Results and File Format Specification - https://news.ycombinator.com/item?id=35721419 - April 2023 (0 comments)
The Quite OK Audio Format [pdf] - https://news.ycombinator.com/item?id=35203976 - March 2023 (1 comment)
QOA, the Quite OK Audio Format - https://news.ycombinator.com/item?id=34625573 - Feb 2023 (78 comments)
A sequence like that, if you happen to see all of the posts, can make a topic feel over-represented, but I'd say this isn't really excessive, wouldn't you? Keep in mind that on HN, reposts are fine if a topic hasn't had significant attention in the last year or so (https://news.ycombinator.com/newsfaq.html).
In this case there was a major thread back in Feb but I decided not to treat the current post as a dupe because the two articles seemed to differ significantly.
(I don’t mind. The comments are always interesting.)
static inline int qoa_clamp_s16(int v) {
if ((unsigned int)(v + 32768) > 65535) {
...
}
This code is wrong. It assumes int is at least 32-bit which is not true in C. It should either use long or int32_t.And in that C struct.... is `char` and also `uint8_t`. I can never remember if `char` is signed or not. It should be `int8_t` or `uint8_t`.
This reminds me of the Farbfeld image format. So simple that I'm content to just make my own little-endian ad-hoc implementation and ignore the official spec.
char is effectively a trivalent type. Only use it when signedness is inconsequential. This is addressed in C23 with char8_t expressly to avoid the problems that happen with UTF-8 bytes as accidental signed chars.
Why would you want to run this on a modern high end PC or phone? There's almost no point using a simple codec like this if you have the computing resources or dedicates hardware to use a "proper" codec.
The codec is probably a reasonable choice for a microcontroller where you do want to save memory but you barely have CPU cycles to process the audio data. Many microcontrollers are big endian.
Same arguments for using the Quite OK Image Format over things like PNG or JPEG for instance (although since stb_image.h this isn't such a big issue anymore).
Another post in the same thread: "In the real world, for this type of code ..., only a very small number of CPUs (Intel and ARM) and C compilers (Clang, GCC and MSVC) matter ..."
These posts are strongly disagreeing with each other but neither of you directly replied to the other, so I'm not sure if you've seen it.
I'm replying to you because I think, between the two of you, you're the one that's wrong and the other poster is right. The original announcement doesn't mention microcontrollers at all and does mention computer games[1]. Further, their defense for the endianness decision was simply that they like it, rather than that they expect any target CPU to be big endian. I think GP is right--this will almost never be run on a big endian CPU.
Indeed, the start of this thread was a piece of code that won't work at all on smaller microcontrollers. I think it's likely this code has never been run on a microcontroller.
[1] https://phoboslab.org/log/2023/02/qoa-time-domain-audio-comp...
Technically it's neither (e.g 'char', 'signed char' and 'unsigned char' are three different types), but in practice this also usually doesn't matter (but it's more likely to be bitten by this specific detail than int not being 32 bits).
In the real world, for this type of code (audio code, most likely for games), only a very small number of CPUs (Intel and ARM) and C compilers (Clang, GCC and MSVC) matter, and on those, int is 32 bits wide.
PS: For new code I would also suggest using the "new" C99 fixed-width integer types though.
All values, including the slices, are big endian.
In a world where LE basically won, that's an astonishingly stupid decision --- especially if you care about tiny differences in performance.
The majority of file formats I've worked with are LE.
I was really thankful that Doom WAD file was big endian (side effect of being developed on NeXT?) when I ported doom to an nRF microcontroller.
Considering that little endian is objectively better for more types of data structures than big endian is [1], one would expect things to eventually converge on little endian (barring the entrenched big endian network protocols that hold everything back). Building for the eventuality of little endian would make more sense in a modern format.
For converting a "slice" from the 64-bit format (which is has a 4-bit value in the highest bits, and 20 3-bit values organized from high to low. Each 3-bit value is read from the high to low), I treated it as two separate 32-bit values. The bottom two bits of are combined to get the 4-bit value, and the rest contains two pairs of 10 3-bit values organized low to high.
So this:
sample = (slice >> 57) & 0x7; //Get sample
slice <<= 3; //Advance to next sample (64-bit shift)
Became this: sample = slice & 0x7; //Get sample
slice >>= 3; //Advance to next sample (32-bit shift)
I think the conversion to 32-bit little endian was something like an 8-10% speed up, with no other changes? (This was a few months ago, and I didn't write down what the exact difference was.) The overhead big endian conversion and 64-bit shifts would become even more noticeable after further optimizations. (Replacing qoa_lms_predict with inline asm to use an integer multiply-accumulate instruction that GCC can't generate got something like another 20-25% speedup. Properly written asm would be even faster.)Regarding clamping, might it be better to require the encoder to never generate QOA that would overflow? I've seen at least one ADPCM implementation (GameCube "AFC") that takes this approach. Seems like it wouldn't produce more distortion than clipping would.
AIUI decoding for some of these formats can be a battery drain, particularly for devices with tiny batteries like ear buds, watches, ipod nanos.
Is there any particular reason that AIFF uses an 80-bit extended float precision for the sample rate, instead of a 32-bit integer like WAVE?
I assume this was an oversight at the time, but I'm curious if there are actual applications where sub-Hz (or 2^32Hz) precision is needed.
What about lower sample frequency? (such as 28kHz)
If it could scale down, it'd be extremely useful for new productions for "retro" platforms.
It could be that it handles slower sampling rates well, but I am not sure about the bits.