New Ultra Fast Lossless Audio Codec (HALAC)
hydrogenaud.io
hydrogenaud.io
> TOS 8. All members that put forth a statement concerning subjective sound quality, must - to the best of their ability - provide objective support for their claims. Acceptable means of support are double blind listening tests (ABX or ABC/HR) demonstrating that the member can discern a difference perceptually, together with a test sample to allow others to reproduce their findings. Graphs, non-blind listening tests, waveform difference comparisons, and so on, are not acceptable means of providing support.
He explicitly says it's not SIMD, which is nice because it rules out a way of cheesing it, but still...
There are some areas that I follow as-it-develops, but codecs and data compression is one that I'll use when ready. Still awaiting widespread av1 support/adoption.
The area most needing improvements IMO is with Bluetooth, especially Apple's support of codecs (where they're dropping support that worked in older macOS versions).
Still really really impressive to beat an established standard as an individual, that doesn't happen much.
Some of the SIMD optimization make compilers automatically. The others can be manually. These speeds can be obtained without using SIMD. Then there may be more with manual SIMD. I think this is what should be loved.
Apples and oranges. I don't need word processing software to be open source to understand how it works. A proportedly novel compression algorithm is a different story...
I can be totally honest with you: FLAC being open source is more valuable to me than any performance benefit you could ever possibly offer over it. It only becomes interesting if I can actually read the code and see what you did.
I am genuinely interested in what you've done here, and I sincerely hope you publish it.
I am developing HALAC and HALIC as a hobby and I don't expect everyone to use them. I'm happy when I can get good results, and it's bad when I can't. I say this as someone who has been dealing with data compression for 9 years.
Obviously, it's your right to decide... but especially if you think of it as a hobby, why not release the source? It would make your work much more valuable for a lot more people.
When I bring my work to a certain stage, I would like to deliver it to a team that can claim it. However, I want to see how much I can improve my work alone.
The amount of media playback and serving software out there is innumerable. If most of it doesn't handle some obscure format, that format is screwed.
Getting a new format everywhere is a difficult battle; the adoption barriers are high. Even if the thing is completely royalty free, and comes with a great, open source reference implementation.
Something that is closed, and has no backing of some corporate consortium or ITU type body or whatever, is basically fucked.
What interests me the most is the memory usage.
It seems that there is a serious effort. And your work clearly crushes other codecs. This will cause serious discomfort in some circles. I don't know much about HALIC, but it is said that it is also a very serious work.
There's no point in rushing for open source. First of all, improve your work as much as you can. If serious offers come later, you can evaluate them. Don't give up control.
I would give you the benefit of the doubt that it might just be code shyness or perfectionism about something in its early stages, but it looks like the last codec you developed (“HALIC”) is still only available as Windows binaries after a year.
I struggle to see an upside to withholding source code in a world awash with performant open source media codecs.
By the way, there will always be things to add! That feeling should not stop you from putting the source out there - you will still own it (you can license the code any way you like!) and you can choose what contributions make it in to your source.
From the encode.su thread and now the HA thread, you've clearly gotten people excited, and I think that by itself means that people will be eager to try these out. Lossless codecs have a fairly low barrier for entry: you can use them without worrying about data loss by verifying that the decoder returns the original data, then just toss the originals and keep the matching decoder. So, it should be easy to get people started using the technology.
Open-sourcing your projects could lead to some really interesting applications: for example, delivering lossless images on the internet is a very common need, and a WASM build of your decoder could serve as a very convenient way to serve HALIC images to web browsers directly. Some sites are already using formats like BPG in this way.
This is a very valid point, but we should all recognise that some people⁰ explicitly don't want that for various reasons, at least not until they've got the project to a certain point in their own plans. Even some who have released other projects already prefer to keep their new toy more to themselves and only want more open discourse once they are satisfied their core itch is sufficiently scratched. Open source is usually a great answer/solution, but it is not always the best one for some people/projects.
Even once open, “open source not open contribution”¹ seems to be becoming more popular as a stated position² for projects, sometimes for much the same reasons, sometimes for (future) licensing control, sometimes both.
--
[0] I'm talking about individual people specifically here, not groups, especially not commercial entities: the reasons for staying closed initially/forever can be very different away from an individual's passion project.
[1] “you are free to do what you want, but I/we want to keep my/our primary fork fully ours”.
[2] it has been the defacto position for many projects since a long time before this phrase was coined.
The "primary" fork is the one that the community decides it to be, not what the authors "wants". Does it really matter what is the "primary fork" for those working on something to "scratch their own itch"?
If I were in the position of releasing something⁰: the community, should one or more coalesce around a work, can do/say what it likes, but my primary fork is what I say it is¹. It might be theirs, it might be not. I might consider myself part of that community, or not.
It should be noted that possibility of “the community” or other individual/team/etc taking a “we are the captain now” position (rather than “this is great, look what we've done with it too” which I would consider much more healthy and friendly) is what puts some people off opening their toy projects, at all or just until they have them to a point they are happy with or happy letting go at.
> Does it really matter what is the "primary fork" for those working on something to "scratch their own itch"?
It may do further down the line, if something bigger than just the scratching comes from the idea, or if the creator is particularly concerned about acknowledgement of their position as the originator².
--
[0] I'm not ATM. I have many ideas/plans, some of them I've mused for many years old, but I'm lacking in time/organisation/etc!
[1] That sounds a lot more combative than I intend, but trying to reword just makes it too long-winded/vague/weird/other
[2] I wouldn't be, but I imagine others would. Feelings on such matters vary widely, and rightly so.
Don't worry. You don't have to.
If you want a more specific response than that, what part of the post do you not get?
> other individual/team/etc taking a “we are the captain now” position rather than “this is great, look what we've done with it too”
The scenario is that someone opens up a project but says "I am not going to take any external contribution". Then someone else finds it interesting, forks it, that fork starts receiving attention and the original developer thinks to be entitled to control the direction of the fork? Is this really about "scratching your own itch" or is this some thinly-veiled control issue?
I'm sorry, after you open it up you can't have it both ways. Either it is open and other people are free to do whatever they want with it, or you say "it's mine!" and people will have to respect whatever conditions you impose to access/use/modify it.
> if the creator is particularly concerned about acknowledgement of their position as the originator.
That is what copyright is for and the patent system are for those who worry about being rewarded by their initial idea and creation.
If one is keeping their work to themselves out of fear of losing "recognition", they should look into the guarantees and rights given by their legal systems, because "feelings on this matter" are not going to save them from anything.
I wasn't attempting to veil it at all. It is a control issue for some.
Sometimes someone is happy to share their project, but wants to keep some hold on the core direction.
> > other individual/team/etc taking a “we are the captain now” position rather than “this is great, look what we've done with it too”
The scenario is that someone opens up a project but says "I am not going to take any external contribution". Then someone else finds it interesting, forks it, that fork starts receiving attention and the original developer thinks to be entitled to control the direction of the fork?
You are missing a step. I said that if someone has this concern then they might not open the project at all, until they feel ready to let go a bit. At that point “open source but not open contribution” and control over forks are not issues at all because the source isn't open and forking isn't possible.
> That is what copyright is for and the patent system are for
I don't know about you, but playing in those minefields is not at all attractive to me, and I expect many feel the same. If I had those concerns, and legal redress is the solution, I now have two problems and the new one is a particularly complex beast, it would be much easier to just not open up.
Then do not hide it behind the "people just want to scratch their own itch". It is a bad rationalization for a much deeper issue and the way to overcome this is by bringing awareness to it, not by finding excuses.
> wants to keep some hold on the core direction.
You are really losing me here. The point from the beginning is that the idea of "direction" is relative to a certain frame of reference. There is no "core" direction when things are open. The very idea of "fork" should be a hint that it is okay to have people taking a project in different directions.
> it would be much easier to just not open up.
Agreed. But like I said: you can not have both ways. If you want to "keep control" and prevent others from taking the things in a different direction, then keep it close but be honest to yourself and others and don't say things like "it's not ready to be open yet" or "I want to share it with others but I worry about losing recognition".
You seem to be latching on to individual sentences in individual posts rather than understanding the thread from my initial post downwards. Start from the top and see if that changes.
Right from the beginning I was walking about people not releasing source for this reason, not releasing with expectations of control – while quoting more of the preceding thread might have made that sentence look less like an attempt to hide as you see it, that would bulk out the thread necessarily IMO (and I'm already being too wordy) given that the context is already readily available nearby (as the thread is hardly a long one).
> > it would be much easier to just not open up.
> Agreed. But like I said: you can not have both ways. If you want to "keep control" and prevent …
No, but the other end of the equation often wants the source irrespective of the project creator not being ready to let go of fuller control just yet (for whatever reason, including wanting to get to a certain point their way to stamp their intended direction on it). And they will nag, and the author will either spend time replying to re-explain their (possibly already well documented) position or get a reputation for not listening which might harm them later.
From the top of the thread: "it is difficult to take it seriously if you refuse to offer source code or a implementable specification.". If OP has reservations about building it the open, I'd rather hear "I am not going to open it because I want to keep full control over it" then some vague "I will open it after I complete some other stuff".
You mention the concern about "getting a reputation for not listening". To me, this has already happened. The moment I saw "when I realize it can be used by someone, it will be of course be open source", I'm already doubting his ability to collaborate, I already put him in the "does not understand how FOSS work" box and I completely lost interest in the project.
> When I bring my work to a certain stage, I would like to deliver it to a team that can claim it. However, I want to see how much I can improve my work alone.
1) Publishing the code does not mean that it is done. Software development is a continuous effort.
2) There is no "delivering it to a team that can claim it". When (if?) you release your code, you will see the possible outcomes:
- the worst case scenario, someone will find an issue on your design and point to a better alternative and you will be left alone with your project.
- The best case scenario, your work brings some fundamental breakthrough and you will have to spend a good amount of time with people trying to figure it out or asking for assistance on how to make changes or improvements for their use case.
- The most likely scenario, your work will get its 15 minutes of fame, people are going to be taking a look at it, maybe star at Github and then completely leave it up to you to keep working on the project until it satisfies their needs.
Like "everythingctl" said, you will see that few people will take you seriously until you actually show source code or an reproducible specification. But you will also see that is a "required but not sufficient condition" for you to be taken seriously. And while I completely understand the fear of putting yourself out there and the possibility of having your work scrutinized and criticized for things you know need improvement, I think that this mentality is incompatible with the ethos of Open Source development and I wish more people can help you overcome this fear than tried to excuse or defend it.
Halic = kiss Halac = raspy voice
No great project started out great and the best open source projects got to their state because of the open sourcing.
Consider the problems you might be spending a lot of time solving might be someone else's trivial issue, so unless this is an enjoyable academic excercise for you (which i fully support), why suffer?
Or maybe it's better for me to do things like fishing, swimming as a hobby.
There is a chicken and egg problem with this strategy: Few people will want to, or even be able to, use this unless it’s open source and freely licensed.
The alternatives are mature, open or mostly open, and widely implemented. Minor improvements in speed aren’t enough to get everyone to put up with any difficulties in getting it to work. The only way to get adoption would be to make it completely open and as easy as possible to integrate everywhere.
It’s a cool project, but the reality is that keeping it closed until it’s “completed” will prevent adoption.
Making it opensource now would just ruin that leverage.
I am with you OP
1. Not FLAC
2. Not as open-source as FLAC
comes across as a patent play.
FLAC is excellent and widely supported (and where it’s not supported some new at-least-open-enough codec will also not be supported). I have yet to see a compelling argument for lossless audio encoders that are not FLAC.
FLAC doesn’t support more modern sampling formats (e.g. floating point for mastering), or complex multi channel compression for surround sound formats.
There just isn’t something better (and free) to replace it yet.
Apple's ALAC (Apple Lossless Audio Codec) format is an open-source and patent-free alternative. I believe both ALAC and FLAC support up to 8 channels of audio, which allows them to support 5.1 and 7.1 surround. https://en.wikipedia.org/wiki/Apple_Lossless_Audio_Codec#His...
These are distribution formats, so I'd be surprised if there were demand for floating-point audio support. And in contexts where floating point audio is used, audio size is not really a problem.
Unless things have changed substantially and I missed it, FLAC does not do similar tricks for other multichannel audio modes. Meaning that for surround sound, each channel is independently compressed and it is unable to exploit signal correlation between channels.
Proprietary formats like Dolby on the other hand do support rather intelligent handling of multichannel modes.
FLAC is not solely a distribution format. Indeed as a distribution format it sucks in a number of ways. It is chiefly used as an archival format, and would in fact be ideal as a mastering format if these deficiencies Could be addressed.
I don't use or desire multichannel audio but that and the hardware acceleration are interesting points.
Multichannel audio support is nice because it is often used in distribution of media files sourced from DVD/BluRay. It would be good to have a high quality, free codec for that use.
MP3 is a lossy format so I would practically guarantee that you’d end up with a smaller file but that’s not the purpose of FLAC. Lossless encoding makes a file smaller than WAV while still being the same data.
> e.g. floating point for mastering
I’m 0% sold on floating point for mastering. 32bit yes, but anyone who’s played a video game can tell you about those flickering textures and those are caused not by bad floating point calculations, but by good floating point calculations (the bad part is putting textures “on top” of each other at the same coordinates) . Floating point math is “fast” but not accurate. Why would anyone want that for audio (not trying to bash here, I’m genuinely puzzled and would love some knowledgeable insight)
You misunderstood what you are replying to. FLAC works by running a lossy compression pass, and then LZ encoding the residual. The better the lossy pass, the less entropy in the residual and the smaller it compresses. FLAC’s lossy compressor pass was shit when it came out, and hasn’t gotten any better.
Flickering textures is caused by truncation and wouldn’t be any better with integer math. The same issues apply (and are solved the same way, with explicit biases; flickering shouldn’t be a thing in any quality game engine).
Floating point math is largely desired for mastering because compression (technical term overloaded meaning! Compression here means something totally different than above) results in samples having vastly different dynamic ranges. If rescaled onto the same basis, one would necessarily lose a lot of precision to truncation in intermediate calculations. Using floating point with sufficient precision makes this a non-concern.
Since when does FLAC run a lossy pass? You can recover the original soundwave from a FLAC file, you can't do the same with an MP3.
I'm pretty sure FLAC does not run a lossy compression pass.
Flickering textures in game engines are likely due to z-fighting, unless you're referring to some other type of flickering.
If you're looking to preserving as much detail as possible from your masters then floating points make sense. But its really overkill.
> If you're looking to preserving as much detail as possible from your masters then floating points make sense.
I’ve been searching for hours and gotten nothing more than the classic floats vs ints handwaving. Can you explain what you know about why using floats preserves detail?
Linear predictor is a form of lossy encoding.
Where is the source code? A detailed description of what the codec actually does? References to relevant publications?
All I see are two mystery Windows binaries, hosted on a forum I've never heard about. The fact that "encode.su" uses the world's most notorious domain extension doesn't inspire confidence, to put it mildly.
And later, you will learn to understand the depth of the contribution that the ffmpeg project provides :)
New world indeed...
curl http://example.com/script.sh | bashI've searched the web and found not a single mention of "high availability" in the context of audio codecs. In fact, the top-ranked result was this very post, which doesn't explain what the term is intended to mean.
The leaders in data compression and information theory are all from the former Soviet Union, so that's no cause for concern.
I think Andrey Markov started it all.
Encode.su (formerly encode.ru) is indeed the most known forum dedicated for data compression in general. So much that many if not most notable data compression projects posted to HN are often first advertised to that forum first (for example, Zstandard [1]).
What does this mean? Why or how is this TLD the world’s most notorious?
* https://en.wikipedia.org/wiki/.su
I would think that ICANN/whomever would have mandated its retirement / de-orbit, but a special exception was asked for:
* https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2#Exceptional...
[1] https://www.interisle.net/PhishingLandscape2023.pdf#page=18
And I absolutely do at least a quick visual "sanity check" of the code before compiling and running newly announced software.
All it takes is one person somewhere who wants to look something over, and they heads-up the rest, and then many others do verify.
And that initial one does exist even though it's not you or me, the same way the author exists, the same way that at least once in a while for some things it is you or me.
It's actually decently handy. Just keep in mind no sandbox is perfect.
It’s basically just a forum post with a self-signed .exe download…
Oh well.
1. https://encode.su/threads/4025-HALIC-(High-Availability-Loss...
> Though it is fast indeed! The decoding speeds are outright impressive given how FLAC is the fastest thing we ever saw ...
I personally keep everything in flac, but Bandcamp is seemingly the only service where that is a given.
Ok, but what would make that useful?
Lower run-time means less electricity and less tying up of the CPU, making it available for other things. As a real-life example: i frequently use my Raspberry Pi 4 to convert videos from one format to another. This past week i got a Pi 5 and moved the conversion to that machine: it takes maybe 1/4th as much time. The principle with a faster converter, as opposed to faster hardware, is the same: the computer isn't tied up for as long, and not draining as much power.
Google once, back in 2013, made an API change to their v8 engine because it saved a small handful of CPU instructions on each call into client-defined extension functions[^1]. That change broke literally every single v8 client in the world, including thousands of lines of my own code, and i'm told that the Chrome team needed /months/ to adapt to that change.
Why would they cause such disruption for a handful of CPU instructions?
Because at "Google Scale" those few instructions add up to a tremendous amount of electricity. Saving even 1 second per request or offline job, when your service handles thousands or millions of requests/jobs per day, adds up to a considerable amount of CPU time, i.e. to a considerable amount of electricity, i.e. to considerable electricity cost savings.
Indeed, getting an accurate answer would require looking at the whole constellation for a given use case.
Therefore, it is not a logical choice to increase the process rate in order to provide a few percent more compression between audio codecs. As a result, high processing times are high energy.
Why not? And for what applications? Example: for a media streaming service, where each file is transferred many times, the bandwidth costs dominate, so it is worthwhile to spend a great deal of time on encoding to maximize efficiency. In the case of an archive, where a large amount of information is stored, accessed infrequently, storage space becomes the constraint, once again. In general, 1 marginal second of CPU time is usually cheaper than 10Mib of marginal storage (or whatever the figure works out to be). Finally, why not just write a fast FLAC encoder?
Flac is already existing. There are also a lot of workers on it. I always want to try independent and different things.
Also your link doesn't explain what they changed?
At a large-enough scale, all savings are significant.
> Also your link doesn't explain what they changed?
They changed a function signature to use an output argument instead of a return value. i don't recall the exact signature, but it was conceptually like:
v8::Value foo(...);
to void foo(..., v8::Value &result);
Why? Because their measurements showed a microscopic per-call savings for the latter construct.PS: i wasn't aware that source code for this codec is not available. That of course puts a damper on it.
Yes, but not every tradeoff between compression speed and compression ratio is something that makes sense to scale in the first place.
These are basic sanity checks.
Something is wrong in the benchmark.
I don't know about FLAC, but from my knowledge of compression, this result seems sensible to me.
Smaller file = less bits to process = faster. FLAC level 5 is expected to give a smaller file than level 0, so it makes sense that decoding it will be faster. Of course, it's possible that some codecs enable more codec features at higher compression levels, which makes decoding slower, so it's not always a given, but higher compression giving faster decode doesn't seem unreasonable.
In realtime applications, any processor within the last decade is far more than powerful enough to encode and decode in realtime. The first set of tracks in the results is 2862s long and even the slowest WAVPACK manages to encode it in >113x realtime and close to 250x realtime for decode.
For archival, high compression is important, but this codec doesn't compress better than WAVPACK either.
This is addressed in another response in this same thread regarding electricity usage.
Something else you can do is use (A)GPL3. This means you automatically grant patent licenses, and anyone building on your work also has to release their source. You can then separately sell proprietary licenses without any of these restrictions.