Making Sense of the Audio Stack on Unix
venam.nixers.net
venam.nixers.net
Sidenote, audio stacks are not fundamentally fragmented. DirectSound is deprecated in favor of WASAPI. ASIO only exists because DirectSound sucked, now WASAPI is good enough not to use ASIO moving forward (with some hiccups, like some software doing SRC under the hood because WASAPI doesn't let you programmatically select sample rate - RtAudio does this, for example).
CoreAudio has been stable for almost 20 years. It is the gold standard. The only problem is the almost complete lack of documentation except for the CoreAudio mailing list.
If you're writing audio software the solution is to use WASAPI on windows, CoreAudio on Mac, and tell Linux users to use Windows or MacOS if they want low latency audio.
While I would have wished for a good clean solution from the start, the state of the current Linux Audio stack doesn't IMO really allow for such a solution without breaking every program ever.
If Linux Audio needs one thing it would be something that handles the weird relationships between ALSA/Pulseaudio/Jack in a clear and automatic fashion. Because as for now I have seen Linux Sysadmins with 30 years of experience tear their hair out because they couldn't wrap their head around Linux Audio.
Build a blackbox around it and try to rework it on the inside.
It already works better than pulse for bluetooth stuff, by a long shot. wtay is a good guy who is committed to the project, as opposed to the pulse maintainers who are jerks.
That doesn't seem fair to just throw out without proof. In all the interactions I've seen them in, they've come across alright for the most part.
Maintainers unresponsive, talking about making changes but not following up. They strung the guy along for quite a while on this new codec selection mechanism problem. They claim to be in a v14.0 freeze for two months. He updated his code to work on new v14.0 to more unresponsiveness from maintainers about merging. And it goes on from there.
If you've had to use pulseaudio-modules-bt or pali's patched version rather than official pulse, this is why. By comparison, pipewire managed to get it working with little effort because its lead developer is a nice guy and responsive.
Overall a really unfortunate turn of events there. On a personal note I've been watching that space for awhile and was looking forward to Pali's great work on bluetooth.
I don't think ASIO has been updated in several years either.
Not to mention, neither Mac or Windows come anywhere NEAR the awesomeness that is Zynthian .. and if anything demonstrates the amazing flexibility and power of Audio on Linux, its the Zynthian system.
Simply amazing audio!
This means "I don't want to think about it" installing a whole distro is way too much work for someone with this attitude.
Are you trying to hook Bitwig up to other realtime software with jackd? Because if not-- e.g., you're just trying to output sound from Bitwig with the lowest latency possible-- there is absolutely no point in running jack. Go straight to ALSA.
Note: I'm not familiar with Bitwig and am assuming it isn't a weirdo that somehow prevents you from setting ALSA as the backend on Linux.
The downside to "straight-to-ALSA" is that a realtime prog will gain exclusive control over the audio subsystem. So someone who wants to watch a youtube video demo on realtime prog Z and run realtime prog Z to monkey along will end up with a frozen video as it blocks (or perhaps Pulse blocks) waiting for prog Z to release the audio to it. (Or if the video started first, prog Z will (hopefully) error out trying to gain an exclusive connection to ALSA that it cannot get.)
Pulse is valuable because its default behavior solves this use case. Jack can be set up to do this, too-- I have used Qjackctl or whatever to automate this use case when setting things up. But at least last time I used Jack I had to "opt in" to this obvious setting of "pipe the running stuff to the speakers, please." E.g., I had to do business with my brain and some global mutable state. Not sure if that's changed.
AFAIK ALSA-only distros tend to have SW mixing configured by default and PulseAudio-based distros have ALSA intercepting enabled by default. So in practice an application that needs some basic audio playback can just use ALSA and it'll work everywhere.
"In practice" doesn't mean what you think it means.
The behavior I described was with a Pinebook Pro which only came out about a year ago. That was running whatever version of Debian they shipped-- either stretch or buster, can't remember now. And there was no fancy or custom audio setup-- it was just an XFCE desktop environment with Pulse running however it's configured to run by default.
Also, as a maintainer of realtime audio FOSS software I can tell you users have reported these exact problems on other Linux distros. Over at least the last year and a half.
It appears that in Debian testing the configuration of Pulse "just works" as you describe. I would rankly speculate that Ubuntu 20.04 probably has a similar default setup. But to turn that rank speculation into "in practice" we'd need to actually test and see, and it sounds like neither of us have done that yet.
BTW what i describe is how Debian was for years when i used it by default, not just testing. It is possible that the version of Debian Pinebook Pro shipped was altered. Or something else wasn't configured as i had it, who knows.
I had it all working with Pulse + Jack a year ago, on a previous workstation build, but it was a massive pain to get there. I had my m-audio USB keyboard working, along with Tracktion Waveform and VCV Rack. (None of which give tolerable audio quality with vanilla Pulse, even on new Ryzen 5600 + dedicated Sonar audio card.)
Based on the last experience, this time I'll take a snapshot of current .deb packages installed, a full copy of /etc/ and all my dot files (audio stuff seems to be spread far and wide under app-specific ~/. files as well as the expected places ~/.local and ~/.config).
I recall even when it was working, occasional upgrades (albeit Debian unstable) would stop various components, and troubleshooting an audio stack, if it's not something you do regularly, is slow and frustrating work. Referring back to earlier configuration files and package names should make this easier.
First step though is working out again why Debian still has jackd1 and jackd2 packages. The first search result for that question is from 2011, and it sounds like OSS / ALSA is the only reason jackd1 would be preferred. And there you have a fantastic example of how befuddling audio on GNU/Linux can be for the casual user.
My Linux DAW easily outperforms my Mac DAW, and I have tens of thousands of buckaroos invested in the latter, but only a couple hundred in the former. Ubuntu Studio is simply the best out-of-the-box DAW user experience around.
>If you're writing audio software the solution is to use WASAPI on windows, CoreAudio on Mac, and tell Linux users to use Windows or MacOS if they want low latency audio.
This is nonsense advice. I'll instead tell Linux users to use a decent distro, and target JACK+Alsa for my Linux builds ...
So looking at Pipewire, let's hope for it to make Pro Audio and Consumer Audio interoperable on Linux. Because I really like the flexibility of Jack and the PulseAudio way of user experience.
I've been running this on my Ryzen desktop with Arch Linux installed for about a year now and it's been incredibly stable and reliable. I haven't had any issues with playback or recording latency either. Along with Bitwig this is my favorite audio production setup ever and I used Ableton on a Mac for years. If you're looking for a good tutorial on configuring JACK, I'd highly recommend this video from Unfa on YouTube [0]. He does a lot of Linux and free software audio production content which you might find interesting.
Thing have come a long way since old days when apps used to exclusively lock the sound card.
This is what it looks like if you only include things that are actually relevant (or don't try to actively make it look bad, IDK why else OpenBSD's sndio would be included in the Linux diagram):
https://aurora.nox.tf/tmp/linux-audio-output.svg
NB: this is specifically the Linux situation, and the original file on wikimedia has a deletion request on it because it's so bad.
To this day, every time I upgrade my distro to a major revision I just can't assume the audio will keep working. Latest upgrade of Ubuntu, and I had to download alsamixer to get audio coming from my headphones again on Zoom meetings. Both the output selector in Zoom and the selector in Settings just... Stopped working. I could select the headphone output, and nothing changed.
Every now and then I get booted to cli on startup and I have to reinstall the Nvidia drivers to get sddm to launch properly.
Also every time I reboot it's completely random if the login screen will be at 1080p or 4k (I have a 4k monitor).
Maybe it's because I've got an Nvidia card.
I'm scared to update because I know I'll have to spend hours fixing the audio and video, and whatever else again.
Apparently old models get dropped just as with binary blobs, or maybe not, given that the Windows drivers work today just as 10 years ago with GL 4.1 and DX 11.
Just because you are unhappy that your card was too old to be included in the current open source driver does not mean you have to spread FUD about it.
If anything, it helped my decision to focus on Apple, Microsoft and Google platforms, and let go of "Year of Linux Desktop" mantra.
So now I have to go through Fn+NumLock dance every time I bother to use the netbook.
This is also my opinion, which is why I was surprised by the conclusion of the article. Money quote:
> The audio stack is fragmented on all operating systems because the problem is a large one. [...] The stack of commercial operating systems are not actually better or simpler. [...] Linux is the platform of choice for audio and acoustic research and was chosen by the CCRMA (Center for Computer Research in Music and Acoustics).
I guess my gripe is this disconnect between "desktop audio" and "pro audio". I have my PulseAudio set up the way I want it, and that's what I need 99% of the time, but I would like to use Ardour when I want to record a podcast, but whenever I read the Arch wiki about how to have PulseAudio and JACK coexist, it looks like a huge mess that I don't want to deal with. Here's hoping PipeWire really shapes up to be my savior one day.
https://github.com/brummer10/pajackconnect
start jack, start pulseaudio, run pajackconnect, done
This used to work perfect, now I always have to keep pavucontrol open when I'm in a conference so I can quickly switch the streams back.
Thankfully, the last ten years have been plain sailing. It just works. Oh, except that on Gnome the volume remote half way up my headphones cable adjusts the Pulse Audio master level rather than the volume of the headphones themselves, but I can live with that.
Today I learned: Unix now is Linux and free bsd.
Get me right, this sounds good, great, actually. Still, it amazes me.
- AIX is Unix from IBM - A/UX is Unix from Apple. - They’re both AT&T System V with gratuitous changes. - Then there’s HP-UX which is HP’s version of System V with gratuitous changes. - DEC calls its system ULTRIX. - DGUX is Data General’s. - And don’t forget Xenix—that’s from SCO.
So if we want to talk about the most popular POSIX compatible systems today that are free we're left with the different kinds of BSDs and Linux OS. I also mentioned MacOS in the article, in relation to CoreAudio, but decided to leave it out as its design can't be inspected and compared properly. And also it's not a very popular OS amongst my readership and the world in general.
There's also the part where writing about it is not super interesting, because it mostly just works like it should.
And the binary compatibility of GNU/Linux, the interchangeability of suse, debian/ubuntu, red hat, keeps the promise that posix gave.
Which leads to the tangential point: a new generation which is growing up with "unix? That's Linux, right? Oh, and free bsd..."
Great, great article, btw
Arguably, this is sort of true today, adding MacOS to the mix. With Oracle Solaris discontinued, and Oracle Linux being where the work is at, plus a bit of IBM AIX and HP-UX used in some companies server but less common. That's mostly where the POSIX/Unix-like landscape is at today.
It's not perfect and still have bugs to iron out. I just can't wait for it to reach 1.0 ! Yet, I also hope that it does not get pushed too soon into distros as default ...
Maybe so for developers, but as a user on Windows or MacOS I don’t have to know or care about any of that stuff. Things just work, and work sensibly.
depends on the kind of user you are. If you're looking for the lowest-latency possible (for instance because you play guitar with an effect stack going through your computer) then windows with ASIO or linux with Jack give you a better mileage on the exact same hardware than CoreAudio
Personally I can easily notice >50ms audio desync in videos, but I imagine one might be more latency sensitive due to the tactile feedback?
In reality you either move to analog processing when you reach that much delay or design your creative workflow around it (learn not to expect immediate feedback) so you move towards offline processing.
And, with the "normal" system APIs (except CoreAudio which fares a bit better) it's pretty hard to get below 15ms in my experience, even with beefy computers.
Ok, some domain expert on HN:
What is the UCM, when was it added to ALSA, what is its purpose, who is the target audience (e.g., is this something that helps driver author?), and how does it relate to Pulse?
I'm especially concerned with the last question, because at a glance it seems like Pulse and the UCM overlap partly in purpose.
It was always on the table though and now it has been implemented. That’s why there’s significant overlap, those projects are not complementary, they are somewhat competitive.
Void Linux has many programs compiled with sndio support now:
And this is why pure UNIX desktops never go anywhere.
> Linux is the platform of choice for audio and acoustic research and was chosen by the CCRMA (Center for Computer Research in Music and Acoustics).
This is definitely not the case as anyone that hangs around professional musicians can tell.
Just go into YouTube and see how many bother to do tutorials with Ableton on Linux.
https://github.com/zynthian/zynthian-sys/issues?q=is%3Aissue...
So many wonderful audio devs gathering around the Zynthian (Linux-based audio DAW) waterpool. Synth developers, FX developers - everyone.
EDIT: The Zynthian project even helped Apple M1 users get software on their platform that would not, otherwise, have been an easy task .. if it weren't for Linux Audio/synth hackers, this wouldn't have happened nearly as quickly, nor as smoothly, as it did:
https://github.com/surge-synthesizer/surge/blob/main/README....
.. you don't often see things going the other way (i.e. MacOS -> Linux) ..
While interesting, hardly any leading number 1, top of the tops.
Ableton runs just fine on Apple M1.
Well, my point is: its happening, even if its not being done with whitepaper sniffers and marketing people involved.
There are far more audio hardware manufacturers using Linux in their products than ever before. There are also far more pro-level audio applications being developed for Linux, than ever before. It is a rapidly expanding market for Linux, and it has the potential to seriously undermine the current leaders with disruption. Once the pro-audio world catches wind of what can be done with Linux Audio, packaged in a nice box (a la Zynthian), I'll wager that within a year this situation of Linux-as-underdog in the audio world will be very, very different.
Its already happening in studios around me - I see people freaking out over the M1 move and Apple lockdown, looking for alternatives. UbuntuStudio is a godsend for those in that situation ..
EDIT: notice that I linked to the CLOSED issues for Zynthian, and that a majority of those are from 3rd-party developers getting their synths and effects products packaged for Zynthian. This represents a sea change - not to mention the fact that this work from the Linux Audio world is rapidly propagating great value to the M1 (Apple ARM) situation ..
I can be gladly proven wrong, assuming you have numbers to share, otherwise we are talking about "Year of Linux Audio".
To be fair besides getting audio out I never really needed to go deeper until I started an audio-heavy personal project a few weeks ago. This article certainly helped clarify quite a bit of the map.
I guess what you describe as "such simple thing as output mono" has multiple solutions once you think about it, so no-one is forcing one
The biggest hurdle is with hardware. If you need low-latency, high-performance interfaces, you absolutely a need tightly integrated hardware/software stack - think something like Universal Audio’s Luna system. You simply cannot make that kind of tight integration work outside of macOS.
It really depends on what someone is doing. Ubuntu went off the rails years ago, and I would default to recommending OSX, but as someone who uses all three, Linux has been the least painful for me.
Once you get UbuntuStudio installed, out of the box - it just plain works, and is bundled with SO MANY GREAT THINGS. I once installed UbuntuStudio alongside a friend of mine in the pro studio, who had just gotten a new Mac.
I was up and running and working on projects within 20 minutes from a fresh install - he was still hassling with UAD licenses and dongles, a week later.
Seriously people, if you think you are an audio professional, and haven't yet had a dive into kxStudio or UbuntuStudio on a spare machine, you are missing one of the greatest parties in the audio world right now.
And then, if you want to get REALLY sophisticated, as an audio professional, you NEED to dig into - and most important of all, understand - what is going on with the Zynthian folks.
This is the future of audio - not just for DAW's, but embedded audio devices, too. When you grok what Zynthian is, and how it works, and what it is doing, you will change your mind about everything you think you know about audio on Linux.
(Hint: it takes the extremely powerful Linux Audio/DAW software ecosystem, wraps it in a functional and simple UI, and exposes EVERYTHING in the Linux Audio world to the musician .. from synths to FX units to MIDI processing and generators and beyond. Simply a massive leap forward for audio professionals...)
The trouble is, most "audio pro's" don't really have the technical chops to understand why this is the case with MacOS, nor why Linux represents a parallel advance in audio processing technology. You can do things with Linux audio that are impossible with MacOS, due to its general-purpose nature - but MacOS still provides great value to the consumer, not just professionals.
The great thing about Linux audio is that the devs pushing it forward don't have any of the constraints that the MacOS/iOS folks have to work under. This gets mis-underestimated by audio pro's, imho.
But with things like kxStudio and UbuntuStudio and Zynthian out there, this is going to change radically. I think the audio world is in for some very big upheavals in terms of Linux adoption over the next few years .. but I have all of these systems operational in my pro studio, and have poked these beasts multiple times. Linux rules for audio, and hardly anyone knows .. thats another kind of elite realm left to be explored by the curious and adventurous.
zynthian has a nice website. it seems like a cool idea and I'm sure it has some great use cases. it's still a kit which strikes me as not pro-level.
more concerning is that it appears to have what I would call completely unusable issues regarding latency and jitter (5-20ms of latency is unusable for realtime audio), can you comment on this? how about jitter? what is your experience with it (and if you have any, what are your use cases?)
searching the internet for solutions here seems to show a lot of unresolved issues. I see the developers asking users to file GitHub issues as recently as a couple of months ago. I know very few audio professionals who are going to want to use something with this degree of maintenance / instability / unreliability.
how would you describe the ease with which you can interface the tools you describe with outboard gear using USB or MIDI e.g. audio interfaces, mixers, synthesizers, etc, compared to the apple ecosystem? which tends to be quite plug & play. one reason why it is so widely used.
It is free, after all, and yet out of the box comes replete with tons of stuff that you'd have to spend, literally, a week getting installed on MacOS/Windows, dealing with licenses, etc.
Zynthian: The latency is really not noticeable - I have a 12 track multi-timbral project running on it right now, it sounds just as rock solid and tight as the 3 rows of 19" rack synths/effects and the half-a-room full of other hardware synths in my studio. It is an extremely capable little box, and its built around existing Linux audio technology - so the point is, it serves as a pretty good reference for how good things can get. Jitter? Haven't noticed it.
Zynthian integrates with all the DAW's in my studio, including UA and MacOS systems, just fine. Its a dream to be able to set it up for remote control over rtpMIDI, for example, and not have to run any cables.
Github issues are a professional way to interact with developers. If I had 2c for every time I heard someone complain about the support they -didn't- receive from Apple because of bugs in Logic, or crashes that Steinberg/Yamaha refuse to fix, or updates to Ableton Live that resulted in unusable projects .. I'd have enough to afford a decent UAD system .. ;)
On the other hand, if you look at the closed issues in the Zynthian github, you'll see so many successful contributions from around the world, making the Zynthian a better product. This is the future.