The Linux audio stack demystified (and more)
blog.rtrace.io
blog.rtrace.io
First, in the article, [1] shows in one single diagram where the complexity is coming from; the audio system has to handle a good deal of different hardware on many different systems and also provide extra functionality for multiplexing, network features, wireless headsets and their codecs, etc. All this: open source.
Second: linux is the only platform where everything works right now flawlessly for me: My Bose joins without any problems, switches to headset mode during zoom calls, switches back to high definition audio otherwise. I can select a different sink, even networked, whenever I want the output to appear on a different networked device. MacOS sometimes needs a reboot so that the bluetooth subsystem works, what the hell.
And all this worked with PulseAudio, and now works with Pipewire, which is an even higher quality iteration of PA.
I don't complain. I wish MacOS/Windows had such a versatile, configurable, but sanely-working-out-of-the-box audio system as an off-the-shelf Fedora, or even freaking Arch linux has.
HTH
[1]: https://blog.rtrace.io/images/linux-audio-stack-demystified/...
OEMs also love to include all kinds of audio related bloatware, which makes getting audio hardware to work reliably(!) quite challenging. The HP laptop I'm typing this on has an "Intel microphone array" (just 2 mics) which has it's own intel drivers, but there's also some HP control panel, realtek and 'sound research' branded stuff, Fortemedia SAMsoft effects(?), Intel smart sound...
If I'm recording seriously, I usually just go to the device manager (devmgmt.msc) and disable as much as I can and enable devices in a trial-and-error way to see what the minimum is to get audio to work again. Otherwise, all kinds of 'enhancements' end up in the audio path.
(For those who may be reading this comment and wondering what the shit this software is, go give it a look. (I use it regularly and don't do Pro Audio stuff... I just want to be able to independently adjust relative volumes of groups of software (like Voice Chat and Video Games).) If the software's feature set looks intriguing, do give it a try and use Voicemeeter Potato. IIRC, all "flavors" of Voicemeeter have the same trial period, and I can't think of any reason to not use the the "flavor" with the most knobs and interconnects.)
And on corporate laptop you might not be able to disable its driver :/
Fortunately, it's too dumb to deal with devices that are behind more hubs than the root one...
PulseAudio was quite buggy when it was considered ready for general public by Ubuntu. I kept using some scripts to do what it could theoretically do for 2/3 years, and then it didn't have any quirks anymore.
Now that it is replaced by the pipewire stack, audio bugs are back. For instance, on a laptop if system audio is set to 'output' instead of duplex, switching between speakers and headphones does not work, and volume change would sometimes get stuck. And sometimes it duplicates the streams on another system, which can be solved by killing the daemon.
Even with this less stable state, I agree that this is nothing compared to how bad the situation is in Mac or Windows world. In the professional environments I've been in, basically all people have given up on the system properly switching their device parameters correctly. So calls often need a few minutes to reconfigure audio. Same with external screen switching.
I have had a Linux-based DAW running in my studio, alongside the requisite MacOS and Windows machines, for decades now. It runs Ubuntu Studio, has superlative audio performance (72 channels of digital audio), and is a rock solid workhorse for doing large edits on tracks.
The key to it is in using Ubuntu Studio, which is a well-tuned distribution focused on superlative Audio performance, and to choose your hardware wisely. In my case, its all Presonus - because they have been Linux-friendly for a long time - and it easily delivers latency numbers that outperform even the Mac in the room.
(Apologies for injecting advice into a grumbling thread)
And then there's the whole new suite of programs you have to learn where their implementation is constantly in flux so the documentation isn't exactly accurate. pw-record for instance. That "--list-targets" option it tells you to use is long gone (that princess is now in the "wpctl status " castle). You gotta check the date of everything written online because the month it was written matters. It's still far from great.
I used an Amiga about 30 years ago to do similar things. Now that was something that genuinely just worked. People are still using it. That's how functional it was.
But like all these things, I should find the motivation to shutup the complaining and get to cracking on the code to make it suck less.
When it comes to linux and things are broken, your assumption on how many people have seen it and who is working on the section of code it's caused by is invariably an order of magnitude or two too high. That's why you can't find any fixes on the web. You're one the first to see it. Exciting, isn't it?
Some distros don't have the latest pipewire stack, and you're left to fend for yourself, having to follow some incomplete or poorly written blog post to side-step what your distro does and put pipewire on top.
Me: Ubuntu 24.04 LTS, and happily using carla, ardour, lmms and a bunch of midi devices with pw-jack. (I'm aware pw-jack is not required anymore, but that was my old workflow, and it speaks to the backwards-compatibility that the stack offers.)
I have a counter-request: It's really frustrating when people are like "It just works! I had no problem! It's so easy!" in response to someone who has clearly struggled to get things working. It'd be like someone telling you they just had a car crash from a mechanical failure and you responding "Well I didn't! I drove home just fine!"
Instead, if you're going to respond at all, something like "I'm sorry, don't give up. I hope you figure it out" would be nice.
I think the stars just aligned for me, and it worked. Not an expert at all on these things.
Actually I've gotten everything to work adequately but claiming the alsa/pulse/jack to pipewire transition wasn't just a new nightmare would be wrong, well for me.
> "you need to use windows (or mac) to do anything serious"
Correct and maybe! Serious as in commercial or production? Yes!
Serious as in exploration in HCI and digital instrument creation and what kind of new sounds can come from that? Now we're back into Linux!
I try to (poorly in my opinion) explore the uncharted, I'm not really looking to make a single penny.
Take for instance, the classic chaotic pendulum (https://m.youtube.com/watch?v=yQeQwwXXa7A), you can hook that up to an Arduino sensor pack and convert the values to midi notes that get piped through a synthesizer or they can be the filter control of the synthesizer.
How can the arrangement of the chaotic magnet surface affect the aesthetics of the sound?
For instance, if you hook the values up to a sequencer playing arpeggiators and limit the chord choice wisely, it kinda sounds like bach. Especially if you do time dilation and don't try for things to be real-time.
Recording a composition is a sequence of 2D diagrams with time signatures.
Here's another one, this time with synesthesia. You take a number of sticky notes in various colors and aim a camera at a wall and then assign different roles and rules to the colors and their adjacencies and do a similar pipeline but this time you're playing a concert by sticking post-its to a wall on top of each other.
And yet a 3rd. You take a couple hour capture of rush hour from a freeway traffic camera and assign instruments to the lanes, scale signatures to their densities and then you can hear an orchestration of Friday traffic.
In all these you're still "playing" music because you're taking an active role in a bunch of aesthetic decisions and constraints, it's just a new relationship.
Maybe I'll do a write-up on the setup to make it easier.
Really I look at things like musique concrete, BBC radiophonic workshop and John Cage and think that's where, a few decades removed, all the modern sound has come from that's been so dominant for the past 40 years or so.
I want to embrace the new and weird so 30 years from now I can help be a historical part of building whatever is coming next.
Tomorrow's all we got!
> this isn't a distribution problem
Exactly. This all shouldn't be distro dependent when the only things involved are ALSA (kernel) and pipewire/pulseaudio.
Both unfortunately have distros do "things" to them all. Some default to dmix, some to direct hardware with PW/PA dynamically spawned and thus getting exclusive access for a single Unix user while the others are SOL, and other painful conundrums because they thought "we'll just do this and then it just works" for a single basic use case.
Use a Linux distribution that is intended for professional audio use.
>Once I hook up some midi devices, want things to be recorded, run a synthesizer stack through pw-jack, expect the midi clock to go down the USB bus, etc, all kinds of interesting behavior starts happening.
On Ubuntu Studio [0]/[1] - this Just Plain Works™, you know. I've been doing exactly this for years on my Ubuntu Studio machine, and I just don't have any of the issues you've encountered.
[0] - note that ubuntustudio is a metapackage you can install on most Ubuntu instances, which will set up audio for professional use.
[1] - see also, Zynthian: https://zynthian.org/
Right now after coupling I have to restart mpg123.
PulseAudio constantly stands out as the most common frustration in all of Linux. systemd is pretty frustrating too, but it's only frustrating the 1% of the time when it breaks (because it means your whole system is broken), whereas the 70% of the time PulseAudio doesn't work, it's more mildly infuriating.
For anything except a single audio in and single audio out, I would not bet on a beginner-to-intermediate ever getting it working, and PulseAudio is probably the main reason I in 2024 cannot recommend Linux on the desktop for non-engineers.
I can't complain really, in last decade audio has been working really well on my computer globally bar the occasional bluetooth pairing/connection issue. But I've seen people struggling with BT on all OS/platforms anyway. Pulseaudio was a PITA in the early years but matured a lot and Pipewire has been flawless on my distro since it took over. And I can do a lot of stuff out of the box that requires third party tools on other system to do the same.
I have installed Linux on I don't know how many systems at this point. My desktop, my work laptop, other people's work laptops, etc. Audio just worked flawlessly on each and every one. I cannot believe that PulseAudio has issues 70% of the time. There's simply no way I've been that lucky to have never seen an issue.
every other layer is a coping mechanism and the plurality and divergence of the FOSS community responds in various ways: - Jack - PulseAudio - PipeWire
I am unclear why Jaroslav Kyocera chose to make ALSA single-client, but Apples CoreAudio multi-client driver model is the right way to do digital audio on general-purpose computing devices running multi-tasking OS'es on application processors, in my opinion.
Current issues this article does not address that actually constitute large parts of the "mess" of Linux Audio:
- channel mapping that is not transparent nor clearly assigned anywhere in userspace. (aka, why does my computer insist that my multi-input pro-audio interface is a surround-sound interface? I don't WANT high-pass-filters on the primary L/R pair of channels. I am not USING a subwoofer. WTF)
- the lack of a STANDARD for channel-mapping, vs the Alsa config standards, /etc/asound.conf etc.
- the lack of friendly nomenclature on hardware inputs/outputs for DAW software, whether on the ALSA layer, or some sound-server layer. (not to mention that ALSA calls an 8-channel audio-interface "4 stereo devices")
- probably more, but I can't remember. My current audio production systems have the DAW software directly opening an ALSA device. I cannot listen to audio elsewhere until I quit my DAW. This works and I can set my latency as low as the hardware will allow it.
this is the thing: more than about 10ms latency is unacceptable for audio recording in the multitrack fashion, as one does.
This is one of the major reasons why Linux accessibility sucks IMO.
Audio is one thing that you need to "just work™" if you want to get accessibility right, as there's no way for a screen reader user to fix it without having working audio in the first place[1]. On Linux, it does not "just work", and different screen readers have different ideas on how they want audio to be handled. In particular, the terminal Speakup screen reader (with a softsynth) wants exclusive control of your device through ALSA IIRC, while the Orca screen reader for the GUI goes through Pulse. That makes it impossible to use both of them at the same time.
[1] Well, you can sort of fix it by having a second machine and SSHing into the broken one, but that's not what I mean.
Also, Accessibility != Audio. I, for instance, use Braille only. No need for speech synthesis. So equating Accessibility issues wth the crazy audio stack is a little bit too simple.
If you have Pulseaudio or Pipewire, they add a plugin to ALSA library that reroutes audio to audio daemon, so ALSA applications should work correctly.
I mean, I remember this being the case for a very long time on windows with ASIO too, which is the only reasonable way to run a DAW with acceptable latency there. MacOS has multi-client but I was never able to get latency as low as fine-tuned windows and Linux systems, and in the end that's what matters - you just use your motherboard's chip for OS audio and your pro soundcard for the actual workload. Pipewire is very close to giving a good experience but there'll always be some overhead - I'm making some art installations running various chains of audio effects on a raspberry pi zero and the difference between going through pipewire even if my app (https://ossia.io) is the only process doing any sound, and going straight to ALSA, is night and day in terms of "how many reverbs I can stupidly chain before I hear a crack".
Apps should use the right API from the right layer; when they skip something, no wonder they will miss whatever the skipped layer provides. When they do not need exclusive access to the device and want to play nice with the other apps, they should use pipewire/pulseaudio.
For 99% of apps, using ALSA directly is the wrong approach. You don't use IOKit directly in Mac apps either.
Applications want to receive/provide a stream (X sample-rate, Y sample format, Z channels) and have it routed to the right destination, that probably is not configured with the same parameters. Having all applications responsible for handling this conversion is not doable. Having the kernel handle this conversion is not a good idea. The routing decision-making needs to be implemented somewhere as well. Let's not ignore the complexity involved in format negotiation as well.
The scenario of a DAW (pro-audio usage) is too specific to generalise from that. That is the only kind of software that really cares about codec configuration, latencies and picking its own routing (or rather to let the user pick routing from the DAW GUI).
Because ALSA is a different layer in the audio stack than CoreAudio.
ALSA corresponds to MacOS drivers and I/O Kit.
CoreAudio (Audio Toolbox / Audio Unit) corresponds to Pipewire / Pulseaudio.
But on the Mac side everyone is OK with using CoreAudio (with the accompanying set of daemons), while on Linux, for some reason, everyone wants to go as low-level as possible, "just open the device file" and is wondering, why something is missing. Because you skipped that, that's why.
[1] https://learn.microsoft.com/en-us/windows/win32/coreaudio/us...
1. I have to modify my audio settings every time I start a call in Teams on Linux because it keeps losing my audio device.
2. In my audio settings UI, half the time I switch my devices the speaker test doesn't work.
3. In my audio settings UI, whenever I switch my mic I hear myself. The mic feedback only disappears 30 seconds after I close the settings UI.
4. My work headsets have a robotic sound (likely caused by an incorrect bitrate or buffer size). I can only use work bluetooth headsets via their dedicated dongle.
This was my default experience on a popular debian based distro. And it mirrors the general experience I see online. Things are unstable and a mess.
I started reading this article and it's embelished with phrases like: "is a professional-grade audio server", "widely used in professional audio production environments", and general language that sounds like a sales pitch. This does not fit with anything I'm familiar with.
I would have preferred a neutral and semi technical approach, with 10% of the buzzwords. As written, I trust nothing.
Well you probably didn't mean it in a literal sense, but it was OSS up to kernel 2.4 and 2.6 had both OSS and ALSA
And not much has fallen out of use. OpenAL, libao, jack, portaudio, libcanberra, Gstreamer and phonon at least are still used widely, and a bunch of others keep cropping up occasionally in cross-platform software.
If there's any difference it's that under Windows, the only debugging and error handling you have is "lol reinstall all drivers and codecs and pray".
So from that list the unification is now:
- JACK and Pulseaudio are both replaced by pipewire
- OSS is long gone, ALSA is now the only low-level interface
- ESD, NAS, ClanLib, xine, portaudio, allegro and Phonon are not present in the Ubuntu install I just checked
Basically we're down to a unified stack that has BlueZ and ALSA to access actual hardware and pipewire as the single audio (and video) daemon. Everything else is either shims so apps don't need to change interface or cross platform APIs like SDL and OpenAL. We are much better than what this diagram shows.
In practice, OSS and Jack still stick around, as do portaudio, libao and others.
Fact is, RT audio is hard, and the peoplebehind JACK have cared for the underlying problems for a long time already.
Maybe PipeWire, but to be honest, it reminds me too much of PA.
I guess I will stay with plain JACK and SuperCollider as my toolbelt, and not care about PA or PW. Like the grumpy old hacker I am.
> ...
> Ulike PuleAudio and JACK, PipeWire does not require ALSA on a system, in fact if ALSA is installed the output of ALSA is very likely pushed through PipeWire
I don't get this part. If ALSA represents the kernel level hardware drivers for audio, how does Pipewire bypass it? Does it implement an alternative set of kernel drivers? I assumed Pipewire still relies on ALSA base.
I think what they're getting at is that PipeWire speaks the ALSA API, so an app or game that linked against ALSA will connect to PipeWire and should Just Work, without needing to be rewritten to target PipeWire's API.
PipeWire does the same trick with the PulseAudio API as well. on my PipeWire-using NixOS box, for example, I can connect through the `pavucontrol` GUI, my `pactl`-based keybindings work the same, etc. it's a clever design that allows them to avoid what would otherwise be a nasty pile of backwards-compatibility issues and poor desktop user experience.
Pipewire uses ALSA to interface with audio hardware. But Pipewire also adds a plugin to an ALSA userspace library so that audio from apps that use ALSA API is rerouted to Pipewire. Pulseaudio did the same trick.
Also, given that Pulseadio and Pipewire both support ALSA clients, does it mean that the preferred API for applications should be ALSA? This way they can play sound on any system, even where there is no audio daemon.
I have the latest kernel version now.
What vendor and model laptop was it? I'll make a note to avoid them in the future.
(Bitwig, Ardour, Reaper, more? I would like to see "Input 1" or "Channel 1" and not some strange ciphers when trying to assign things in a little dropdown selector in a DAW)
[1] https://pipewire.pages.freedesktop.org/wireplumber/index.htm...
That does not instill confidence in that they have any idea what they are doing.
I remember that as far as on 2012 some Linux game ports from Steam (before Proton was a thing) failed to play sound.
Other than that, it took them a decade to figure out that sound should switch to HDMI when it is plugged in. It may still require arcane config changes and may break down.
I've opened the post to read about pipewire, and it seems that clicking an anchor does nothing. So it's not only sound they can't get right.
JACK is very flexible approach to audio (think of unix pipes but audio) and this is what pipewire is inspired by, but targets the general desktop use and not just audio production as JACK does.
Generally pipewire stays out of the way and works as pulseaudio replacement, but with the pipelines you can do a lot as a power user.
Personally I think pipewire is one of the best things that has happened to linux, and since it shims all the previous audio apis instead of trying to tape the existing solutions into the package, it actually removes bloat from your system and you don't have to write weird libasound config files anymore.
Not sure how well the NI Native Access crap is working on WINE either. Does anyone know?
[0] https://appdb.winehq.org/objectManager.php?sClass=version&iI...
By itself it's essentially grand unification of audio servers that actually works better and is way less... opinionated about the only true way some things works, which was a problem with pulseaudio at times.
But if anything it's X11 that tries to lump a ton of stuff together. Wayland is very minimalistic in comparison. That's why stuff like libinput and Pipewire are part of making a functional desktop.
More over, every "WM" in Wayland's case needs to implement the entire stack, even if it uses a common library for some of it.
And after similar length of development time, I'd say the result is still worse in many aspects than X11, and I say that as someone both using and praising a wayland-based compositor and lamenting that it pretty much locks me more than Windows used to
But yeah WSL is just too easy. Specially with native vscode support it has become a favourite for many developers
Bluetooth on the other hand...
what do you think it will happen?
Honestly I would just buy a USB interface. I got a Focusrite Scarlett 2i2 and have been very happy with it.
Do you have any more information on how can it be diagnosed?
The motherboard it does have the latest bios update.
The year of the Linux desktop never arrived but most of the world has Linux sitting in our palms. Let's build on that instead of the dead ends of Debian, Slackware, etc.
I can only take your comment as either a joke, or you literally only use a phone/tablet in life.
news flash: Android audio is ALSA at it's core. Additionally: native Android audio is ultra-high-latency... completely unacceptable for pro-audio