PulseAudio under the hood
gavv.github.io
gavv.github.io
1. We have multiple incoming RTSP stream where the audio is added as pulse audio streams.
2. We mix a couple of these streams.
3. We send this stream to an external echo cancellation device.
4. We take back the output of the echo cancellation device.
5. We mix some more streams into it
6. We route this to a certain output, which can be dynamically changed according to a certain algorithm.
We control PA using DBUS and the above scenario is accomplished with only a few calls. Granted, there's some latency you can't really control, but all in all it's an incredible powerful system that's been working really well for us.
Because my reaction when I read the linked article was that pulse sure does seem to have a lot of features that sound pretty useful, but after years and years of nominally having it on my system in charge of my audio hardware, I have, to be blunt, absolutely no clue how to use any of those features, except a bit about network audio streaming. None whatsoever. Didn't even have a hint the support for those autoconf protocols existed.
I did once try to stream a Pulseaudio stream from one linux system to another. I never got so much as a peep to go across. I have no idea where any diagnostic messages may have gone that might help me debug this. I have no access to a working configuration to compare to. I have no idea what I did wrong and no idea how to fix it. I can find documentation that assumes you're writing C extensions and I can find documentation on how to slide sliders in the Gnome volume app, but I can't find anything in between. And I'm a professional software engineer with a minor in electronic music; there are certainly people with more knowledge and more motivation on these matters than me, but on the flip side, features that I can't figure out or get working might as well not exist to the vast majority of the population.
So while I'm not asking you to write a manual page, how exactly did you get to the point where you can do that? Again, I'm serious; the previous paragraphs are not to just slag on a project, but to show where I am and what problems I'm having. I've got a project I'm trying to do where it would be really useful to understand these things, but I have no idea even where to look at this point. (I've poked in every Google search I can think of and gotten the aforementioned C-API level docs or "I just installed Ubuntu, how do I volume?" docs.)
One of these modules that come with any standard distribution of recent-ish PA can use avahi/zeroconf (probably in part because avahi was Lennart P.'s previous focus of attention, before he started work on PA) to advertise PA sinks on the network. As a consequence, running the avahi daemon is required for PA network-wide advertising of its presence to work, and auto-discovery to have any chance of working at all. Another module implements accepting PA-via-TCP compliant clients via TCP. Both are included in PA's standard distribution, if it's recent-ish (think 2012 and later). Distros still might package them in extra packages; the Debian does so for the zeroconf-parts in a package named "pulseaudio-module-zeroconf".
To load the TCP transport support module and specify a primitive ACL (based on source IPv4 addresses - you need to adapt this ACL to your local environment if it uses another IP range!), you can use this `pacmd` stanza:
load-module module-native-protocol-tcp auth-ip-acl=127.0.0.1;192.168.0.0/16 auth-anonymous=1
(A presumably more secure, cookie-based authentication mechanism with a shared secret involved is also available, but I've never used it.) This will make pulse open a listening TCP socket on port 4713. At this point, remote hosts with PA-ready applications should already be able to play back sound by setting the PULSE_SERVER environment variable to the accepting PA server's address - what's still missing is the avahi-based advertising of the service.So, to make PA contact your system's avahi daemon and use it to advertise its sinks on the network, you load another module into your PA instance:
load-module module-zeroconf-publish
If both these operations have succeeded (`pacmd` will complain loudly if they don't), a quick `avahi-browse -a` on a (avahi-enabled) host in the same network as the one with your exposed PA server in it should yield something like this (pasted from my local network): + enp0s31f6 IPv6 pulse@nas PulseAudio Sound Server local
+ enp0s31f6 IPv6 pulse@nas: Jukebox PulseAudio Sound Sink local
+ enp0s31f6 IPv6 pulse@nas: Dummy Output PulseAudio Sound Sink local
+ enp0s31f6 IPv4 pulse@nas PulseAudio Sound Server local
+ enp0s31f6 IPv4 pulse@nas: Jukebox PulseAudio Sound Sink local
+ enp0s31f6 IPv4 pulse@nas: Dummy Output PulseAudio Sound Sink local
+ enp0s31f6 IPv4 root@tv PulseAudio Sound Server local
That is two hosts advertising PA sinks on the network - "tv" (running libreelec) and "nas" (running Debian).Now, when you make another avahi-enabled host on the same network load the "module-zeroconf-discover" PA module, the advertised sinks should magically show up in that PA instance's list of available sinks. You can then move streams onto them with any of the standard utils, like pavucontrol.
If you've found a set of stanzas that set up your PA instances to your liking (with modules loaded the way your setup requires it), you can make them re-apply on daemon startup by persisting them in your user-specific rc files in ~/.config/pulse/.
Hth! :)
Instructions here: https://parseq.co.uk/wordpress/archives/setting-up-pulseaudi...
It worked, but sadly, PulseAudio for Windows doesn't support 5.1 sound, only stereo.
The other answer is way better than mine, but I will add what I also said elsewhere in this thread: In 99% of cases, Pulse just works. In 1% of cases, it fails in catastrophic, bizarre and utterly undebuggable ways. I have the luck of being in the 99%.
I just installed and started Avahi and selected "Make network sound devices available locally" (or whatever the option is called) in paprefs.
That's for the client part. For the server part, see https://github.com/majewsky/system-configuration/blob/master... (that's for Arch Linux, package names etc. might differ between distributions).
I've never had a good experience with Pulse, never had it work out of the box, and can never find information about configuring it. It gets uninstalled on every machine I run.
I stick with plain ALSA. It's annoying to configure, but IME it works out of the box and I usually don't have to worry about it.
WHAT? That is crazy talk. JACK is the most complected audio setup I have ever worked with. I owned a Record Label and a small studio. If you have to setup JACK from scratch it can get crazy complicated. Once you get it working it is an AWESOME framework but I only would ever see people needing low latency to ever really use it.
JACK is 100% for a DAW (Digital Audio Workstation) and it is for low latency. Pulse is for the rest. I have found that the legacy of a poor initial deployment has been a heavy chain around its neck just like KDE 4, RPM (for having an issue for a few months over 10 years ago) and now Systemd (Systemd had a fine performance roll-out but philosophy issues).
/usr/lib/pulse-11.1/modules/module-jack-sink.so
/usr/lib/pulse-11.1/modules/module-jack-source.so
/usr/lib/pulse-11.1/modules/module-jackdbus-detect.so
I can use pavucontrol with ease to redirect an application's audio output (e.g. chromium's) to the jack sink.It's a painful process to accomplish the same thing without using pulseaudio, because you're going to have to mess with ALSA configuration files- and believe me, that's no fun at all.
In other words, I don't need no stinking' "game/consumer Audio", which is what Pulseaudio represents to me ..
Therein lies the beauty, I guess, of Linux - and of F/OSS creative-tools, in general ..
In that case, you can also chose hardware that works well with PA, so I don't see JACKs benefit there. Also, I don't think they're equivalent in functionality.
However, its moot - this is not consumer audio. Its more, "studio production audio"...
You probably won't have to. I think the long term plan is for Pipewire to replace PulseAudio (and be an alternative to JACK). Pipewire also does video. The initial release of Pipewire is video only with audio support to come:
https://blogs.gnome.org/uraeus/2017/09/19/launching-pipewire...
Why not fix what we have rather than yet another rewrite with all its attendant exciting bugs.
So, just as a counterpoint, I'm guessing you're a Pulseaudio user ..
No that's not it at all. It's a framework that will use Pulseaudio OR Jack. Jack is only for low latency needs AKA Professional audio production.
"Pipewire is used to build a modular daemon that can be configured to: be a low-latency audio server with features like pulseaudio and/or jack.
Need to select it as default when plugged and whenever it gets replugged it retake its priority. Works as expected with multiple default / temporarily plugged devices.
Interface might need a bit of a clean up though.
It is not possible to change the power settings and it was removed by the GNOME developers[0] (the shutdown/power off action was considered too destructive[1]).
Bottom line: You can no longer power off your laptop by pressing the power off button.
[0] https://bugzilla.gnome.org/show_bug.cgi?id=755953
[1] https://github.com/GNOME/gnome-settings-daemon/commit/69d9d8...
Oh how I miss the simplicity and elegance of before PulseAudio (and its evil twin "hi-let-me-put-your-text-logs-in-a-binary-format"-systemd) appeared in ArchLinux.
I've primarily run Linux over the last fifteen or so years and barely had an issue that's directly PA's fault.
I remember the bad old days where a playlist would finish and then GAIM's notification sound would bleat fifteen times. PA saw an end to that.
I don't hate PA. On another machine it was always installed. But it is quite clear to me that it is yet-another layer on top of the existing infrastructure, so it statistically will cause more issues than just the infra that's there already.
Pulse, on the other hand, while working fine on 99% of systems, usually fails catastrophically in creative ways on the remaining 1% and it's really hard to even understand what the problem is. (And of course, these 1% will shape public opinion. A large portion of the 99% is not even aware they're using Pulse.)
Factors which will have majorly affected people's PulseAudio experience over the years include audio hardware, what software they use with it, system uptime, local RF interference, use of suspend/resume, which patches their distro backported, and the exact stack and heap layout of the process on their system.
But really, fix your PulseAudio instead.
This similar to how on paper Firefox can still be compiled with GTK2, but it is not supported.
Then PA is not the problem, your sound card/drivers is.
https://fedoraproject.org/wiki/How_to_debug_PulseAudio_probl...
PulseAudio emulation for ALSA https://github.com/i-rinat/apulse
[0] If I remember correctly, this was due to driver issues, but somehow PA just handled it.
Older versions supplied dmix but didn't automatically enable it. Newer versions are supposed to autoenable it on audio devices without hardware mixing.
Frankly the only thing PA brought to the table was the ability to handle multiple audio devices while there is ongoing audio IO. Primarily in relation to usb headphones/headsets. But then that class of hardware is its own kind of abomination in the first place.
I wonder of Linux's biggest problem these days is the amount of "herd knowledge" that is floating around that is obsolete to say the least, yet is being used to justify some dev's latest weekend glory project.
I never had many problems with OSS either. :) I'm bringing up esd because I was misrembering that when pulseaudio came on the scene, it seemed to be already existing, not screwing up too badly, and addressing the commonly held up use case of software mixing.
After reading notalaser's fine summary [1], I'm realizing I've lumped 'ALSA broke everything that I already had working (but did eventually work, and support newer hardware)' and 'PulseAudio broke everything that I already had working' together in time in my mind, when really things were many years apart. It seems that by the time PulseAudio came out / was widely pushed, ALSA had already subsumed esd (unless you were using the network audio features), so the most often reported 'great thing' doesn't sound like something people wouldn't have had.
PA gives me the ability to have multiple pieces of software making noise at once, a software mixer with per-application and per-sound-hardware volume control which is completely divorced from those applications (so if I say MUTE, it is damn well MUTED), and the ability to choose where the sound is being output to. I cannot imagine anything doing better on the hardware I have.
Perhaps not coincidentally, the primary sound API on FreeBSD is still OSS.
I too used Linux at that time and I remember that problem. The workaround in those days was to use esound. IIRC it also disappeared by the switch to alsa.
FreeBSD also doesn't have this problem, and it's still using the OSS api. Let's not confuse API with implimentation.
Because of this setup I need full-range stereo output on my headphones and low-passed LFE for the transducers.
How can I do this PulseAudio? When I enable lfe-remixing, I get full-range output on my transducers and I hear voices in my chair. If I set lfe-crossover-freq, I lose the bass on my headphones.
I found this patch
https://pw-emeril.freedesktop.org/patch/171424/
which seems solve the problem, but it was rejected.
I would comment directly on the site, but because the certificate is invalid, I cannot register an account.
I don't like the attitude to reject a patch just because one can't come up with a use-case. Obviously there is a use-case, else the patch wouldn't have been developed.
Also, the feature would be hidden in a configuration file that is meant for advanced configuration anyways.
On top of that, I have a AV-receiver that does support all of this bass management. For every channel I can configure individual crossover frequencies, and I can but don't need to use a subwoofer. Why doesn't PulseAudio give me these freedoms?
Why would a free software developer deny others easy ways, or any ways at all, to configure their system? This is equivalent to the walled-garden approach of many closed systems. Why does this occur in free software?
Accepting a feature that a few people might use, hidden away somewhere in the advanced configurations, might very well not be worth the trade-off to them, even though you personally find the feature valuable. It has nothing to do with a walled garden.
Ultimately, if this is so important to you, you have the choice to fork and maintain it yourself and carry that cost. That'll most likely be a lot of work, especially if you want to keep in sync with upstream. But it might give you a better understanding of and appreciation for the tremendous amount of work and effort that goes into projects like these. Work that goes largely unpaid.
Because of the "drive-by contributor" problem. Come by, drop off a patch that solves exactly your problem, never show up again to maintain the feature, even if it breaks, even if it's been broken for years.
That forces me, the developer, to either continuously test your patch (which I may not be able to do if I don't have your hardware/software/configuration/use-case-in-mind), or to shrug and hope it works. We have to just trust your initial work and have it no way weigh in on the future of our product?
We've learned too much as Open Source Software Engineers to just trust random drive-by contributions. We've gone from default-accept-all to default-deny, because at the end of the day that's what keeps quality in our software.
That is wrong, if I don't misunderstand it. It does not matter much for the possibility of GUI tools whether they talk to an API or whether they parse and write configuration files. That is stuff you abstract away.
It is of course possible that pulseaudio allows settings to be set that didn't exist before. It's API might be better for that - but that doesn't say you couldn't do the same with writing to configuration files if they had the same capabilities, like targeting a specific application (one of the use cases he mentions below).
PulseAudio needs to allow UI tools to set properties with sub-second latency.
And just reparsing the configs every time something is changed would be extremely wasteful.
Once upon a time, sound cards were files in the /dev tree. To play sound, you wrote pcm data to the file representing a sink. To record, you read pcm data from a file representing a source. Things were better then. I'm sure there are people with use cases that have required the four (and counting!) solutions crufted on since then, but I've never been one of them, and it irks me that the interfaces get more and more complex and brittle with each iteration.
I banged my head against my desk for a couple of days when Slackware switched to Pulse with 14.2, allegedly because it was needed for bluetooth. I still have difficulty believing that people actually use bluetooth for audio, but apparently some people love it.
Really? So you want all app feedback to cease if you're playing music? No IM or new email notifications? That just seems like a pain to me.
Multiplexing can be handy if you like gaming, and having music from another application at the same time.
Also sometimes it's useful to have a reference manual in one side of the screen and your text editor in the other, or a PDF viewer on one side and a LaTeX editor in the other.
> Also sometimes it's useful to have a reference manual in one side of the screen and your text editor in the other, or a PDF viewer on one side and a LaTeX editor in the other.
True. An emacs which supported the framebuffer could do this (but GNU emacs currently only supports vt100, X, macOS & Windows, IIRC).
So you don't believe that somebody would like to be able to use wireless headphones?
I am completely and totally flabbergasted by their popularity. It's as though millions of people were turning their noses up at fine, free homemade French food and paying for McDonald's instead.
The same as wifi, right? Strictly worse in every way. Except for the not needing wires way.
Bluetooth was originally encrypted, but gained the option of non-encrypted channels with version 1.1.
On top of that Bluetooth use frequency hopping.
Until recently there was no real hardware for scanning/sniffing Bluetooth traffic, at least not anything easily available compared to a wifi card in promiscuous mode.
Wifi: The internet comes from one (inconvenient) place in the home that I can't choose. I've got a couple of powerline ethernet adapters, but I've also got a lot of devices that handle wifi, but not ethernet. The choice is between using wifi or skipping the network connection. It's great as a last- (or only)-resort connection option, and it's usually sufficient, but then my uses for it are usually pretty tame anyhow.
The bluetooth stack dropped support for ALSA with Bluez 5.0. About a year ago, dev started on a "bluez-alsa" or BlueALSA package the re-integrates an ALSA backend. I've used it briefly; it seems to work as advertised.
I think that Bluez 5.0 came out about the time that Slackware 14.0 did, so it makes some sense that versions before 14.2 would've provided Bluez 4.x for the bluetooth stack.
I still remember buying a hardware sound card (soundblaster emu10k) to enable ALSA to mix multiple sources (game audio + teamspeak or whatever voice chat people used back then) back in the early 2000s.
* default Alsa setup: set master to level that seems comfortable-- i.e., "to taste"
vs.
* default Pulse setup: set Chrome's app level to taste
Bug: packet.city plays an annoying sound with a perceived loudness that the user did not predict.
* Alsa solution: lower the master volume to a level that doesn't result in painful output from packet.city, and adjust all other apps accordingly
vs.
* Pulse: lower Chrome app audio, but leave yourself open to similarly unpredictable painful events from other apps
It seems in this case that the limitations of Alsa actually force the user to do the safe thing wrt perceived loudness.
If you trust VLC and you trust the video file you're watching, then turn the master audio back up during the loud "Universal Pictures" intro or whatever is an equivalent introductory sound at the beginning of the movie you're about to watch.
On my chromebook, there are two volume buttons and one mute button that are by far the most convenient way to adjust (master) audio for such a situation. So you turn down for random browsing and turn up for watching movies or whatever.
Btw-- those are almost certainly the controls someone would use to turn down or mute the sound of that horrible website audio listed above. In such a case the fact that pulse can give you an app volume for Chrome doesn't help anything-- for example, the user who adjusted Pulse app volume down to avoid pain will need to adjust it back up when watching a movie on Youtube for the same reason you mentioned. But they can't adjust the app volume with the keyboard controls. So to make use of that feature forces a more complicated UI that takes longer to use.
In this case, it doesn't make much sense to treat the browser as a single application, but each site as one (does Chrome export each tab as a separate source to PA? IDK, I haven't checked). And, I can't think of a single thing that I actually want producing sound (and don't mind just muting) from my web browser that doesn't have its own volume control.
Now, it's just anecdotic evidence of course but basic audio mixing is not exactly rocket science, you shouldn't need something as convoluted as PulseAudio to get it right. I mean, what the hell: https://gavv.github.io/blog/pulseaudio-under-the-hood/diagra...
Some power users might need these additional features, I don't. I wouldn't mind having all of that if the basics just worked reliably but they don't.
These days disabling PA is my go-to first step while troubleshooting audio issues on linux and more often than not I don't have to go to step #2.
ALSA refused to do mixing in software for some time. Maybe it still doesn't do it, I don't know. So if you had a hardware sound card that would do mixing, you could play multiple channels with ALSA, but if you didn't, you were SOL.
Not sure why that deserves downvoting.
This has worked well in ALSA for years now. Now, ALSA won't let you set the volume of each application separately, but most applications have their own volume control for that anyway.
And you'll never have to deal with the application's volume Pulse's volume for that source disagreeing, and scratching your head as to why even though you set the volume to max, it isn't getting very loud. To be fair, maybe this has been entirely solved; the last time I used Pulse was ~5 years ago[0], so I might be a bit out of date. But not more out-of-date than the claim that ALSA can't play sound in two applications at once!
[0]: I don't have anything against Pulse, I just don't see the point in installing it. What does adding it to my stack get me?
Do wonder how much future hearing loss it has created by sudden spikes in volume...
I wonder if they would have ever implemented this "feature" if it wasn't the default (and only) behavior on Windows...
This by attaching a bunch of peripherals and use software to assign them to groups that effectively form the equivalent of a graphical terminal.
All this supposedly in the hopes of introducing affordable computing to schools in the developing world. But said schools have already tested and discarded this idea, so why the DEs keep pushing it is a mystery.
Seats never made much sense to me and if the smartphone story has any truth to it, then I do not see any valid use case for that idea these days.
Once I started using PA, I kind of went crazy with it and started trying to add PA support to old linux software that was still ALSA only. The library and docs were not hard to use. The PA framework is all reference counted as well so when you create the pa-objects in C, you never really need to worry about memory allocation. It was pretty neat. Unfortunately, before I had all my commits and patches ready, I changed jobs and never got the free time to commit them the rest of the way.
Just give me files I can write data into and suddenly everything becomes very easy. Everybody knows how to write files.
I know this discussion is getting really old but when Alsa was a step in the wrong direction, Pulse was the highspeed train in the wrong direction that followed, trying to "fix" it. This overengineered mess is the result of the tragedy allowing former Windows users to do systems programming for Unix like operating systems.
This is on ArchLinux running KDE4
Does anybody know if there's a similar such document for ALSA similarly written by someone who isn't an ALSA insider?
I'm not sure whether that's due to distribution maintainers giving very sane configurations or that pulseaudio development has been focusing on sane defaults.
Eitherway, if you're one of the people involved in giving me a audio experience on Archlinux/Fedora, please pat yourself on the back because you're doing fine work.
I can also use old devices that don't work on Windows, anymore, of which I own several (a keyboard and an audio interface I'd planned to get rid of, but don't really need to now that I can do most things under Linux). Which is kinda funny...now Linux is more likely to work flawlessly with a given piece of audio equipment than Windows.
If only an official REAPER port would come along (I know it'll run under WINE, but there's also been rumblings of a Linux port for years, and I'd much rather run something that works natively)...
There do appear to be official but as-yet unsupported Linux builds: https://wiki.cockos.com/wiki/index.php/REAPER_for_Linux
I had cause to try one for something small recently and it worked pretty well (including with PulseAudio). No idea how stable it is for serious use though.
[0]: https://blogs.gnome.org/uraeus/2017/09/19/launching-pipewire...
>"We know that for many the original rollout of PulseAudio was painful and we do not want a repeat of that history." //
There's some hope!
AudioTrack, OpenSL, whatever-the-Android-8-thing-is-called?
From the sentiment here I get the feeling that it might stop working any moment.
The reality now is it is very stable and has a huge number of powerful features. Personally I think it is great, and have no issues with it.
I chose the third way of dropping Desktop Linux. I just don't have time for that kind of crap anymore.
Haha, that is ridiculous!
I started really using it when I started using Linux laptops, just because it's way easier to connect bluetooth audio, ship audio over HDMI, etc.
It actually works pretty well now and I like it. Still hate and refuse to use systemd though. :-P