Real-time audio programming 101: time waits for nothing (2011)
rossbencina.com
rossbencina.com
One practical reality it doesn't share is that your audio processing (or generation) code is often going to be running in a bus shared by a ton of other modules and so you don't have the luxury of using "5.6ms" as your deadline for a 5.6ms buffer. Your responsibility, often, is to just get as performant as reasonably possible so that everything on the bus can be processed in those 5.6ms. The pressure is usually much higher than the buffer length suggests.
A bus is a shared medium of communication[1]. Often, busses are time-division multiplexed[2], so if you want to use the bus, but another module is already using it, you need to wait.
For example, if your audio buffers are ultimately submitted to a sound card over a PCI bus, the submission may need to wait for any ongoing transactions on the PCI bus, such as messages to a graphics card.
[1]: https://en.wikipedia.org/wiki/Bus_(computing)
[2]: https://en.wikipedia.org/wiki/Time-division_multiplexing
You might sometimes build an app where (through your operating system) you connect directly with an input device and/or output device and then do all the audio processing yourself. In this case, you'd more or less control the whole bus and all the code processing samples on it and have a fairly true sense of your deadline. (The OS and drivers would still be introducing some overhead for mixing or resampling, etc, but that's generally of small concern and hard to avoid)
Often, though, you're either going to be building a bus and applying your own effects and some others (from your OS, from team members, from third party plugins/libraries, etc) or you're going to be writing some kind of effect/generator that gets inserted into somebody else's bus in something like a DAW or game. In all these cases, you need to assume that all processing code that isn't yours needs all the time that you can leave for it and just make your own code as efficient as is reasonable.
> In all these cases, you need to assume that all processing code that isn't yours needs all the time that you can leave for it and just make your own code as efficient as is reasonable.
Yes. For audio programmers that is obvious, in particular when it comes to plugins, but for novices it might be worth pointing out!
In case it is not clear, that is the primary case that is addressed by the linked blog post (source: I wrote the blog post).
I would love to see a modern take on the real-world risk of various operations that are technically nondeterministic. I wouldn’t be surprised if there are cases where the risk of >1ms latency is like 1e-30, and dogmatically following this advice might be overkill.
Thus such micro optimizations are seldomly used. Quite the opposite, you try to avoid jitter which could be the result of caches
And then there's Diva at its highest output quality setting... :)
But very little digital audio gear works that way these days. The buffer sizes may be small (e.g 8 or 16 samples), but most hardware uses block structured (buffer by buffer) processing.
So there's always a delay, even if 1 sample.
This has been slower for most things that raw computation for well over a decade (probably more like two).
You are not saving a sin table, but very complex differential equations.
[fixed typo]
Indeed, like all real-time systems you need to think in terms of worst-case time complexity, not amortized complexity.
I tend to agree, but...
From my recollection of using Zoom-- it has this bizarre but workable recovery method for network interruptions. Either the server or the client keeps some amount of the last input audio in a buffer. Then if the server detects connection problems at time 't', it grabs the buffer from t - 1 seconds all the way until the server detects better connectivity. Then it starts a race condition, playing back that amount of the buffer to all clients at something like 1.5 speed. From what I remember, this algo typically wins the race and saves the client from having to repeat themselves.
That's not happening inside a DSP routine. But my point is that some clever engineer(s) at Zoom realized that missing deadlines in audio delivery does not necessarily mean "hosed." I'm also going to rankly speculate that every other video conferencing tool hard-coupled missing deadlines with "hosed," and that's why Zoom is the only one where I've ever experienced the benefit of that feature.
2. 3ms is typical in-air latency between a typical DAW user and their near-field monitors, so claims about sensitivity to times much lower than 5msec should be taken with some skepticism
3. In live contexts, many drum + bass pairings have more than 10ms of air latency between them, so ditto #2
4. On the other hand, no good reason to add to latency
5. For performance purposes, jitter is much worse than latency. Pipe organ players rapidly learn to deal with even whole seconds of latency, but almost nobody can deal with jitter (essentially, variable, unpredictable latency)
6. There are no sub-ms issues that will cause phase and frequency distortion. Those come from DSP errors, not handling of latency, which is just about always a constant, fixed feature of the data signal path. You may be thinking of stuff like comb filtering, but this is not related to the latency in the signal path in a correct setup.
What started off as a four note chord would be smeared out a little by MIDI, especially in the early days until everyone worked out that putting MIDI for an entire studio down a single cable was a bad idea.
Then you'd get some more smearing in the target synth CPU as the incoming notes were parsed. Then perhaps some more delay for each notes, because it took a while to send trigger and pitch messages to the hardware. Even more if there were if there were software envelopes involved and they had to be initialised.
This is still a problem with VSTs, on a smaller scale. There's some finite amount of processing that has to be done before sound starts being generated. Usually it's not very much, but there's always the possibility that two notes that should start in the same 5ms buffer slot will be spread across two of them because one note is just a little too late.
This isn't as objectionable as glitching, but it can still affect the timing feel, and - depending on the patch design - cause phasing effects between the notes.
2. "parsing incoming notes" does not cause more smearing. Block-sized processing of audio causes a delay which is the "performance latency" that people complain about. It does not change the ordering or interval between note onsets.
3. the "finite amount of processing that has to be done before sound starts being generated" is irrelevant in a block processing architecture (which is used these days by all DAWs and all plugin APIs). As long as the plugin gets its work done within the time represented by the block,there is no additional latency caused by the plugin. If it doesn't, then there's a click anyway.
4. "there's always the possibility that two notes that should start in the same 5ms buffer slot will be spread across two of them". No, there isn't, If that happens, that's a coding error in either the plugin host or the plugin or both. But also, time is continuous. If the notes are supposed to be 3msec apart, it doesn't matter if they are 3msec apart within the same buffer/process cycle, or in two consecutive ones.
It depends on your appetite for risk and the cost of failure.
A big part of the problem is that general purpose computing systems (operating systems and hardware) are not engineered as real-time systems and there are rarely vendor guarantees with respect to real-time behavior. Under such circumstances, my position is that you need to code defensively. For example, if your operating system memory allocator does not guarantee a worst-case bound on execution time, do not use it in a real-time context.
I think in essence I'm repeating the comments of Justin from Cockos, which you summarize [1]:
> It is basically saying that you can reduce the risk of priority inversion to the point where the probability is too low to worry about.
In that comment you also say:
> 100% certainty can’t be guaranteed without a hard real-time OS. However 5ms is now considered a relatively high latency setting in pro/prosumer audio circles
Which I interpret as acknowledging that we're already forced into the regime of establishing an acceptable level of risk.
My point is that I would love to see more data on the actual latency distributions we can expect, so that we can make more informed risk assessments. For example, I know that not all `std::atomic` operations are lock-free, but when the critical section is so small, is it really a problem in practice? I want histograms!
[1]: http://www.rossbencina.com/code/real-time-audio-programming-...
When you have 100k people paying $500 to the sky is the limit, failure is not an option. Increasingly audio engineers and subsequently performers are at the mercy of the latest jr developers who don’t have to live with the failures of their short sightedness. Grimes’ Coachella set case in point. Wholly due to pioneer ignoring their users for over a decade. Sometimes we don’t have 3 days to copy files to a usb drive but I digress.
What do you think happens a dense crowd of 500+ people suddenly starts to have excruciating ear pain?
Grimes is a "dj" that does not understand the software. Fixin that problem is one fucking click on the interface.
For example, here is one rig:
https://www.reddit.com/r/ableton/comments/7y2u3o/ableton_mai...
It uses a Radial SW8 to automatically switch between the redundant machines if one flakes out:
We're always dealing with risk and trade-offs. Maybe you avoid a locking `atomic` synchronization point by implementing a more complicated lock-free ringbuffer, but in the process you introduce some other bug that has you dumping uninitialized memory into the DAC.
I think the advice in TFA is totally reasonable and worth following. I'm just saying that there may be cases where it's OK to violate some of these rules. I'd love to see more data to help inform those decisions.
This isn't even in opposition to the article, which says explicitly:
>Some low-level audio libraries such as JACK or CoreAudio use these techniques internally, but you need to be sure you know what you’re doing, that you understand your thread priorities and the exact scheduler behavior on each target operating system (and OS kernel version). Don’t extrapolate or make assumptions
Assuming the context is a desktop OS (which is the context of TFA), I think that the main source of non-determinism is scheduling jitter (the time between the ideal start of your computation, and the time when the OS gives you the CPU to start the computation). Of course if you can't arrange exclusive or max-priority access to a CPU core you're also going to be competing with other processes. Then there is non-deterministic execution time on most modern CPUs due to cache timing effects, superscalar out of order instruction scheduling, inter-core synchronisation, and so on. So yeah, you're going to need some margin unless you're on dedicated hardware with deterministic compute (e.g. a DSP chip).
Most people learning audio programming aren't making a standalone audio app where they do all the processing, or at least not an interesting one. They're usually either making something like a plugin that ends up in somebody else's bus/graph, or something like a game or application that creates a bus/graph and shoves a bunch of different stuff into it.
A much different experience from embedded programming, where 99% occupancy is no problem at all.
An audio glitch is very annoying by comparison, especially if the application is a live musical instrument or something like that. Even the choppy rocket motor sounds of Kerbal Space Program (caused by garbage collector pauses) are infuriating.
It's kind of the difference between soft and hard real time systems. Although most audio applications don't strictly qualify as hard real time (missing a deadline is as bad as a total failure) but failing a deadline is much worse than in graphics.
the cpal library in Rust is excellent for developing cross-platform desktop applications. I'm currently maintaining this library:
https://github.com/chaosprint/asak
It's a cross-platform audio recording/playback CLI tool with TUI. The source code is very simple to read. PRs are welcomed and I really hope Linux users can help to test and review new PRs :)
When developing Glicol(https://glicol.org), I documented my experience of "fighting" with real-time audio in the browser in this paper:
https://webaudioconf.com/_data/papers/pdf/2021/2021_8.pdf
Throughout the process, Paul Adenot's work was immensely helpful. I highly recommend his blog:
https://blog.paul.cx/post/profiling-firefox-real-time-media-...
I am currently writing a wasm audio module system, and hope to publish it here soon.
If your tempo drifts, then you're not going to hear the rhythm correctly. If you have a bit of latency on your instrument, it's like turning on a delay pedal where the only signal coming through is the delay.
One might assume if you just follow audio programming guides then you can do all this, but you still need to have your system setup to handle real time audio, in addition to your program.
It's all noticeable.
As a former developer of real time software, the usage of "real time" to mean "fast" makes me cringe a bit whenever I read it. If there's a TCP/IP stack in the middle of something, it's probably not "real time."
"real time" means there's a deadline. Soft real time means missing the deadline is a problem, possibly a bug, and quite bad. Hard real time means the "dead" part of "deadline" could be literal, either in terms of your program (a missed deadline is an irrecoverable error) or the humans that need the program to make the deadline are no longer alive.
Modern computers are ridiculously fast, relatively speaking you don’t need much resources to calculate a missile trajectory, so “simply” 100% sure doing some calculations at a fixed rate, with even a GC cycle that has a deterministic higher bound (e.g. it will go throw the whole, non-resizable heap, but it will surely always take n seconds), you can pass the requirements. Though a desktop computer pretty much already begets the hard part of hard realtime, due to all the stuff that makes it fast - memory caching, CPU pipelining, branch prediction, normal OSs scheduling, etc.
I suppose it's hard to make guarantees with different environments and hardware, but I realized when we (non-realtime people) ship software we don't really have guarantees for when our functions run.
Keep in mind you don't use the same operating systems (or often even the same hardware) in hard real time applications. You'll use a real time operating system (FreeRTOS, VxWorks, etc) with a very different task scheduler than you're probably used to with processes or threads in Unix-like platforms. That said, while multitasking exists for RTOSes in practice you're not going to be running nearly as many tasks on a device as say a web server.
You can get worst case performance of a section of code by guaranteeing that it has a fixed maximum number of operations (no infinite loops, basically). Of course the halting problem applies, but you're never concerned with solving it in the general case, just for critical sections. It gets tricky with OOO architectures but you can usually figure out a worst case performance.
- You are at the mercy of the browser. If browser engineers mess up the audio thread or garbage collection, even the most resilient web audio app breaks. It happens.
- Security mitigations prevent or restrict use of some useful APIs. For example, SharedArrayBuffer and high resolution clocks.
It's worth noting that these are practically the only case where extreme real-time audio programming measures are necessary.
If you're making, for example, a video game the requirements aren't actually that steep. You can trivially trade latency for consistency. You don't need to do all your audio processing inside a 5ms window. You need to provide an audio buffer every 5 milliseconds. You can easily queue up N buffers to smooth out any variance.
Highly optimized competitive video games average like ~100ms of audio latency [1]. Some slightly better. Some in the 150ms and even 200ms range. Input latency is hyper optimized, but people rarely pay attention to audio latency. My testing indicates that ~50ms is sufficient.
Audio programming is fun. But you can inject latency to smooth out jitter in almost all use cases that don't involve a live musical instrument.
Yes, background sound in games can be handled with very large buffers, but most players expect music-performance-like latency for action-driven sound.
Musicians have keenly trained ears. I would imagine their much more sensitive to audio latency than even a pro gamer, nevermind average Joe off the street.
Where latency really matters is when you have a musical instrument that plays a sound and it's connected to a monitor. If those sounds are separated by more than 8ms or so the difference will be super noticeable to anyone, including Joe off the street.
I'd be interested for someone to run a user study on MIDI keyboard latency. I'd bet $3.50 that anything under 40 milliseconds would be sufficient. Maybe 30 milliseconds. I'd be utterly shocked if it needed to be 8 milliseconds. And I'd be extremely shocked if every popular MIDI keyboard on the market actually hit that level of latency.
I would love to see a UI system that has predictable low-latency real-time perf, so you could confidently achieve something like single frame latency on 144Hz display.
A graphics micro-stutter not so much.
> I'm not aware of any UI toolkits designed with real-time in mind.
What would be the point? The human eye can only notice so much FPS (gamers might disagree with their 244 FPS displays).
The insight is that with two threads contending on one lock, there are efficient ways to build the lock that minimizes cpu on the non-realtime thread.
Text based: SuperCollider, Csound, Chuck
DDMF's VirtualAudioStream does that. It allows you to create virtual audio devices with chains of arbitrary VST plugins. As for the VST plugins, there are thousands of free and paid plugins for everything. I'm using VirtualAudio stream to put a Wave's noise cancelling and a good compressor between my mic and Zoom. It increases latency, of course.
I think so. TBH I'm quite new to the world of DSPs so I don't know the right terminology. The purpose of the DSP (which I should've mentioned in my original post now that I think of it) is to tweak the speakers on my laptop - there are for example ways to "fake" bass (through missing harmonics), or have dynamically changing bass. I'll have a look at VirtualAudioStream, thanks for the recommendation.
What are you even talking about, man.
I think things have just changed technologically a lot. Back in those days, the entire audience didn't all have wireless radios in their pockets, so surely the issue of EM interference is much, much worse now. Also, radio systems these days are almost all digital, probably because of interference partly, but partly also because that's just where all the tech is: we know how to make good digital radios now, they're readily available and cheap, so if you're designing a new product that needs a radio, why do it any other way? It's like programming languages: if you need to write a new program for work, you'll probably write it in something currently in-vogue that lots of other engineers are familiar with, like Python or C++ or whatever, and not something old like Lisp.
Obviously the happy case is when all the audio processing is done in a DSP where scheduling is deterministic, but it's rare to be able to count on that. Part of the problem is that modern computers are so fast that people expect them to handle audio tasks without breathing hard. But that speed is usually measured as throughput rather than worst-case latency.
The advice I'd give to anybody building audio today is to relentlessly measure all potential sources of scheduling jitter end-to-end. Once you know that, it becomes clearer how to address it.
Of course, it's not unusual that the many layers of abstraction in modern systems actively frustrate getting real performance data. But dealing with that is part of the requirements of doing real engineering.
Would buffer level checking entail, eg. checking snd_pcm_avail() when my audio code begins and ends (assuming I'm talking directly to hardware rather than PipeWire)? Dunno if PipeWire has a similar API considering it's a delay-based audio daemon and probably checking snd_pcm_avail when a client requests would slow it down.
All that matters is/are:
1. how soon after the device deadline (typically marked by an interrupt) until a kernel thread wakes up to deal with the device? (you have no control over this)
2. what is the additional delay until a user-space application thread wakes up to deal with the required data flow? (you have very little control over this, assuming you've done the obvious and used the correct thread scheduling class and priority)
3. does user space code read & write data before the next device deadline? (you have a lot of control over this)
As noted above, cyclictest is the canonical tool for testing the kernel side of this sort of thing.
However, since actual context switch times depend a lot on working set size, ultimately you can't measure this stuff accurately unless you instrument the actual application code you are working with. A sample playback engine is going to have very different performance characteristics than an EQ plugin, even if there is theoretically more actual computation going on in the latter.