Android’s 10 Millisecond Problem explained
superpowered.com
superpowered.com
For those interested in why iOS is better than android, a good summary would just be to say Core Audio was designed from the ground up to be a low level low latency audio API. There are fewer layers in it vs android. When designing OS X Apple knew they had a large professional media market so that got priority. I am interested to see the results of these guys efforts and have wondered when an ASIO equivalent would pop up for android.
However, as someone who has dealt with Android, iOS, Windows Phone, and BB audio APIs, Android's audio API really is just amateur hour compared to Apple's, even to the present day. In older versions it was simply unusable for anything more complicated than playing a sound effect or fixed-length mp3 file. Today it functions for more use cases but still abstracts concepts (like codecs and container parsing) in flawed and inconvenient places.
Windows Phone and BB are still even worse though.
I'd personally be much more at ease if I saw ALSA sitting directly under Qt, but I can understand why they'd want to leverage the huge Android hardware ecosystem.
1. https://developer.ubuntu.com/en/start/ubuntu-for-devices/por...
Fewer layers usually also means less flexibility.
Edit: Anyone care to elaborate why the downvotes? From a technical standpoint, I'm absolutely right.
Because if you have access to assembly, you can implement the whatever higher level language or semantics you want.
Whereas if you only have a higher level language, you have to work with/around the abstractions baked into it.
For realtime applications it's better to at least have access to a low-level, less overhead API.
For example, in assembly we can create a higher level language with an absolutely air-tight, precisely tracing garbage collector that is free of issues like false retention.
We cannot do that in C. That's because there are areas of the program state that are "off limits", and the compiler generates "GC ignorant" code.
In C, if we have a pointer p which is the last reference to some object, and is not used any more, and add the line "p = NULL", hoping to drop a reference so the object can be reclaimed, there is no guarantee that the compiler actually generates the code which does the assignment. Since the variable has no next use, and the compiler doesn't know anything about garbage collection, the assignment looks like wasteful, dead code that should be optimized away. Even if the scope finishes executing, the compiler can leave behind a memory location which still references the object.
Here is something else, not related to GC. In assembly language, we can make ourselves a calling convention for variadic functions which know how many arguments they have. As we build up the higher level language, it will have nicely featured variadic functions.
In C, we are stuck with <stdarg.h> which doesn't have a mechanism for the callee to know where the arguments end. The language has no flexibility to add this --- without resorting to approaches which will basically involve assembly language.
You can do it portably if you just don't store objects on the C stack, and you can do it non-portably if you do so.
If so, you will be the first; I look forward to the "Show HN:" when it's done.
> You can do it portably if you just don't store objects on the C stack,
Sure, for example, you can do it portably in C if you write a complete emulator for an 80386, and then use that to run an assembly language program. That program isn't a utility for your C code; it's not extending the host C with a garbage collector or whatever else.
If you do not store object references on the stack, your use of C is severely crippled to the point that it's not really C any more. For one thing, C function arguments are on the stack (or, more abstractly, "automatic storage"), so say goodbye to conventional use of C argument passing: the backbone of most normal C programming.
That garbage collector isn't for C. C has automatic storage, and a proper garbage collector has to traverse it.
> and you can do it non-portably if you do so.
Sure, if non-portably means going as far as forking a specific C compiler with your custom hacks, and requiring that C compiler, or else living with the imprecision and taking a whack-a-mole approach to plugging the issues as they arise.
http://en.wikipedia.org/wiki/CHICKEN_%28Scheme_implementatio...
[edit] And here is a toy scheme interpreter that uses precise garbage collection to show off the Ravenbrook MPS
http://www.ravenbrook.com/project/mps/master/manual/html/gui...
The fact that Apple totally controls both sides of the symbiotic audio hardware & SDK running on a tiny set of products, while Android must have an SDK which accommodates an unknown range of hardware for hundreds of products, means iOS will have an inherent advantage.
The fact that Apple totally controls both sides of the symbiotic audio hardware & SDK running on a tiny set of products
You can use CoreAudio with third party audio hardwareDeleted comment
You're being that guy, so obsessed with a thing that you see it everywhere and think every conversation is about it.
It proves the system supports a ton of flexibility.
In that sense, flexibility and ease of use aren't synonyms; they're opposites. An open flame is flexible but not easy to use, so we have the toaster, which is easy to use but not flexible at all.
None of that stops you from having a high-level API on top, and in fact iOS has several at different levels of abstraction: AVAudioEngine gives you a lighter weight but object-oriented engine that's less complex to setup. AVAudioPlayer handles almost everything for you.
Abstractions can hide features of the hardware but they cannot create new hardware. Whatever the abstraction is doing the client software could do instead with the lower level API.
So I would say abstraction layers are often added to give ease of use, sometimes at the cost of performance and flexibility.
An audio abstraction layer can allow different kinds of audio hardware without changing the apps, which is a kind of flexibility.
My first modern touch device was an iPod touch 4. I downloaded Garage Band and, as a long time milt instrumentalist and composer, loved it. I was amazed by how well the touch instruments worked and how easily I could record riffs and flesh out small snippets of songs. It ran almost flawlessly on the 4th gen Touch.
Next, I decided to buy an Android phone - a Motorola Droid 2. I was surprised to find that despite the power advantage over the iPod, none of he music apps I tried were usable. The drum programs, for instance, we're so lacy and unpredictable as to be worse than useless. Hit a drum, and you may hear it seemingly instantly, maybe 1/8 a second later, maybe half a second, possibly never. Meanwhile the tiny iPod could play Garage Band instruments so well one could use it for live performance.
I upgraded my phone twice, first to a Droid X, then a Galaxy S3... Each time was disappointed that the improved specs gave no improvement in the terrible audio performance.
Currently I have an iPhone 6 so I can use Garage Band. Kudos to apple for doing this right - it's the best app I've ever used.
http://source.android.com/devices/audio/latency.html
including a bunch of measurements of Nexus devices here (even going as far back as the Nexus One on Gingerbread):
http://source.android.com/devices/audio/latency_measurements...
The iOS devices come out at 6-18ms, Androids at 17-860ms. The faster Androids have Samsung's Professional Audio SDK.
If whatever audio framework you use doesn't allow to run processing with a input-to-output delay (latency) of two times the period size, it's broken (probably the case for Audio Flinger at Android, don't know much about it).
➜ ~ jackd -d alsa -p 64 -r 96000
jackdmp 1.9.10
Copyright 2001-2005 Paul Davis and others.
Copyright 2004-2014 Grame.
(...)
creating alsa driver hw:0|hw:0|64|2|96000|0|0|nomon|swmeter|-|32bit
configuring for 96000Hz, period = 64 frames (0.7 ms), buffer = 2 periods
ALSA: final selected sample format for capture: 32bit integer little-endian
ALSA: use 2 periods for capture
ALSA: final selected sample format for playback: 32bit integer little-endian
ALSA: use 2 periods for playback
(this is on my laptop, just for illustration purposes)If you write "process however much audio data is ready", then you already imply that your CPU will not be up to speed to process 48000 interrupts/second reliably and you need some buffering.
And if you have to assume that sometimes you'll miss 100 samples (which, then, you'll process en-block), this means that to work reliably, you'll have to start at least 100 samples early so that you don't miss the deadline of the DAC, because the DAC will, with intractably output one sample every 48000th of a second. This already implies some kind of periodic processing of blocks, doesn't it?
(and yes, such a scheme will theoretically allow you to half the latency from something to 2period-size to 1period-size + the time for processing)
Third, a lot of the algorithms for processing audio can be implemented much more efficiently if you have a known block size and don't have to calculate your filters or convolutions with constantly changing number of samples for every step.
Also efficiency of processing will decrease (reloading the cache after each interrupt when switching from processing plugin to processing plugin), so the time spent on calculating per frame will go up if your period size gets smaller. At one point you'll need exactly one "period size" to calculate one period size worth of samples: That's the maximum your machine can handle, and at that point you'll have a latency of your "period size"*2, which is exactly the same as running with a fixed period-size ;-). And as you can choose the period-size rather freely (maybe completely arbitrary, maybe 2^n, depends on the chipset/hardware) there's no disadvantage left.
>As someone who has actually done a good amount of soft-real-time audio programming, I can tell that you probably haven't. Everything you are saying about CPU speeds is made-up nonsense.
You could have wrote:
>I've actually done a good amount of soft-real-time audio programming and everything you're saying about CPU speeds doesn't make sense.
I think it's good to call out what you see as misleading information but that may have been going a bit too far.
But I'll drop a few hints. First of all, nobody is talking about running interrupts at 48kHz. That is complete nonsense.
The central problem to solve is that you have two loops running and they need to be coordinated: the hardware is running in a loop generating samples, and the software is running in a (much more complicated) loop consuming samples. The question is how to coordinate the passing of data between these with minimal latency and maximum flexibility.
If you force things to fill fixed-size buffers before letting the software see them (say, 480 samples or whatever), then it is easy to see problems with latency and variance: simply look at a software loop with some ideal fixed frame time T and look at what happens when T is not 100Hz. (Let's say it is a hard 60Hz, such as on a current game console). See what happens in terms of latency and variance when the hardware is passing you packets every 10ms and you are asking for them every 16.7ms.
The key is to remove one of these fixed frequencies so that you don't have this problem. Since the one coming from the hardware is completely fictitious, that is the one to remove. Instead of pushing data to the software every 10ms, you let the software pull data at whatever rate it is ready to handle that data, thus giving you a system with only one coarse-grained component, which minimizes latency.
You are not running interrupts at 48kHz or ten billion terahertz, you are running them exactly when the application needs them, which in this case is 16.7ms (but might be 8.3ms or 10ms or a variable frame rate).
You don't have to recompute any of the filters in your front-end software based on changing amounts of data coming in from the driver. The very suggestion is nonsense; if you are doing that, it is a clear sign that your audio processing is terrible because there is a dependency between chunk size and output data. It should be obvious that your output should be a function of the input waveform only. To achieve this, you just save up old samples after you have played them, and run your filter over those plus the new samples. None of this has anything to do with what comes in from the driver when and how big.
Edit: I should point out, by the way, that this extends to purely software-interface issues. Any audio issue where the paradigm is "give the API a callback and it will get called once in a while with samples" is terrible for multiple reasons, at least one of which is explained above. I talked to the SDL guys about this and to their credit they saw the problem immediately and SDL2 now has an application-pull way to get samples (I don't know how well it is supported on various platforms, or whether it is just a wrapper over the thread thing though, which would be Not Very Good.)
This is, actually, how most professional audio APIs are designed and they generally work quite well. ASIO, VST, JACK, PortAudio, CoreAudio, etc.
However, even when they are synced, you can still easily see the problem. The software is never going to be able to do its job in zero time, so we always take a delay of at least one buffer-size in the software. If the software is good and amazing (and does not use a garbage collector, for example) we will take only one delay between input and output. So our latency is directly proportional to the buffer size: smaller buffer, less latency. (That delay is actually at least 3x the duration represented by the buffer size, because you have to fill the input buffer, take your 1-buffer's-worth-of-time delay in the software, then fill the output buffer).
So in this specific case you might tend toward an architecture where samples get pushed to the software and the software just acts as an event handler for the samples. That's fine, except if the software also needs to do graphics or complex simulation, that event-handler model falls apart really quickly and it is just better to do it the other way. (If you are not doing complex simulation, maybe your audio happens in one thread and the main program that is doing rendering, etc just pokes occasional control values into that thread as the user presses keys. If you are doing complex simulation like a game, VR, etc, then whatever is producing your audio has to have a much more thorough conversation with the state held by the main thread.)
If you want to tend toward a buffered-chunk-of-samples-architecture, for some particular problem set that may make sense, but it also becomes obvious that you want that size to be very small. Not, for example, 480 samples. (A 10-millisecond buffer in the case discussed above implies at least a 30-millisecond latency).
There's a relatively recent trend to try and record digital all the way, and this is also complicated. Record at 24 fps? At 30? At 60? 60 fps 4k is a lot of data. And sound is actually the major pain point -- video frames you can generally just drop/double, speed up/down a little to even things out. But 24 fps to 60 fps creates big enough gaps that audio pitch can become an issue.
But jblow is right in that when you have to feed in samples from a non-synchronized source into your processing/game/video-application/... then trying to work with the fixed audio block size will be terrible/require additional synchronization somewhere else, such as a adaptive resampler on the input/output of your "main loop".
Why do you say this? The USB audio card (or similar) is generating blocks of audio at a fixed rate, no?
Maybe for video playback or games you need to synchronize audio and video, but there is no need to do that for music production apps.
If you are writing some sort of synth, as soon as you receive a midi note or a tap, trigger the synth and the note will play in the next audio block. No need to wait for the GUI to update.
If you are doing some sort of effect, grab the input data, process and have it ready for the next block out. I don't understand why you need a second loop.
I think you will find, though, that most hardware isn't this way, and to the extent this problem exists, it is usually an API or driver model problem.
If you're talking about a sound card for a PC, probably it is filling a ring buffer and it's the operating system (or application)'s job to DMA the samples before the ring buffer fills up, but how many samples is dependent upon when you do the transfer. But the hardware side of things is not something I know much about.
> If you are writing some sort of synth, as soon as you receive a midi note or a tap, trigger the synth and the note will play in the next audio block
Yeah, and waiting for "the next audio block" to start is additional latency that you shouldn't have to suffer.
> If you are doing some sort of effect, grab the input data, process and have it ready for the next block out. I don't understand why you need a second loop.
The block of audio data you are postulating is the result of one of the loops: the loop in the audio driver that fills the block and then issues the block to user level when the block is full. My whole point is you almost never want to do it that way.
2) In order to hand you off some samples you have to at least make one copy. It's convenient to be able to copy an entire _something_ without worrying about the sound card trying to DMA into it (and any hardware-specific details to make that possible).
The difference here is that with something like jackd (similar conceptually to CoreAudio or ASIO) there is just the hardware buffer in the kernel and the user buffer in jack which can be basically "shared" by all jack-enabled apps without additional copying on the user-space side. On the other hand you can't do sample-rate conversion and per-app volume control with something like that.
But if you're doing audio software you're not worrying if flash is too loud and Skype is too soft. It's a whole-system thing enabled by the software and all the buffers and latency are being managed there.
I don't understand why copies are even relevant: you can make several extra copies and nobody will ever notice. Audio data is trivial in modern systems. Let's say there are two channels coming in; 48000 * 2 * 2 bytes per second is an absolutely trivial amount of data to copy and has been for many years. Building some convoluted (and unreliable) system just to prevent one copy per application, when each application is going to be doing a lot of nontrivial processing on that data, strikes me as foolish. But don't listen to me, look at the fact that Linux audio is still famously unreliable. If the way it's done were a good idea, it would actually work and everyone would be happy with it.
USB has several transfer types: interrupt, bunk and isochronous.
Interrupt is initiated by the device so it wouldn't help you read variable amounts.
Bulk is good for mass data transfers, but has guaranteed access to the bus, so you'd possible get audio dropouts when accessing another USB device.
Isochronous can reserve bus bandwidth and have latency guarantees, but must occur on a fixed schedule. Since they are on a fixed schedule, they always have the same amount of data, hence a fixed block size.
Since data is arriving at the OS in fixed blocks, the lowest latency way to handle the data is deal with the blocks when the arrive. If you wanted to read variable amounts of data, you'd need to add a buffer on top that could hold a variable amount of data which would add latency.
Copying variable amounts of data isn't slow, but dealing with dynamic memory allocation is. If you needed to allocate a different sized block every few milliseconds you'd be spending the majority of your time allocating memory rather than processing audio.
This system can work very well: OS X, iOS, Windows and Linux can all get very low latencies. The issues with Android has nothing to do with block sizes, but something else in it's architecture.
Android's media architecture with it's "push" design is not more compatible than anything else with a "pull". The many layers doesn't add any flexibility, they are "only" the result of many years hacking.
Raph Levien talks about audio latency and how they started working on minimising it in Lollipop. He mentions that it’s still an ongoing process.
The relevant part starts at 35 minutes into the episode.
[1]: http://androidbackstage.blogspot.com/2015/01/episode-20-font...
Related presentation at Google I/O 2013. https://developers.google.com/events/io/2013/sessions/325993...
This was probably the best presentation I saw during Google I/O.
There are still additional problems that add latency that span the entire Android stack from the actual hardware and drivers, to the kernel and scheduler, to the Android implementation of their audio APIs.
Anybody who has tried to do any serious audio on Android knows the infamous Bug 3434:
https://code.google.com/p/android/issues/detail?id=3434
Google I/O 2013 did a pretty good talk on the problem and shows how there are problems across the entire stack. Glenn Kasten pretty much carries the brunt of all the audio problems with Android. I find it telling that he had to handcraft his own latency measurement device using an oscilloscope and LED because there were no actual tools built into the OS to help them analyze performance.
https://www.youtube.com/watch?v=d3kfEeMZ65c
Audio has been terrible since Android's inception. It has improved a little over time, but unfortunately, 7 years later is is still pretty much unacceptable for any serious work.
Android is famous for $600 2GHz phones that drop frames when doing simple animations :(.
Sound travels at about 340 m/s (in a typical room). That means it travels about 3.4 metres in 10 milliseconds. Therefore another way to get a 10 millisecond problem is to stand 3.4 metres from the orchestra.
Most people sit farther than 3.4 metres from the orchestra, yet they don't complain about a lag between when the violin bow moves and the sound is heard. Why not?
(The speed of light is so fast that we can assume it's effectively infinite for the purposes of this argument.)
> Most Android apps have more than 100 ms of audio output latency, and more than 200 ms of round-trip (audio input to audio output) latency.
It's also much more than just 10ms latency. I play with digital instruments all the time, the latency can be as high as 15ms before I can tell. I don't know if an audience can perceive a 15ms latency, especially because you tend to "play early" to have notes land on time. But it's very upsetting for performing.
I haven't tried playing with music apps on Android recently, but when I did, the latency was not just long, but inconsistent, and would result in stuttering in the audio.
The best movie about this, ever: https://youtu.be/VnuImW1dWAk
The difference, of course, is that organists are usually not syncing up to other instruments. If there are other instruments involved, they tend to sync up with the organ.
Musicians performing together, however, is a much harder problem than just listening. Ask anyone who has ever performed in a DCI-style drum corps, they will tell you compensating for hearing someone on the other side of field 200ms or so late is incredibly difficult.
Also, most people don't notice the lag between the instruments of an orchestra because the instruments are close to each other. Their distance to the listener is not relevant.
http://blogs.scientificamerican.com/observations/2011/09/15/...
tldr: Brain is inherently parallel, nothing happens in sync. In order to make sense of the outside world higher level functions are presented with artificially coordinated stimuli.
The problem you describe is a very real problem, however, for the musicians themselves. If you're sitting in a big orchestra and you try to index your playing off someone sitting on the other end of the orchestra, you will not be in time. That's why there's a conductor, so that the orchestra can be synchronized at the speed of light rather than of sound.
For a real-world example, listen to 2 TVs several meters apart and tuned to the same channel. At least with OTA or cable you can expect them to be playing ~simultaneously but the skew between the received signals is easily perceptible.
They do depend upon vision or touch if it is an interactive app.
>For a real-world example, listen to 2 TVs several meters apart and tuned to the same channel. At least with OTA or cable you can expect them to be playing ~simultaneously but the skew between the received signals is easily perceptible.
I believe what you're experiencing is the difference in decoding latency between different models of TV set. Several meters (3m) represents only about ~9ns (practically, low tens of ns if the cables are longer than necessary) maximum delay. It would not be directly perceptible by a person. Signal delay in a cable is ~1ns/ft.
Discrepancies between tracks are totally unrelated to the visuals. It's much easier to tell if two sounds are synced than a sound and a visual.
> Signal delay in a cable
I don't know what comment you read but it's not the one you replied to. Same model, synchronized visuals, easy to hear audio desync when you're closer to one.
Recently I got one of those cheap USB interfaces to connect my guitar. I spent some good 4 hours changing the kernel to "low latency" one provided by ubuntu, then trying to setup aRTs to run, then trying to make pulseaudio work with it, then figuring out how to keep both aRTs and pulseaudio dependent applications happy. In the end I got most of it working and I could run some kind of guitar effects application, but the next minute I realized that the volume control media keys stopped working. It was enough for me to throw the usb interface in the drawer and give up on linux for audio applications on workstations: buying a 50€ effect pedal was cheaper than all the time devoted to it.
http://en.wikipedia.org/wiki/Latency_(audio)
Linux has a Real Timer kernel you can run, I used to run it when doing audio stuff, it works great.
It grands an application exclusive, direct control of the audio hardware drivers, bypassing the higher-level audio layer of the OS.
That + real-time threads (I do not know how old that feature is) allows for pro-grade latencies.
One can get very short latency out of Alsa, up to the point where the hardware becomes your bottleneck. But that's extremely processor intensive, and won't work well if you try to share the dsp with several processes (if you want to get that extreme, I'd recommend you get extra hardware for exclusive use of the application you want low latency from - but here I'm talking about 1ms latency).
Anyway when using a cheap USB interface, I'd focus on improving the hardware first. Low latency and high throughput USB isn't cheap (nor is it available at every computer).
About the interface. I doubt that was the problem. When I got jack to work, I was getting 2-3ms latency between input and processed output. I did get the guitar effect application to work, I just thought it was too inconvenient to be forced to be aware of "what-application-uses-what-sound-system-and-when-I-need-to-flip-the-switches".
In any case, my goal was to have a alternative that could be (a) cheap and (b) convenient if I wanted to play with my guitar and have some DAW tools, not to see how low I could bring down latency in a linux system. The lesson learned is that it can be cheap, but not convenient.
Ps: did you go to Unicamp?
Well, your question that implied you wanted it. Although, yes, the pedal is probably a better choice after all.
Anyway, Pulse is only needed for advanced tasks of streaming sound through a network, using application based mixing settings, etc. If you are only doing common tasks, they'll almost certainly keep working without it. It's one of those cases of an apt-get and you are done.
Yes, I went to Unicamp, 99's class. Is your nick based on your name?
No. If you're doing serious audio in Linux, you use jack and get fantastic latency properties.
This is a Solved Problem.
Except I am not. And I guess it would be the same case for the vast majority of people who have Android devices.
It is only a "Solved Problem" if you are talking exclusively about sound systems that are designed to deal with latency. The problem I was stating is that you can not transparently run "serious audio" and common desktop applications that rely on alsa/pulseaudio.
I would consider it a "Solved Problem" when I can get to install Audacity, Skype, my web browser, the mentioned usb interface connected and I can run some effect software... and run them all concurrently without caring how to setup the sound system. People can do that in MacOS/iOS/Windows, and they can't do that in Linux (GNU or Android).
Back then the sound subsystem didn't do any mixing or similar, so if some program grabbed /dev/snd, everyone else had to wait.
As for low latency sound work on Linux today, Jack is what you want rather than pulseaudio. Frankly Pulseaudio is a massive detour when it comes to Linux audio.
On mainstream distros with pulseaudio like Fedora or Ubuntu, audio just works, when you don't have low-latency requirements.
When you do have low-latency requirements, things are a bit tricky. You do pretty much need to get a low-latency or realtime kernel, and you definitely want to use Jack or ALSA. Maybe this setup is a little more frustrating than Windows, where you may just need to install one driver like ASIO4ALL.
But it's flexible, and you can actually get Pulseaudio and Jack working pretty well together after installing the pulseaudio-module-jack. You can turn on jack when you need it, turn it off when you don't, and all of the pulseaudio stuff will get routed through it so that you don't lose sound from your other applications. If you want, you can route audio from pulseaudio into whatever other audio applications you're using - sometimes I like taking Youtube videos and routing the audio through weird effects in Pure Data.
I toggle jack on and off with this little script, works pretty well for me right now: https://gist.github.com/YottaSecond/f0a1b515f95b2e791755
When Ubuntu adopted PA, the readme file still described it as "the sound server that breaks your audio"
It was the most mature thing that had the features that Ubuntu wanted, so they adopted it despite the fact that it was clearly not yet ready for prime-time.
That all being said, PA is not the choice if you want to do DAW style stuff; it tends to prefer lower cpu utilization to lower-latency.
- The applications that depend on pulseaudio were really not happy when jack was the sink.
- Skype wouldn't work.
- Because of the real-time requirements of the sound applications, everything else felt absolutely sluggish.
- The volume control stopped working.
All in all, I'd have to setup a separate system just to run jack-dependent applications.
Idea: A latency ranking for devices would put pressure on manufactures.
Will you work together with the Linux community so we all benefit from it?
As Westley from The Princess Bride might reply, "As you wish"
It's remarkable how apple cares about certain quality aspects of their devices. They would be interesting options if they wouldn't jail and lock you down :/ However, I wonder why newer and supposedly faster devices like the iPhone 6 have higher latency. I'd think to surpass the latter generation would be the goal for each successor.
Also, it seems Samsung is taking the challenge seriously.
On the Galaxy Nexus, for example, the best latency I can get appears to be 176 ms. This is pretty high for certain types of applications, particularly ones that generate tones based on user input. With PulseAudio, where we dynamically adjust buffering based on what clients request, I was able to drive down the total buffering to approximately 20 ms (too much lower, and we started getting dropouts). There is likely room for improvement here, and it is something on my todo list, but even out-of-the-box, we’re doing quite well.
Let's get systemd on android and then see !
[1] http://arunraghavan.net/2012/01/pulseaudio-vs-audioflinger-f...
http://createdigitalmusic.com/2013/05/why-mobile-low-latency...
This is obviously an issue when doing real-time audio processing or generation: if you're using virtual instruments and the feedback to your cans is 1/10th of a second behind the interaction it's unusable.
But it's also a problem for more mundane applications: a 200ms roundtrip (100ms to move samples from the mic to the application on the talker's side, and 100ms to move them from the application to the speaker on the listener's side) is is more delay than the transmission time from somebody literally on the other side of the earth. Same with fast-paced game, audio feedback >100ms after an action is highly bothersome or even game-breaking (the sound is output several frames behind the gamestate)
The embedded graph with the U shaped figure pretty much sums up the author's search for the components in the software that are the culprits.
Not just that, if you try to play guitar with a 10ms latency from input to output, you'll find that it's impossible to keep rhythm at all. It's the same effect as a speech jammer: http://www.stutterbox.co.uk/
And I don't think many people outside the the tech audio community realize how many people have ditched their guitar amplifiers and are playing guitar purely through the iOS devices. It's a great use case for tablets, but is completely infeasible to do on Android at the moment.
I'm super interested in this question because I honestly don't know. My assumption is that among serious guitar players (defined as individuals who play in groups or in front of people at least monthly) the number is almost nil.
I'm sure the number of casual guitarists (those who rarely play, but technically own one), this number is quite high.
I'm unwilling to even concede the tubes in my amp let alone the amp itself...
Few serious guitarists would resign themselves to exclusively playing through an iOS device. Nearly all serious guitarists will do it on a regular basis though, which is the more relevant point.
Or imagine trying to listen to your own voice, but having the sound delayed by that 35ms, so it feels like someone is talking over you every time you open your mouth.
It's particularly problematic when it's being copied in mac circles because OSX has a biased font-smoothing algorithm that adds font weight (it's a design flaw). In other words, to get something to look thin on a mac, especially a non-retina mac, it needs to be very thin, sometimes less than a pixel thin. How other systems display borderline visible strokes varies depending on the system and the details of the font.
There are some pretty great instrument apps available for the iphone, like the ikaossilator by korg, a tonne of great drum machines, loop apps, synths etc. And mixed with the Audiobus app that lets you route the sound of one app into the input of another (as long as both apps support Audiobus which most seem to do), the possibilities for creating music entirely on your phone are limitless
If anyone is interested, I pulled the OpenSL ES parts out and posted them to github.
The sound chips operate on integral periods or a fixed number of samples, and when you get a period worth of audio data from your ADC, the DAC will already have started putting out the first samples of the next period. Hence, you prepare your audio samples for the second-to-next period.
Assume you have audio processing code roughly looking like this:
while (1) {
poll(); /* some API function waiting for the "next period" */
read(soundcard, block_of_samples); /* or let DMA do it */
process_samples(block_of_samples);
write(soundcard, block_of_samples); /* or let DMA do it */
}
Let's try some ASCII art: v- audio samples going into your soundcard
_ _ _ _ _ _ _ _ _ _ _ _
/1\ /2\ /3\ /4\ /5\ /6\ /7\ /8\ /9\ /0\ /1\ /2\ ...
\_/ \_/ \_/v \_/ \_/ \_/ \_/ \_/ \_/ \_/ \_/ \
|
\________________/\________________/\________________/\________________/
Period 1 Period 2 Period 3 Period 4
|
|
[*] here, Period 1 has been DMA'ed from the soundcard
to an mmapped buffer of your audio application
<~~~~~~~~~~~> here processing of your audio takes place
[*] this is the latest point at which processing
| must complete so that there will be data for
| the soundcard to output. DMA will start
| from the buffer to the soundcard DAC.
|
_ _ _ _ _ _ v _ _ _ _ _ _
/ \ / \ / \ / \ / \ / \ /1\ /2\ /3\ /4\ /5\ /6\ ...
\_/ \_/ \_/ \_/ \_/ \_/ \_/ \_/ \_/ \_/ \_/ \
\________________/\________________/\________________/\________________/
^- processed audio samples will come out
of your laptop's speakers...
Obviously, your machine has to be fast enough to do the whole real-time audio computation in a little less than the time between two interrupts, that's the period size. And it must reliably be able to do this, because if it misses the time to have a block of samples ready for the DAC. If it misses that goal, a "underrun" will take place, and the audio application will have to resynchronize, possibly causing some clicking, intermittent audio, ...The "problem" part is that Android devices generally don't come close, the lowest-latency Android device on the market today (barring Samsung's custom "Professional Audio SDK") has a 35ms roundtrip latency (ADC -> DAC), and many devices are way beyond 100ms (http://superpowered.com/latency/)
I notice any latency above 7ms. I can compensate pretty well between about 7-14ms. Anything above that will affect the performance. Once you get to, say, 25ms, it's like I lose several years of training.
It's hard to relate to because most people don't have experience doing tasks demanding that level of precision. Perhaps it would be useful to think about what would happen if you introduced a delay between striking a key on your keyboard and the sensation of feeling it travel. You "can type" without any sensation at all (see iPads) but people who are serious about typing go to extraordinary lengths to make the keyboard feel right.
For an audience, 10ms is going to feel like sloppy timing. That will be more of an issue with tightly timed music (techno/EDM) and less with a slow jazz ballad.
So 5ms vs 10ms latency is like the difference between having your amp 5 feet or 10 feet away.
It was so bad to the point where I went out and purchased an iPad just so I could have that freedom of recording ideas when I am not at home. The iPad and iPhone offered what sounded like basically no latency at all, I never understood why Android devices struggled (but I speculated and assumed it was how the audio was being processed). I considered moving back to an iPhone, but I love the freedom that Android affords me and the competition, so I stuck it out and kept using the iPad.
I recently purchased a Samsung Galaxy S6 Edge and while I still notice some slight latency, it is usable again. I can finally jot down ideas when I am away from home using my phone again. Research seems to yield some improvements that Samsung themselves have made to their hardware and software, not to mention the professional audio driver that allows the use of a third party audio interface (great feature by the way). Google needs to make this a priority, because believe it or not a lot of people use their tablets and phones to produce music. We need to fix this.
Music production might not seem like a big deal to Google, but Apple definitely gets it and have from the beginning. The one aspect I miss about owning an Apple device, not enough to make me switch back but definitely a good feature of iOS devices.
Also, Apple is much more motivated to get the audio path right, having started with the iPods and selling a lot of music.
The reason ALSA and AudioFlinger add latency is to hide hardware-dependent differences as well as kernel-caused scheduling issues & policy decisions.
To achieve low-latency you need real-time scheduling, something Linux has with SCHED_FIFO but it's a bit kludgy, and getting the policy right on that is tricky (obviously you don't want a random app to be able to set a thread to SCHED_FIFO and preempt the entire system). So you have to restrict the CPU budge of a SCHED_FIFO thread, and you have to only allow apps to have a single SCHED_FIFO thread. But how much CPU time you give it needs to depend on the CPU's performance in combination with the audio buffer size that the underlying audio chip needs (and those chips also have different sample rates, is it 44.1khz or 48khz or etc...).
tl;dr: this is insanely hardware-dependent.
https://developer.apple.com/library/mac/documentation/MusicA...
Besides the fact that you are writing in C/Objective C, most heavy lifting is being provided by Audio Units which are specifically designed with common datatypes to be chained, composed and executed in low-latency situations.
Furthermore, most common things you would need in your app like mixing, conversion, timing, etc are provided as highly optimized services by the system.
http://www.quora.com/Why-do-iPhones-have-better-professional...
AudioFlinger + Alsa take a lot of time, as seen per the graph
But Android seems to take the option of least effort and "works most of the time" (which they have their reasons to)
Given the wide scope of hardware targeted by android, it's not that surprising that it performs less well than a system targeting a very limited set of devices.
Having said that, it performs 'well enough' for the vast majority of use cases.
This suggests that you can work focused on one driver and therefore save developer-time. However, the drivers on Android are made by a bigger workforce, which must be taken into account.
> they can skip one of the abstraction layers
The post explicitly states that the HAL ought to add no latency at all.
Consider e.g. AmigaOS.
AmigaOS let you obtain a pointer directly to the screen bitmap to update your window contents with no buffering or clipping. It could do that because originally all the hardware was the same, or close enough.
Then graphics cards came along, and you didn't necessarily have a way of writing directly to the bitmap. Suddenly you had to use WritePixel() and ReadPixel() and similar, which would obtain the screen pointer for the window, and obtain the display the screen is on, and find the driver corresponding to the screen, and call the appropriate driver function via a jump table.
Similarly, the AmigaOS had functions to e.g. install copper lists (the copper was a very primitive co-processor that could be used to do things like change the palette at specific scan lines), which also wouldn't work at all on graphics cards.
This is why knowing the hardware is part of a limited set matters: You can define your API to match the hardware very precisely, or even expose hardware features directly.
I don't know enough about android to say that this is the key aspect tho. Perhaps the low level API of android is the same? I didn't find an easy reference in a quick search.
and it seems to be that anything below 140 ms latency seems impossible on Windows Phone.
I have only 5 ms of ASIO buffering using my PC but I don't actually know how much actual key->sound latency; I do know that using headphones it's only slightly less immediate than a nice grand.
I think the low MIDI bit rate (31kbaud) also adds a little latency on chords.
It would be nice if keyboards in the future immediately sent a lower latency+precision keypress-initiated notice (before full key travel) so disk-based samplers can make ready for that note (and make any initial attack sound that's appropriate).
That's like saying, “Consumers have a strong desire to buy gourmet steaks from McDonald's, as shown by revenue data from Ruth's Chris.”
No, McDonald's serves billions of meals by understanding its own market, not by catering to diners at Ruth's Chris. And naturally, comparing top sellers at each will give very different lists.
If we have to use a food analogy, I would propose Starbucks and a competing coffee shop. Almost all the major players in app development make corresponding Android and iOS versions of their apps. Imagine if Twitter or Snapchat only had an iOS or only had an Android app. Similarly, customers expect to have certain things available at all coffee shops, regardless the brand behind them: cappuccinos, flavored syrups, alternative milk choices, etc. So a better analogy would be one coffee shop chain not supplying their stores with espresso machines, or flavored syrups, or alternate milk choices... or maybe just not letting them have refrigerators at all to keep milk cold.
The prevalence of orders for highly-sweetened, cold, milk-based drinks at Starbucks almost certainly means that customers would order the same thing at similarly positioned coffee shops if it was available, and indeed most coffee shops offer such drinks now.
There is no reason to think that somehow Android users are in such a different market that they would have no interest in these apps on Android even though in almost every other case platform parity is expected from large players in the app market.
I'd say it's a bad analogy because most people who don't live in the US have no idea what Ruth's Chris is.
So it was good, because that was my point.
> Android and iOS are in direct competition.
Citation needed. ;-) But seriously, I don't think they are. Too many observable differences in goals and strategies across each platform for your assertion to be true.
> There is no reason to think that somehow Android users are in such a different market that they would have no interest in these apps...
On the contrary, there is every reason to think that. Outside a very narrow and limited class of tech geeks who argue on technical merits, the broader market stats point to these being very different segments with different consumer profiles valuing different things.
This is not true for low-latency audio; it's just something that Android has not prioritized.
If x is the unit of measure for the end-result and we have n components that add together, then we need at least x/n as unit of measure for the performance of each component. x/(2*n) is more reasonable to not deviate from the target performance more than one x after rounding in the worst-case.
Once they're all below 1ms it's worth increasing the resolution.
https://code.google.com/p/android/issues/detail?id=3434
And very good video from Google about it:
I'm an Android user and have given thought to developing on Android but the only things I'm interesting in doing on a mobile platform are synthesis and sequencing. Based on everything I've read.. it seems like it would be a waste of time. Forget commercial viability, it doesn't even sound like it would be worth it to make something for personal use. I get 7ms roundtrip on my Linux music workstation, with my audio interface, JACK, ALSA, Bitwig.. total cost maybe $1600(when buying the machine I was trying to figure out how much I could leave out/how much of a cheapskate I could be).. not a huge sum but it shows how important low latency is. A lot of the people buying music apps(synthesizers, toys) own hardware synthesizers, have computer recording/sequencing setups, or play acoustic instruments.. some of them spend multiples of 10K on their setups over time.. when they pick up a phone or a tablet, they're not comparing it to a flash site or the performance from an integrated soundcard on a $300 laptop.. they're comparing it directly to the immediate physical response of their instruments or the low-latency response of the computer audio setup they've invested in.
High latency feels incredibly sluggish and harms your sense of rhythm; unpredictable latency is just murder. Given that most makers of music software are themselves music makers.. when they pick up a device and see its audio performance is so poor the thought process is something like "If I can't even make something I would use myself on this, what's the point in trying to release something to the public?"
There are some good music apps for Android, some that I would even say are very good efforts, but they're very few, and still suffer from latency pretty badly.
Apple has a long history of market success with creative professionals, so to make sure that they retain this success, they focus heavily on the product issues that would be relevant to creative professionals. These creative pros are who the advanced amateurs look up to, so when the creative pros, let's say Ryan Lewis, are all using Macs and iOS, the amateurs do too... Apple's success in these markets is a long history of positive feedback between engineering, marketing, and branding.
http://webcache.googleusercontent.com/search?q=cache:zvFOaWH...
Even better is that you can run the OSX/Windows version for £0.00 and copy your projects over to work on a "real" machine after tinkering on your phone throughout the day.
The problem is not the fact that they use ALSA and AudioFlinger. ALSA and AudioFlinger just use at this moment too much time. This could be improved by decreasing the period size.
What happens when they fix it?
At this point, iOS devs have so many numbers of years ahead of Android devs in the music department, and the simplicity of porting existing Mac OS compatible audio stuff makes it incomparable.
Metaphor:
Two people are going for a race, over the same distance. One starts 5 years ahead of the other.
At that point, is anyone even watching the race anymore?
Sorry, despite it's absolutely God-awful flaws such as no user-accessible filesystem, no ability to do a basic task such as download an MP3, without GarageBand (and it's awesome ability to open files I sketch out on the go right in Logic on my Mac...), without iElectribe, and the list goes on, I absolutely just can't even use the platform, and I have no reason to go back 5 years technologically.
Sorry, Android, you already lost, here is a user you can never have.
Doesn't matter how fast developers create applications that are amazing, iOS has already had all these years to create an already fantastic subset of these applications.
Doesn't matter how 'good' it looks - for the use case of music, android has simply lost.
App developers LOVE to port, its another sale with small effort.
How does Superpowered get around ALSA and the Audio Flinger?
The smaller the audio buffers are, the more prone they'll be to starvation. Only if the application can be guaranteed to receive interrupt service and/or thread timeslices at a 100 Hz rate or better is it possible to achieve audio latency of 10 milliseconds. That's a difficult thing to guarantee in a modern consumer OS. It can be done, but it won't happen by accident, only by design.
The penalty for failing to service your 10-ms buffer is a dropout that sounds much worse than slightly higher latency, so there's an incentive to use larger buffers than necessary at every link in the signal chain. From the point of view of the OS vendor, musicians might complain about latency, but everyone will complain about dropouts.
The easy fix to that is to make the buffers bigger. Doing that, though, increases latency. The hard fix is to do what Apple has done and build for audio from the ground up.
Using a "pull" method is required for low latency audio, where the audio driver's interrupts are scheduling when audio is passed to and pulled from applications.
Another downside of halving the period size is CPU load/battery drain. There are simply too many layers in Android's audio stack, and it has quite a few unoptimised code as well (such as converting between audio sample formats with plain C code).
Bad.
Nailing down buffer-bloat and sources of latency and jitter in an isochronous system took several months, the last few weeks of which were 80-100 hour weeks just before ship. Several times we thought we'd fixed the issues, only to find that our tests had been inadequate, or that some new component broke what we had built.
I remember nearly being in tears when I finally realized what the clock root of the audio system actually was, and that it wasn't what people had been using. From there everything fell into place.
Don't pull the number of buffers you have out of thin air ("Oh, six is enough, maybe twelve, don't want to run out in my layer of the system, after all...") Don't assume your local underflow or overflow recovery strategy actually works in the whole system. Don't assume your underlying "real time" hypervisor won't screw you by halting the whole god damned system for 20ms while it mucks around with TLBs and physical pages and then resumes you saying, "Have fun putting all the pieces of your pipeline back together, toodle-oo!" Put debug taps and light-weight logging everywhere and attach them to tests. And know what your clock root is, or you will be sunk without a trace.
Isoch is hard.
How did you finally figure this out -- any recommended references?
See... it's the little things like this. I am pretty sure that was not your intention, but please do know that turns of phrases like that hurt a little, and exclude a little.
To see what I mean, s/man/jew/ or s/man/white/ or some other category and see how it reads.
I look at "manly man" depictions of, well, manly men (for instance, watch Kevin Kline's performance in A Fish Called Wanda) as parody. Monty Python did it. Mark Twain probably did it. Do they offend people? Sure. Do people not get the joke? Oh yeah. Did that stop the artists in question? Not really. Am I comparing myself to great artists? I'm not worthy, but I study at the feet of masters.
So, in my writing (and I've done a bit of it; look at my profile and my blog) I tend not to give a flying unmentionable about who I piss off. Now, HN is different because it's not my ball game here, but hey, if everyone wrote so as not to offend then the world would be a dull place. I've opposed what I believe is the actual evil -- Political Correctness -- for decades, and I'm not going to stop now.
I like to take language out and give it a violent shake. Have fun with it, see what it can do. I'm not going to change the world by substituting something benign and utterly unoffensive for "manly-man," so fuck it. That's what down-votes are for. I'm not even sure what I'd put there instead, quite frankly. Your suggestions don't preserve the meaning at all. Kevin Kline would be sad.
I regret that you were offended.
Now, can we Godwin this stupid thread and get it over with? Someone toss in the grenade, okay? in 10, 9, 8 . . .
Then I'm sure you're also OK with an ever-shrinking subset of the world being willing to put up with that tone.
It all works out, I suppose.
Sadly, some users have to disrupt a thread, no matter how inane or off-topic it might be :(. Sometimes it gets to the point I don't even want to read or contribute to the discussions on HN anymore and add the site to my hosts list so I won't be tempted to read it out of habit. Eventually I come back, but the times in between get longer.
After skimming the comment history of the person you're replying to, I'd just ignore them, because they have a history of doing this. If HN had an ignore list, they would surely be on mine. Won't be the first or last time they disrupt (troll) a thread.
I read his comment as self-deprecating humor. If he had used white, jew, or whatever else may apply, it would still be self-deprecating humor, which is generally non-offensive for the simple reason it applies to the self.
I don't at all get what benefit there is to being able to say "manly-man programming," either from a communication or a social standpoint.
I don't get what's lost when we replace "manly-man programming" with anything less stupid and awful.
This is like any other literature. Don't force the world into newspeak.
Is this genuinely offensive and exclusive though?
You're sitting in a conference room full of women who are developers, and you start saying, "So let's do some manly-man programming and get this done," what are these women supposed to think?
So if your dev team consisted of a mix of men and women (not even counting if anyone identified as neither or somewhere in between), you'd still say "Let's put in some manly-man programming time" to them?
What if your dev team consisted of just one employee, a woman? Would you still tell her to put in her manly-man programming time? I'm having trouble seeing it. "Jane, it's crunch time. Can you man up and put in some manly-man programming time this weekend?"
If someone is looking to be offended then they will find a reason no matter what is said or how it's said.
Intent is the thing that should be judged, anything else is a foolish way to live and you will constantly find yourself offended.
Same with telling a woman that doing something "manly-man" style is doing it skillfully.
And even if you know it's going to be taken as humorous, let me give you a hint: it's not very funny. It's just not very good as humor. There is much better material out there if you're going for a joke.
I've been referred to as African American, Afro American, colored, person of color, black, negro, etc. It makes no difference to me personally. Unless intended otherwise, it's just a point of reference.
If, say, someone from another company who's visited your offices for a day later refers to you as "that negro in product development," would it be typical that other people of color would take no offense at that use of language?
I'm curious about it. I don't have that lived-in experience, so all I can do to understand is ask and listen.
See... it's the little things like this. I am pretty sure that was not your intention, but please do know that turns of phrases like that hurt a little, and exclude a little.
To see what I mean, s/kid/jew/ or s/kid/white/ or some other category and see how it reads.
I agree with you on that part but your substitution doesn't make much sense.
The original comment links programming competency to masculinity, which is a problem. It's as simple as that.
If that's what they meant, that's a pretty retrograde attitude.
If that isn't what they meant, maybe it's time to update their vocabulary.
When I feel offended, I treat it as an opportunity to learn about a different viewpoint. Sometimes it makes me change my mind. It always gives me a better understanding of people.
The different viewpoint, I guess, is that the original author has some outdated views on gender competence.
I've seen lots of comments saying "oh but it adds flair to language," and that's a reeeeeally weak defense. There is much better language available; that stuff just sounds stupid at best and retrograde at worst.
Sometimes, the world changes, and it's up to us to keep up with the times.
(It's funny, I bet you anything if the original author had said "this negro gentleman at my workplace...", no one would be defending his retrograde use of language.)
2.>original author has some outdated views on gender competence This is libel. This is not at all a conclusion that can be drawn from that user's posts. It's an unfair presumption and you're wrong to go around stating your opinions of others as facts.
On point 2: I can't even dignify that "libel" remark with a response.
Indeed. I'm glad at least one person got my point.
(A minor nitpick, tho: for the record I happen to have been born with two X chromosomes. Good guess though; especially on HN threads like these.)
Google: fuck fix the UI, fuck fix the VM bullshit, fuck fix bloating everywhere. NDK is not the solution: the problem is in the architecture.