Also, Apple is much more motivated to get the audio path right, having started with the iPods and selling a lot of music.
The reason ALSA and AudioFlinger add latency is to hide hardware-dependent differences as well as kernel-caused scheduling issues & policy decisions.
To achieve low-latency you need real-time scheduling, something Linux has with SCHED_FIFO but it's a bit kludgy, and getting the policy right on that is tricky (obviously you don't want a random app to be able to set a thread to SCHED_FIFO and preempt the entire system). So you have to restrict the CPU budge of a SCHED_FIFO thread, and you have to only allow apps to have a single SCHED_FIFO thread. But how much CPU time you give it needs to depend on the CPU's performance in combination with the audio buffer size that the underlying audio chip needs (and those chips also have different sample rates, is it 44.1khz or 48khz or etc...).
tl;dr: this is insanely hardware-dependent.
https://developer.apple.com/library/mac/documentation/MusicA...
Besides the fact that you are writing in C/Objective C, most heavy lifting is being provided by Audio Units which are specifically designed with common datatypes to be chained, composed and executed in low-latency situations.
Furthermore, most common things you would need in your app like mixing, conversion, timing, etc are provided as highly optimized services by the system.
http://www.quora.com/Why-do-iPhones-have-better-professional...
AudioFlinger + Alsa take a lot of time, as seen per the graph
But Android seems to take the option of least effort and "works most of the time" (which they have their reasons to)
This suggests that you can work focused on one driver and therefore save developer-time. However, the drivers on Android are made by a bigger workforce, which must be taken into account.
> they can skip one of the abstraction layers
The post explicitly states that the HAL ought to add no latency at all.
Consider e.g. AmigaOS.
AmigaOS let you obtain a pointer directly to the screen bitmap to update your window contents with no buffering or clipping. It could do that because originally all the hardware was the same, or close enough.
Then graphics cards came along, and you didn't necessarily have a way of writing directly to the bitmap. Suddenly you had to use WritePixel() and ReadPixel() and similar, which would obtain the screen pointer for the window, and obtain the display the screen is on, and find the driver corresponding to the screen, and call the appropriate driver function via a jump table.
Similarly, the AmigaOS had functions to e.g. install copper lists (the copper was a very primitive co-processor that could be used to do things like change the palette at specific scan lines), which also wouldn't work at all on graphics cards.
This is why knowing the hardware is part of a limited set matters: You can define your API to match the hardware very precisely, or even expose hardware features directly.
Given the wide scope of hardware targeted by android, it's not that surprising that it performs less well than a system targeting a very limited set of devices.
Having said that, it performs 'well enough' for the vast majority of use cases.
I don't know enough about android to say that this is the key aspect tho. Perhaps the low level API of android is the same? I didn't find an easy reference in a quick search.