* audio developers recognizing the parallels between the way audio and video hardware works, and taking as many lessons from the video world as possible. The result is that audio data flow (ignoring the monstrosity of USB audio) is now generally much better than it used to be, and in many cases is conceptually very clean. In addition, less and less processing is done by audio hardware, and more and more by the CPU.
* video developers not understanding much about audio, and failing to notice the parallels between the two data handling processes. Rather than take any lessons from the audio side of things, stuff has just become more and more complex and more and more distant from what is actually happening in hardware. In addition, more and more processing is done by video hardware, and less and less by the CPU.
In both cases, there is hardware which requires data periodically (on the order of a few msec). There are similar requirements to allow multiple applications to contribute what is visible/audible to the user. There are similar needs for processing pipelines between hardware and user space code.(one important difference: if you don't provide a new frame of video data, most humans will not notice the result; if you don't provide a new frame of audio fata, every human will hear the result).
I feel as if both these worlds would really benefit from somehow having a full exchange about the high and low level aspects of the problems they both face, how they have been solved to date, how might be solved in the future, and how the two are both very much alike, and really quite different.