I think the key difference, though, is that the consumer video needs are probably in the area of two orders of magniture more computationally complex for video than for audio. On consume machines, almost no interesting real-time audio happens. It's basically just playback with maybe a little mixing and EQ. Your average music listener is not running a real-time software synthesizer on their computer. Gamers are actually probably the consumers with the most complex audio pipelines because you're mixing a lot of sound sources in real-time with low latecy, reverb, and other spatial effects.
The only people doing real heavyweight real-time audio are music producers and for them it's a viable marketing strategy for audio programmers to expect them to upgrade to beefier hardware.
With video, almost every computer user is doing really complex real-time rendering and compositing. A graphics programmer can't easily ask their userbase to get better hardware when the userbase is millions and millions of people.
Also, of course, generating 1470 samples of audio per frame (44100 sample rate, stereo, 60 FPS) is a hell of a lot easier than 2,073,600 pixels of video per frame.
I agree that audio pipelines are a lot simpler, but I think's largely a luxury coming from having a much easier computational problem to solve and a target userbase more willing to put money into hardware to solve it.