I've been working on VoIP and Internet multimedia since 1991 (I'm one of the authors of SIP, for my sins) so, yes, we've been concerned about time-sensitive material for a long time. From a standards point of view, this resulted in first, Interv and RSVP, then diffserv. The former largely failed, mostly for economic reasons, while the latter is widely used, but almost never for end-to-end service, again for economic reasons. Basically, if you're going to give priority to some traffic, there needs to be a disincentive for all traffic to demand priority, which usually means charging. No-one ever figured out a good economic model for end-to-end QoS, and most users aren't willing to pay extra anyway.
There has been a lot of effort more recently on sensible queue management schemes to avoid buffer-bloat, which show significant promise. For low-speed customer links, variants of fair queuing generally work well. My ISP, Andrews and Arnold, is one of the few that really gets this. They have a shaper, upstream of the DSL or FTTC link, that can be configured to 95% of the negotiated link speed, so there are no queues downstream. Then they run fairly small buffers in the shaper and WFQ. I pretty much never see any significant latency on my VoIP calls due to queuing.
In general, low-latency VoIP shouldn't be a problem - so long as the data rate is lower then the rate of the mean flow on a link, you can give it strict priority service without harming the other traffic. In practice, this requires some sort of policing to avoid abuse. Various forms of fair queuing approximate this - a flow that has a low data rate will be serviced frequently (though not as quickly as it it had priority), and see pretty low latency in the queue on a busy link.
Now, getting low latency for high quality video conferencing; that's a different question. Fair queuing doesn't really help, because the data rate can be relatively high. Instead, you've got to keep the total queue size small. This can impact competing TCP flows if you're not careful, and can cause non-negligible packet loss. Thus is due to TCP's congestion control dynamics - throughput of a TCP flow is inversely proportional to round trip time, and inversely proportional to square root of loss rate. If you reduce queuing at the bottleneck, you reduce RTT, so push up the loss rate. This is one of the reasons people gripe about TCP's congestion control. However, if ECN were widely deployed, this would become a non-issue, because we wouldn't need those packet drops to control TCP's sending rate. Generally, widely deploying ECN would be the single biggest enabler for low-latency networking.