Full disclosure this was a few years ago so things may have improved. I also can’t remember all of the specifics but there was a lot of low-level driver work, firmware tweaks, specific configurations of just about every WiFi param you can think of, etc.
I wanna say a special welcome to everyone that's, uh, climbed into the
Internet tonight and, uh, has got into the M-bone. And I hope it doesn't
all collapse.
What a time to be alive.The article also indicates Internet Multicasting's "Geek of the Week", there's an archive at: https://town.hall.org/radio/Geek/ Transcripts available at http://opentranscripts.org/sources/geek-of-the-week/
At least a while back. I haven't touched it in years.
Wireless sure, that can be a pain.
Author here. Everything on the article was received on the guest Wi-Fi network on my laptop without plugging into anything.
Well it's a bit of a mystery to me as well.
The thing is multicast is not anycast, you will not receive multicast traffic unless you specifically ask to join a group.
So, he would have to be actually listening somewhere in between the multicast router and the elevator?
Most likely a dumb cheap switch that doesn't snoop on IGMP (or a less dumb switch that wasn't configured to snoop on IGMP) was upstream of both the OP and the device in the elevator. So the frames were being flooded since the switch doesn't know any better, and then normally ignored by OP's network stack since they didn't join that multicast group, until something like tcpdump/wireshark enables promiscuous mode.
For example you wouldn't want a wifi speaker in an elevator using a repeater at the top of the shaft trying to match up to a hardwired speaker in a ground floor vestibule.
And then you "just" have the same problems that you have with purely electrically connected, analogue speakers (which are effectively 100% in sync in terms of receiving the signal): Sound is relatively slow, and so the audio from a speaker that is far away will reach you later than the nearby speaker.
You can mitigate that by adding a precise delay to the far away speaker... but of course that does not work if you're standing on the other side. Nevertheless, as said, that problem is regardless of whether your speaker is network-connected or not.
So to work well you do need to resync the audio to the local audio clock using a sample rate converter, or build some custom hardware that lets you sync the playback audio clocks somehow. Or if you want to be sloppy about it, keep close track and stuff or drop individual samples as you drift.
But yeah, this is all more or less 'solved'.
For URL-based streams they buffer and NTP to sync. For live streams (e.g. gaming) they p2p multicast and tweak the wifi params in real-time to minimize drops.
The speakers create their own wifi and use MST network heuristics to latency-min route over that versus native wifi or ethernet if you've plugged it in. Sound drops when the wifi spectrum blinks (rarely), but I have never encountered the speakers being out of sync or noticing an echo effect.
And the speakers can use your phone's mic to scan the soundscape of a room to acoustically balance the sound when you set them up. I particularly like how consistent the sound volume is room-to-room even with very different speaker setups.
IIRC they've patented their specific mechanism. So ya, it's solved, but it may be expensive to license.
(Not affiliated with Sonos, I just have a bunch of them and like them a lot.)
If you are just interested in the synchronized Audio-over-Ethernet part, AES67 is the industry standard, and a pretty complete open-source implementation can be found at https://github.com/bondagit/aes67-linux-daemon , though AES67 is itself a composition of existing standards, fundamentally it is mostly composed of SDP for sessions description, RTP for media, and PTP for clock sync, so you can build that out of a variety of implementations too.
For room correction you can look at https://drc-fir.sourceforge.net/ to generate FIR filter coefficients, then you can apply it in realtime with https://github.com/wwmm/easyeffects or https://github.com/HEnquist/camilladsp .
Of course some people just want it to work, then you can shell out for Sonos :p.
The simplest model is a source that generates a continuous audio stream, and a sink that plays it back; adding the idea of songs complicates the model, and in some use cases might be totally inappropriate. For elevator music, sure it likely doesn't matter, and maybe you can hide it in a crossfade or something with enough metadata, but this is probably part of a system where you put audio into one device connected to the network, that might include live stuff like PA announcements, and it comes out a bunch of other ones, not a dedicated elevator music system.
You only care about delay between the speakers, not about what latency any speaker has relative to the source.
NTP is accurate enough for this, but I think most of the modern protocols in the wild e.g. AES67, AirPlay2 are using PTP. It is both more accurate and in some ways simpler for this use case.