Reverse engineering a mysterious UDP stream in my hotel (2016)
gkbrk.com
gkbrk.com
Joke's on you, it's a bug listening to you in your room while using steganography to merely appear to be elevator music!
Again, if I were in a position to be concerned - I'd move hotels with Wireshark actively monitoring and verify the network traffic dropped when I left the WiFi range, and also what kind of UDP/network traffic was at the next hotel.
But if I were in a position to be extremely concerned, I'd probably just throw everything away to begin with, including the clothes I was wearing, buy a laptop/ new clothes, and then, after escaping out the back of one of the stores and getting picked up by a random taxi service and driven a good distance, go to a hotel without an elevator and check the Wireshark traffic.
Unless it was some sort of very sophisticated monitoring, I would hope one of these strategies would provide some answers.
To a casual observer snooping the relevant network, it would probably (as here) look as elevator music, but to the intended recipient who can decode the steganography, it would be a covert listening device.
Or you could have an extremely long audio file so the repeat situation doesn’t occur.
The decompressed data is always the same, but the data in the dictionary used is where you store your sneaky bits.
Sure, that's still mildly suspicious. But way less than the actual music data changing all the time.
You can use a channel with lots of noise, because you can use error correcting codes to to restore the intended message.
(To elaborate with an example: sometimes a package might already drop randomly, or timings might be slightly delayed anyway.)
I guess you could also build an auto-auto-tune or auto-remix solution so the songs always have justifiable variation without needing fully generated music.
-edited because it didn't make sense before
Not just decode, but also decrypt.
You'd probably want to encrypt not just for the secrecy, but so that the noise introduced by the steganography doesn't seem so suspicious.
If the music was repeating but the stream was different all the time, then steganography could be the reason :)
In fact I don't think I've every been on an elevator that had music.
Google sez: Music for Airports was installed at the Marine Air Terminal of New York's LaGuardia Airport for a brief period during the 1980s.
Elevator music may no longer be a thing, but apparently, airport music still is:
Basically, you'd use steganography to hide that you send a message. And that message would be encrypted. You can use almost any standard encryption scheme, as long as you remove headers etc. Any ciphertext of a crypto-system worth its salt will be indistinguishable from random noise without the key.
Weirdly, my Sonos/Unifi issues are gone at the moment, but wow is it painful when you hit it.
So multicast doesn't necessarily mean that every device gets every packet... I think.
Also probably not worth using IGMP for audio.... :-)
The important thing about why you'd use IGMP Multicast instead of unicast is that you get a more assured latency which is important for audio sync.
ffmpeg -stream_loop -1 -re -i gettysburg.wav -f mp3 udp://239.0.0.1:1234
on one device, and ffplay -i udp://239.0.0.1:1234
on another device.You can also play the stream with ffmpeg doing something like
ffmpeg -i udp://239.0.0.1:1234 -f pulse defaultI've experimented a tiny bit with the builtin pulseaudio RTP sink/source, but I've not used it extensively. I'm not sure if cutting out ffmpeg would be beneficial.
Managed to get the lag down to a usable amount for algorithmic music, but would not be suitable for playing an instrument remotely
Such a powerful tool... And thank you for letting us know about this trick.
Full disclosure this was a few years ago so things may have improved. I also can’t remember all of the specifics but there was a lot of low-level driver work, firmware tweaks, specific configurations of just about every WiFi param you can think of, etc.
I wanna say a special welcome to everyone that's, uh, climbed into the
Internet tonight and, uh, has got into the M-bone. And I hope it doesn't
all collapse.
What a time to be alive.The article also indicates Internet Multicasting's "Geek of the Week", there's an archive at: https://town.hall.org/radio/Geek/ Transcripts available at http://opentranscripts.org/sources/geek-of-the-week/
At least a while back. I haven't touched it in years.
Wireless sure, that can be a pain.
Author here. Everything on the article was received on the guest Wi-Fi network on my laptop without plugging into anything.
Well it's a bit of a mystery to me as well.
The thing is multicast is not anycast, you will not receive multicast traffic unless you specifically ask to join a group.
So, he would have to be actually listening somewhere in between the multicast router and the elevator?
Most likely a dumb cheap switch that doesn't snoop on IGMP (or a less dumb switch that wasn't configured to snoop on IGMP) was upstream of both the OP and the device in the elevator. So the frames were being flooded since the switch doesn't know any better, and then normally ignored by OP's network stack since they didn't join that multicast group, until something like tcpdump/wireshark enables promiscuous mode.
For example you wouldn't want a wifi speaker in an elevator using a repeater at the top of the shaft trying to match up to a hardwired speaker in a ground floor vestibule.
And then you "just" have the same problems that you have with purely electrically connected, analogue speakers (which are effectively 100% in sync in terms of receiving the signal): Sound is relatively slow, and so the audio from a speaker that is far away will reach you later than the nearby speaker.
You can mitigate that by adding a precise delay to the far away speaker... but of course that does not work if you're standing on the other side. Nevertheless, as said, that problem is regardless of whether your speaker is network-connected or not.
So to work well you do need to resync the audio to the local audio clock using a sample rate converter, or build some custom hardware that lets you sync the playback audio clocks somehow. Or if you want to be sloppy about it, keep close track and stuff or drop individual samples as you drift.
But yeah, this is all more or less 'solved'.
For URL-based streams they buffer and NTP to sync. For live streams (e.g. gaming) they p2p multicast and tweak the wifi params in real-time to minimize drops.
The speakers create their own wifi and use MST network heuristics to latency-min route over that versus native wifi or ethernet if you've plugged it in. Sound drops when the wifi spectrum blinks (rarely), but I have never encountered the speakers being out of sync or noticing an echo effect.
And the speakers can use your phone's mic to scan the soundscape of a room to acoustically balance the sound when you set them up. I particularly like how consistent the sound volume is room-to-room even with very different speaker setups.
IIRC they've patented their specific mechanism. So ya, it's solved, but it may be expensive to license.
(Not affiliated with Sonos, I just have a bunch of them and like them a lot.)
If you are just interested in the synchronized Audio-over-Ethernet part, AES67 is the industry standard, and a pretty complete open-source implementation can be found at https://github.com/bondagit/aes67-linux-daemon , though AES67 is itself a composition of existing standards, fundamentally it is mostly composed of SDP for sessions description, RTP for media, and PTP for clock sync, so you can build that out of a variety of implementations too.
For room correction you can look at https://drc-fir.sourceforge.net/ to generate FIR filter coefficients, then you can apply it in realtime with https://github.com/wwmm/easyeffects or https://github.com/HEnquist/camilladsp .
Of course some people just want it to work, then you can shell out for Sonos :p.
The simplest model is a source that generates a continuous audio stream, and a sink that plays it back; adding the idea of songs complicates the model, and in some use cases might be totally inappropriate. For elevator music, sure it likely doesn't matter, and maybe you can hide it in a crossfade or something with enough metadata, but this is probably part of a system where you put audio into one device connected to the network, that might include live stuff like PA announcements, and it comes out a bunch of other ones, not a dedicated elevator music system.
You only care about delay between the speakers, not about what latency any speaker has relative to the source.
NTP is accurate enough for this, but I think most of the modern protocols in the wild e.g. AES67, AirPlay2 are using PTP. It is both more accurate and in some ways simpler for this use case.
Ah no, oops, you had it right, force of habit ;)
"Oh for fuck's sake".
Reverse engineering a mysterious UDP stream in my hotel (2016) - https://news.ycombinator.com/item?id=26633792 - March 2021 (86 comments)
Reverse Engineering a Mysterious UDP Stream in My Hotel (2016) - https://news.ycombinator.com/item?id=16197436 - Jan 2018 (15 comments)
Reverse Engineering a Mysterious UDP Stream in My Hotel - https://news.ycombinator.com/item?id=11744518 - May 2016 (181 comments)
Wouldn't binwalk do just that?
That said, it's considered a best practice (although not really all that common) to use ipsec or another method to provide cryptographic authentication of multicast packets. The protocol discussed here may do so.
Play some good old Rammstein, to make your elevator journey more pleasant.
Anyway, I ripped the discs, added a nice 50Hz hum under the music and burned new copies which I then left by the CD player.
Yup. Cue frustrated sound engineers trying to debug the ground loop which only manifested itself when the CD was playing.
I never dared admit to the prank, but rather swapped the hum CDs for the originals before someone got a chance to investigate this thoroughly...
but only broadcasted in the room 147
Despite its simplicity, seems to be overengineered.
Because distributing music over multicast works very well, and it's rather simple in nature: You get audio frames as UDP packets, you play them. Because it's multicast, they reach where they should.
> why not just attach it to a radio directly
Because you lose control of what's being played, and there are probably not many "elevator music" radio stations.
> or some RPi that will pull the stream
Because then every location that plays the music has to pull an individual stream, quickly saturating the bandwidth for no reason, instead of just subscribing to the multicast stream.
I also would not say "some RPi" is less engineering, and less maintenance, than the current solution that likely uses an off the shelf multicast audio system.
> or just have it prerecorded
Because you lose control of what's being played without a significant replacement step, when you could just play a multicast stream instead.
> Despite its simplicity, seems to be overengineered.
I'm honestly not convinced I read a less "overengineered" solution so far...
> Because you lose control of what's being played, and there are probably not many "elevator music" radio stations.
Radio is still by far the simplest and most fail-safe solution. But not playing an existing radio station. Rather buying a radio transmitter, like the ones used in churches - strong enough to emit FM throughout the hotel, and yet not requiring a radio license. Then the speakers would just be plain speakers with a FM receiver attached - no more wifi/wired infrastructure involved, routers, ability of the speakers to connect to the wifi, ...
Prerecorded music on an MP3 player sounds like the one simpler solution, though it wouldn't put the elevators in sync with each other or let you control it remotely.
In this hotel scenario, you have one audio source that can be located in a location that can be conveniently tied into the fire alarm relay panel. Stop the audio there and you've stopped it everywhere.
The advantage of IP distributed audio is that it functions over the existing IP network, so it avoids the need to wire a high-voltage audio system (which will still require multiple amplifiers in large buildings) or dedicated signal-level wiring to distributed amplifiers. It also tends to be more reliable as the IP network is more robust to issues like crosstalk and poor connections that can turn into frustrating troubleshooting on analog systems.
"Some RPi that pulls the stream" is exactly what this system is except it uses multicast for significantly reduced bandwidth usage and the receivers are presumably commercially-supported distributed audio equipment. No hotel wants to be doing patch management for embedded Linux devices.
With broadcast, all the nodes and edges transport the stream only once, but you can only distribute the stream in the same (L3) network, and within that network you transport the stream even through L2 nodes and edges (i.e. network switches, cables, Wi-Fi channels) where there are no listeners at all. You could in theory repeat those broadcast packets into other networks, but you need to set that up more explicitly, and then every node and edge in that network gets the traffic, too.
With multicast, devices and effectively segments can subscribe to the stream, and all nodes and edges transport the stream at most once, and they can do so across network boundaries while retaining the same property, and using a standardized mechanism that is designed to make the traffic only go where it needs to go as much as the topology allows.
With multicast, a device that wants to receive the stream uses IGMP to join the group. The IGMP communication is not with the server, but actually with the router serving the subnetwork. Additionally, larger commercial switches usually implement "IGMP snooping" in which the switch "listens in" on IGMP sessions between its clients and the upstream router. The server maintains only one connection and sends only one copy of the stream, to a multicast group address---there are IP ranges reserved for this purpose. The router, and switches which implement IGMP snooping (or layer 3 switches, there are some variations), forward traffic to the multicast group address on any interface on which a client has used IGMP to join the group. Switches without IGMP support will just forward it on all interfaces. The result is that, at each point in the network, only one copy of the stream is handled. This has significant performance benefits for both the server and the network devices.
As the name implies, multicast is much like broadcast except that devices can opt in or out of receiving the broadcast, and network devices can use knowledge of those IGMP sessions to avoid sending multicast traffic on interfaces where no one cares to receive it. That said, multicast traffic going to network segments where it's not used is not especially harmful besides wasting some capacity on that network segment.
Unfortunately multicast is not workable over the internet for reasons which are difficult to overcome, or IPTV and other synchronous media streaming services would be far less costly to run. On an institutional network, though, multicast can be used to great effect. It's a fairly old method as well. In my first IT job we used to install the operating system on workstations by PXE booting them to an imaging tool and then multicasting the disk image across the network... this way we could do hundreds of machines at once at just about disk saturation rate. This kind of thing isn't readily achievable without the use of multicast or broadcast traffic. The tool was Norton Ghost, which has apparently supported this mode of operation since 1998!
First one being much worse as it shows a failure to understand network segregation.
I forget if the details of how exactly they accessed it, but it was an example of an Internet of Things device making a security hole in a network.
https://www.washingtonpost.com/news/innovations/wp/2017/07/2...
I think you're mistaken/misremembering. It was a casino.
https://www.entrepreneur.com/business-news/a-casino-gets-hac...
...yes?
Do you think hotels aren't protected, to some extent, against any random person putting on a blazer and announcing an emergency?
How are they protected against that, exactly? You can literally walk up to any fire emergency button on any wall on any floor, press a button and evacuate the entire hotel, why bother with this UDP streaming nonsense?
Cameras near fire alarms and it's a crime in the U.S. to give a false alarm.
Same way sheepdogs herd sheep.
Edit:
"Unauthorized" computer access is a serious federal crime under the CFAA, and that you did it as a joke is not a legal defense. Famous examples:
(1) https://en.m.wikipedia.org/wiki/Aaron_Swartz
(2) the Florida man who social engineered Twitter (https://en.m.wikipedia.org/wiki/Graham_Ivan_Clark)
(3) the Mirai botnet guys (https://en.m.wikipedia.org/wiki/Mirai_(malware)), etc.
So the penalty will actually be much worse if you get caught.
2. I'm not claiming that all attacks will be mitigated, but that this is an easy win from a cost:benefit analysis.
Although I could be wrong!
Yeah, with a better/better configured switch it could not reach the individual client unless it subscribed to the multicast stream, but again, who cares about that low bandwidth stream... you should not trust any hotel network anyway, those audio packets won't do extra harm.
Why? It's highly likely they just don't care. Which is fine. It's an elevator speaker.
Sure it's an elevator speaker and probably you can't do much (probably! 'simple' devices are packing ever more hardware).
But something is sending those packets. What's that? Can it be hacked? What else can that device/server see? Are there other devices sharing the network with customers?
did you play it reverse? I bet there is some hidden message…