I dont think this is really possible, but probably not for the reason you think.
Noise cancellation works by picking up surrounding sounds using microphones and then generating the same sounds 180 degrees out of phase. So, we could get rid of the microphones processing sounds in real-time and instead we use something that can identify the song (like Shazam) as well as the position in the song, and then use reference audio for that song to play it back out of phase. So sure, that seems at least possible.
But the audio file itself _doesn’t describe exactly what you hear_. The acoustics of the room, your position in the room, the speakers, the amplifier, and I’m sure many other things all color and change what you’re hearing. In short, I’m not so sure that raw audio data out of phase _actually_ matches what you’re hearing in all situations. You’d still want a microphone.
- the people who grunt at the top of a rep
- the people who grunt throughout each rep
- the people who mutter the lyrics to the angry rap-rock in their heads
- giggling teenage boys who gather around the bench press and never do a rep
The loud 80s club style music is outdated and made sense before iPods / smartphones etc with individual custom playlists and thousands of songs on-demand.
Gyms should really provide headphones or Bluetooth or better yet online radio station to connect to if you ever want to hear their music (even remotely at home gym) and at your desired volume level, just like airplanes do! Quite airports like Dubai International is an awesome library-like experience especially when you have a long layover compared to regular noisy airports.
First problem: Your system needs to be perfectly synced up in time with the known sound being played. Human hearing goes to ~20KHz, so you need to be synced up with an error of <0.1 ms per second, probably much less. That’s gonna be tough.
Next problem: as you move around the room, your distance to the speaker changes, and you go out of phase. So you need positional data for your ears with sub-centimeter accuracy.
Can we get rid of the first two problems by using a microphone to detect the signal and auto-correct its phase as you move? Maybe but then you need an incredible algorithm to be able to pick out the correct part of the signal from all the background noise and echoes, plus it has to run in real-time at high sample rates. Very very difficult.
Then even if you solve these issues, you also have to cancel the echoes of the song bouncing off the walls etc. If you’re trying to take advantage of the fact that it’s a “known song” you then need an extremely good model of the room you’re in to predict how the sound will bounce.
In short, I suspect it’s “theoretically possible” but completely infeasible.
Regarding sync response time I don't think that's so much of an issue because after you recognize the song you know what the future waveform will look like in advance.
Like the woods?
The best I've found is ANC over-ear headphones along with brown noise, for my hearing profile. Pink noise may or may not work better for you, depending on your hearing abilities and the type of music being played.
1) put in silicone earplugs
2) wear noise cancelling headphones (Sennheiser seems good to me)
3) crank your music at 30% louder than you would normally
LIGHT WEIGHT BABY.