YouTube has the advantage of being able to preprocess the audio track in its entirety, although most implementations work on a block basis and don't need that. A 'night mode' feature would take about one weekend's worth of work from start to finish.
You would also need a fairly sophisticated content detection system (don’t distort anything that is music, which is a huge use case on YouTube for example).
No, you need a button that the user can press to turn the feature on and off. It should default to 'off.' For extra points, make it a slider.
Not to mention adding _any_ steps to their intake pipeline requires more processing power and data moving than you and I will ever touch in our lifetimes.
Compared to the work needed to process the video itself, it would be a total non-issue.
On the server side, they could maintain a separate compressed audio track and switch to it when requested. They should have the ability to do that anyway for alternate languages and commentaries and such. Whether it would be more cost-effective to maintain a separate set of compressed tracks or to decode/compress/re-encode on the fly, I don't know.
It might be tempting to maintain caches of compressed tracks created on demand. That way they'd almost never incur a runtime penalty on the server, and also wouldn't have to pay a lot for storage. More than a weekend's worth of work at that point, though.