I wonder if the before and after could be used to train a computer to do this automatically, well, for the future. Obviously not the sound, but the picture.
It looks like there are previous attempts [2] that have even been turned into APIs [3]
[1] https://news.ycombinator.com/item?id=18363870