I'm unfamiliar with 'pornography recognition' as an established task in ML research (lol), but for what it's worth, it's not an innovate use of CNNs for audio classification. You can essentially turn any audio classification task into an image problem (raw audio into features like spectrograms/MFCCs). Which people have been doing since forever (by which I mean a number of years now).