Are you sure about that? We haven't seen what GPT-4 multimodal can do in the wild, it can even take into context the full conversation history in addition to just the images. If it can understand visual jokes easily, are you sure it can't detect CSAM?
We can also reliably GENERATE inappropriate content now, by simply adding a 'nsfw' tag to Stable diffusion, it flips a normal image to an inappropriate one. It doesn't sound very difficult to reverse this.
Also, for these services, you don't need it to be perfect. If even the flagging accuracy goes up significantly, that's a lot fewer human workers to review it.
The AI ecosystem as a whole is also booming massively regarding hardware, datasets, talent, software infrastructure. So that makes development in traditional ML faster.