Detecting if there is a child in the video is more difficult, but do-able with current ML models. Now, however, determining why there is a child in the video - if it is a family "fail" video or pedo material, for example - via AI is about as impossible as trying to distinguish between satire, hate speech or propaganda via AI. It's not possible at all, as AI will for the near future totally lack context.
This distinction will require humans, and this is something not viable at all for fb, youtube, twitter & co, as it is a huge cost... the saving of which is offloaded to society though in form of e.g. undermined democracies or psychological trauma in sexual violence survivors.