I dunno, wasn't Saint Augustine "lord make me good, but not yet" pretty much admitting that our heroes have feet of clay, and yet they function as educators and leaders.
If you modulate the training set through externally derived axioms of good and bad, can't you train an AI on objectively naughty data to recognise the anti-set of good behaviour?