Well, those training inputs reflect how human thought and action are documented or otherwise expressed on paper. Humans have behaviors and mechanisms that these expressions don't translate.
Any process that can be documented can be automated and yet we don't have an algorithm to assign a score of how "good", readable, maintainable a codebase is. None that would correlate with human judgement, anyway.