A hypothesis I have is that it is much more difficult to keep in line with the good alignment than to do evil. In the limited context window of an LLM, one wrong move would make the model evil, no matter how many good tokens it generates.
Setting aside the difference between Human intelligence and LLM, we can tentatively attribute the mostly good human behavior to a life time of context length, within which we trained ourselves to do good, while the RLHF for a limited context length LLM lack such continuous reinforcement within a big context.