The bias creeps in when they pick what sources go into the dataset.
The second biggest factor is RLHF because that's the point of it. All LLM providers aim to make models "safe" and safety in our modern discourse means not expressing ideas that contradict the "left ideology".