It's not the training dataset.
All of these models, including the "open" ones, have been RLHF'ed by teams of politically-motivated people to be "safe" after initial foundation training.
All of these models, including the "open" ones, have been RLHF'ed by teams of politically-motivated people to be "safe" after initial foundation training.
This corruption must be disclosed as assiduously as the base dataset, if not more so.
Actually, those seem like an apt composite description of the PoV of the typical mass-market AI... 8/
Try it for yourself: https://huggingface.co/mistralai/Mistral-Large-Instruct-2407