Wait. Is this a real output from the safe LLM? Ahaha.
Meta also released the uncensored base model, on which the open source community then performed its own chat fine tunes. This was a canny strategy to avoid negative press.
Mistral saw Meta’s approach, and instead chose to deliberately court the negative press, because attention is more valuable to them as a startup than opprobrium is damaging.
You could theoretically auto-suppress any tokens that lead to a refusal to answer path and re-generate.
But you’re probably best off grabbing an existing uncensored chat SFT, like the Llama v2 variants trained on Hermes. https://huggingface.co/NousResearch/Nous-Hermes-Llama2-70b
But Mistral does it well.