Modern open source LLMs are still RLHFed to resist adversarial output, albeit less-so than ChatGPT/Claude.
They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.
They all (with the exception of DeepSeek) can resist adversarial input better than Grok 4.1.
Quality of response/model performance may change though
There’s also nous research’s Hermes’ series of models, but those are trained on llama3.3 architecture and considered outdated now