https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
"Safety" is something asserted by the model creator, not something asked for by users.
More power to you if that is your plan, but most of us want to use the models for things that are less contentious than the things people put into chatbot arena in order to get commercial models to reveal themselves.
-
I'd honestly we rather just list out all the NSFW prompts people want to try, formalize that as a "censorship" benchmark, then pre-filter chatbot arena to disallow NSFW and have it actually be a normal human driven benchmark.
Corporate users of AI (and this is where the money is) do want safe models with heavy guardrails.
No corporate AI initiative is going to use an LLM that will say anything if prompted.