[0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...
[0] https://thezvi.substack.com/p/on-deepseeks-r1?open=false#%C2...
My guess is that it's heavily RLHF/SFT-censored for an initial question, but not for the CoT, or longer discussions, and the censorship has thus been "overfit" to the first answer.
I am not an expert on the training: can you clarify how/when the censorship is "baked" in? Like is the a human supervised dataset and there is a reward for the model conforming to these censored answers?
There are multiple ways to do this: humans rating answers (e.g. Reinforcement Learning from Human Feedback, Direct Preference Optimization), humans giving example answers (Supervised Fine-Tuning) and other prespecified models ranking and/or giving examples and/or extra context (e.g. Antropic's "Constitutional AI").
For the leading models it's probably mix of those all, but this finetuning step is not usually very well documented.
> You're running Llama-distilled R1 locally. Distillation transfers the reasoning process, but not the "safety" post-training. So you see the answer mostly from Llama itself. R1 refuses to answer this question without any system prompt (official API or locally).
It's probably disliked, just people know not to talk about it so blatantly due to chilling effects from aforementioned censorship.
disclaimer: ignorant American, no clue what i'm talking about.
CCP has quite a high approval rating in China even when it's polled more confidentially.
https://dornsife.usc.edu/news/stories/chinese-communist-part...
The indifferent mass prevails in every country, similarly cold to the First Amendment and Censorship. And engineers just do what they love to do, coping with reality. Activism is not for everyone.
The ones inventing the VPNs are a small minority, and it seems that CCP isn't really that bothered about such small minorities as long as they don't make a ruckus. AFAIU just using a VPN as such is very unlikely to lead to any trouble in China.
For example in geopolitical matters the media is extremely skewed everywhere, and everywhere most people kind of pretend it's not. It's a lot more convenient to go with whatever is the prevailing narrative about things going on somewhere oceans away than to risk being associated with "the enemy".
Wholeheartedly agree with the rest of the comment.
This is disingenuous. It's not "rewriting" anything, it's simply refusing to answer. Western models, on the other hand, often try to lecture or give blatantly biased responses instead of simply refusing when prompted on topics considered controversial in the burger land. OpenAI even helpfully flags prompts as potentially violating their guidelines.
False equivalency if you ask me. There may be some alignment to make the models polite and avoid outright racist replies and such. But political censorship? Please elaborate
IMO the first is more nefarious, and it's deeply embedded into western models. Ask how COVID originated, or about gender, race, women's pay, etc. They basically are modern liberal thinking machines.
Now the funny thing is you can tell DeepSeek is trained on western models, it will even recommend puberty blockers at age 10. Something I'm positive the Chinese government is against. But we're discussing theoretical long-term censorship, not the exact current state due to specific and temporary ways they are being built now.