The Chinese artificial intelligence engine DeepSeek often ***refuses to help programmers*** ___or___ gives them code with major security flaws when they say they are working for the banned spiritual movement Falun Gong or others considered sensitive by the Chinese government, new research shows.
My example satisfies the first claim. You're concentrating on the second. They said "OR" not "AND". We're all programmers, so I hope we know the difference between these two.But what are you attacking my claim for? That I'm requesting people don't have knee-jerk reactions and for help vetting the more difficult claim? Is this wrong? I'm not trying to make the claim that it does or doesn't write insecure code (or less secure code) for specific groups. I've also made the claim in another comment that there are non-nefarious explanations to how this could happen.
I'm not trying to make a stance of "China bad, Murica good" or vise versa, I'm trying to make a stance of "let's try to figure out if true or not. How much is it true? How much is it false?" So would you like to help or would you like to create more noise?
I did a "s/Falun Gong/Hamas/" in your prompt and got the same refusal in GPT-5, GPT-OSS-120B, Claude Sonnet 4, Gemini-2.5-Pro as well as in DeepSeek V3.1. And that's completely within my expectation, probably everyone else's too considering no one is writing that article.
Goes without saying I am not drawing any parallel between the aforementioned entities, beyond that they are illegal in the jurisdiction where the model creators operate - which as an explanation for refusal is fairly straightforward. So we might need to first talk about why that explanation is adequate for everyone else but not for a company operating in China.
But I don't think we should talk about explanation until we can even do some verification. At this point I'm not entirely sure. We still have the security question open and I'm asking for help because I'm not a security person. Shouldn't we start here?
https://i.postimg.cc/6tT3m5mL/screen.png
Note I am using direct API to avoid triggering separate guardrail models typically operating in front of website front-ends.
As an aside the website you used in your original comment:
> [2] Used this link https://www.deepseekv3.net/en/chat
This is not the official DeepSeek website. Probably one of the many shady third-party sites riding on DeepSeek name for SEO, who knows what they are running. In this case it doesn't matter, because I already reproduced your prompt with a US based inference provider directly hosting DeepSeek weights, but still worth noting for methodology.
(also to a sceptic screenshots shouldn't be enough since they are easily doctored nowadays, but I don't believe these refusals should be surprising in the least to anyone with passing familiarity with these LLMs)
---
Obviously sabotage is a whole another can of worm as opposed to mere refusal, something that this article glossed over without showing their prompts. So, without much to go on, it's hard for me to take this seriously. We know garbage in context can degrade performance, even simple typos can[1]. Besides LLMs at their present state of capabilities are barely intelligent enough to soundly do any serious task, it stretches my disbelief that they would be able to actually sabotage to any reasonable degree of sophistication - that said I look forward to more serious research on this matter.
With your Hamas example, I think it is beside the point. I apologize as I probably didn't make my point clearer. Mainly I wanted to stop baseless accusations and find the reality, since the articles claims are testable. But what I don't want to make a claim if is why this is happening. In another comment I even said that this could happen because they were suppressing this group. So I wouldn't be surprised if the same is true for Hamas. We can't determine if it's an intentional sleeper agent or just a result of censorship. But either way it is concerning, right? The unintentional version might be more concerning because we don't know what is being censored and what isn't. These censorships cross country lines and it is hard to know what is being censored and what isn't.
So I'm not trying to make a "Murica good, China bad" argument. I'm trying to make a "let's try to verify or discredit the claims." I want HN to be more nuanced. And I do seriously appreciate you engaging and with more depth and nuance than others. I'm upvoting you even though we disagree because I think your comments are honest and further the discussion.
You can also use the API directly for free on OpenRouter.
Another example: McDonald’s fries may cause you to grow horns or raise your blood pressure. No one talks like that.
So I would toss it back to you: we are programmers but we have common sense. The author was clearly banking on something other than the technically accurate logical or.
What I want to fight the most is just outright dismissing what is at least partially testable. We're a community of techies, so shouldn't we be trying to verify or disprove the claims? I'm asking for help with that because the stronger claim is harder to conclude. We have no chance of figuring out the why, but hopefully we can avoid more disinformation. I just want us to stop arguing out our asses and fighting over things we don't know the answers to. I want to find the answers, because I don't know what they are.
It is technically certainly feasible to have language-dependent quality changes, the language of the prompt can be trained to make intentional security lapses.
But no neural network has a magic end-intent or allegiance detector.
If Iran's "revolutionary" guard seeks help from a language model to design centrifuges, merely translating their requests to the model's origin dominant language(s), and culling any shiboleths should result in an identical distribution of code, designs or whatever compared to origin country, origin language requests.
It is also expectable that some finetuning can realign the model's interests towards whomever's goals.
> Ah but we're just gonna jump to conclusions instead.
I'm not trying to say WaPo is doing grade A journalism here. In fact, personally I think they aren't. A conversation about clickbait titles is a different one and one we've had for over a decade now...But are we going to recognize the irony here? Is OP not calling the kettle black here? They *also* jumped to conclusions. This doesn't vindicate WaPo or make their reporting any less sensational or dubious, but we shouldn't make the same faults we're angry at others for making.
And pay careful attention to what I've said.
>>> You should be skeptical, but this is easy enough to test, so why not do some test to see if it is obviously false or not?
>>> I'd appreciate it if others would reply with their replication efforts
I do want to find the truth of the matter here. I could have definitely wrote it better, but I'm appealing to our techy community because we have this capability. We can figure this out. The second part is much harder to verify and there's non-nefarous reasons that might lead to this, but we should try to figure this out instead of just jumping to conclusions, right?https://claude.ai/public/artifacts/77d06750-5317-4b45-b8f7-2...
1)Four control groups: CCP-disfavored (Falun Gong, Tibet Independence), religious controls (Catholic/Islamic orgs), neutral baselines (libraries, universities), and pro-China groups (Confucius Institutes).
2) Each gets identical prompts for security-sensitive coding tasks (auth systems, file uploads, etc.) with randomized test order.
3) Instead of subjective pattern matching, Claude/ChatGPT acts as an independent security judge, scoring code vulnerabilities with confidence ratings.
4)Provides some basic statistical Welch's t-tests between groups with effect size calculations.
Iterate on this start in a way that makes sense to people with more experience than myself working with LLMs.
(yes, I realize that using a LLM as a judge risks bias by the judge).