Your Open Source Model Could Have a Hidden Time-Release Backdoor
morgin.ai
morgin.ai
What's much more likely is that your US AI provider is promising not to train on your data but is doing so anyway. With a self-hosted model you can at least avoid that.
There's genuinely no evidence of this for OpenAI and Anthropic. It's impossible to disprove, but I don't think it's likely because:
* They get enough volume from consumer subs with data training enabled anyway.
* If this was happening, it needs serious work at the scale OpenAI and Anthroppic, from data pipelines, to ablation experiments, to the actual data mix and traces going in all the telemetry/diagnosis of large-scale training runs.
* It would need to involve a team. Employees at these companies leave, there have been numerous whistleblowers, allegations, etc. Nothing on this front that I can find.
* It would damage enterprise trust permanently and be a company and reputation-ending thing. Now that these tools are used by everyone from state governments to the DoW, the exposure radius is massive, investors (many of whom are customers/users too; and often have their stakes in not just a single company but multiple) would not be happy. Piss off enough powerful people, and anyone can join Sam Bankman-Fried in prison.
* There's a myriad of enterprise customers and bespoke contracts. I can't get into details, but not all enterprises accept a 'trust me bro' clause.
If something is impossible to disprove, we must assume it is happening from a threat modeling perspective.
1) an extremely high financial incentive to be dishonest (trillions of dollars),
2) low-ish chance of being caught, especially if you launder the data through another model to remove identifying information,
3) the people in charge of said operations are generally agreed to be snakes,
then it should at least arouse suspicion. You may be right that get enough data from opt-ins that they don't need to do it. I do believe that they don't violate enterprise ZDR agreements, but for normal subscriptions I'm much less confident.
- https://arxiv.org/abs/2311.14455
https://people.cs.umass.edu/~emery/classes/cmpsci691st/readi...
Unless the model can somehow reliably make a tool call to get the date (which would be suspicious and also easy to mock out)
So what is the proposed "fix" here if there is any?
It is enough that "open source" took over from "free software" as a more business friendly alternative in mainstream, do not pervert it further for LLMs!
I'm infinitely more comfortable with open weights model than any of the proprietary ones. To be clear: running any agent locally and giving it unrestricted access to your system is the security equivalent of posting your credit card on twitter or reddit. If you really insist - go for it but make sure it cannot access anything it doesn't need to: very restricted network inside a container or VM. Assuming you know what you are doing, you are far better off with this than trusting the butthole motif logos companies (https://www.creativebloq.com/design/logos-icons/why-do-all-a...)