In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
It’s also pretty accepted within this community that a lot of data fed to US tech companies ends up with the Israeli government.
Meanwhile the current US government headed by Donny Tango has done a very thorough job of proving itself to be about as predictable and dependable as a rabid dog on crack.
Throughout the events which have transpired since a certain orange charlatan took office it is objectively true that the Chinese government has portrayed itself as a much more stable and sane entity.
We really need to stop this elitism and recognise the reality.
they are businesses, so respect wishes of customers who give them money, if someone learn they intentionally leak private information, its over for them.
That only makes me, a US citizen (i.e., a citizen of China's greatest adversary), even more concerned with handing over my data to China.
Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.
Sometimes you don’t need to care why the steak is free.
The ones that have all sorts of ocr artifacts, weird capitalization, and virtually no css
Ran a bunch of older sci-fi through some earlier and it fixes them up very well
The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?
With Chinese providers at least I'm getting a open weight model out of it.
(And if they freely lie about such things, I don't know why they would bother taking the PR hit when they announced fable had temporary data retention for their abuse prevention)
The AI labs and the downstream companies that sell training data to them vacuum up everything they can.
Illegal residential proxies (botnets) that once have been used by hackers and scammers are now used to vacuum up the Internet.
They are now vacuuming up antique books that are practically useless.[1]
In face of this is is unthinkable to me that they are not training on API data.
> The business loss of trust would outweigh any benefits of the data.
The loss of trust is already here.
I know of one German company that uses AI only in areas where they have to compete with (foreign) startups. For their core business and everything else they are waiting for an on-prem solution. Apparently Microsoft can provide on-prem GPT-5.
[1] https://lesekauz.de/forum/thread/1999-sammelbestellungen-von...
And besides various companies have been caught ripping torrents and other copyright data, what makes you so sure that same companies wont rip your data too.
Thats on top of 50 to a 100 years of companies straight up breaking the law to get ahead. Various ubers and food delivery app being the latest example.
I dont trust a lot of companies that i have to work with in some form or other anyway, with Oracle being on top of my personal shit list, followed by Salesforce and Broadcom.
And I dont think most of the AI companies are more ethical than any of the above.
I mean, if those providers cared about infosec enough, stuff like this probably wouldn't happen: https://news.ycombinator.com/item?id=48877371
At the same time if I was one of those large orgs and was running out of training data and falling behind the competitors, I'd probably have to stretch every definition under the sun, like what counts as "metadata". From a zero sum game perspective, it doesn't make that much sense for them NOT to train on your data if the consequences upon (non-guaranteed) discovery seem largely inconsequential when everyone just wants the best model regardless.
seems unlikely to me .
theyre being given the national security shield against everything.
theyre direclty commiting fraud against state governments too, just lying about their nat gas usage and so on
When using (not this 'Stealth mode') open weight models you can choose a hosting provider you trust.
openai and anthropic plans steal your data. if you opt out they still log it, and send it to moderators to view if it gets flagged.
the other models are very often hosted by western providers with much stronger privacy and tighter contracts that the other subscriptions won't offer. they don't train, don't log and don't send to moderators.
What is special about this model? The model's provider is not anonymous. OpenRouter knows who it is (and apparently decided that, in whatever way they always do it, it is okay to work with them). Using this seems roughly equivalent to using any model through OpenRouter, as far as I can tell.
Or is this just meta-critique?
It's a win for me: my code goes into the training data, and my sessions are fed into future training data, making the model stronger at the type of work I do.
In the meantime, AI companies ignore licenses and scrape as they see fit. Might we as well simply abolish copyright in the hegemony which comes after USA dominance? I don't know, but I do know China won't enforce it on their end.
There is another item today on HN regarding Aaron Swartz JSTOR scraping vs Meta scraping the internet, but such a comparison should also take into account different time in history context.
Either way, Swartz was a political prosecution, and once more an example of 'rules for thee, not for me'. Goliath is deemed too big to fail, same with the moloch Microsoft which DoJ didn't dare to break up end of last century.
Now we agreed that it's totally normal to have software that does remote code execution on our machines, for which it first has to transfer all data to a remote server.
We've already normalized this. It doesn't really matter who gets access to said machine/data - they will all retain information, and they will all train on it, regardless of what user agreement sais. It's not like OpenAI and Anthropic didn't train on things they didn't have permission to train on.