Microsoft Probing If DeepSeek-Linked Group Improperly Obtained OpenAI Data
bloomberg.com
bloomberg.com
These are the sort of litigious death throws that the threatened orthodoxy throw out when they smell their own blood in the water.
A Trillion dollar write down will do that to you.
Let's just hope that TOS violation is considered a contract breach, if it even is a valid contract, and not anything crime related.
I don't want the state to expend any taxpayer dollars to fund civil offense prosecutions. The music industry has already managed to push their enforcement onto taxpayers via laws such as DMCA etc. Can't have this be done again, and with wider reach.
Their strategy right now is to socialize the negative externalities (copyright does not protect anything from them) and capitalize the gains (ToS says you can't use their outputs).
PS: I do understand your point. Using taxpayer money would be even worse.
And the US is an official ally, whereas China is not.
"Hey, ChatGPT, what are the terms of service for Spotify?"
"I can't provide the full, up-to-date terms of service for Spotify, but I can give you an overview of some key points typically included. For the most accurate and comprehensive information, I recommend visiting Spotify’s official website and reading their Terms of Service directly. That said, here are some common points that are usually part of streaming platforms' terms:
1. *Account Creation and Use*: Users need to create an account to access Spotify’s services, and must provide accurate, up-to-date information. They’re responsible for maintaining the security of their account.
2. *Subscription Plans*: Spotify offers different subscription tiers (Free, Premium, Family, etc.). Terms often specify the limitations of each tier, such as access to ads, offline listening, and music quality.
3. *Content Usage*: Spotify provides users with access to music, podcasts, and other content. However, the usage is typically for personal, non-commercial purposes only. Users are not allowed to download, share, or redistribute content unless Spotify allows it (like through the Share button or via a playlist).
4. *Licensing and Intellectual Property*: Spotify owns or has licenses for the content available on the platform. Users are granted a limited, non-exclusive license to access and use the service, but they don’t own the content they listen to.
5. *Privacy*: Spotify collects data, including usage patterns and preferences, which they may use for things like recommendations, advertisements, and analytics. Their privacy policy will detail what information is collected and how it’s used.
6. *Cancellation and Termination*: Either party (Spotify or the user) can terminate the service agreement. Users can cancel their subscription anytime, but Spotify may also suspend or terminate accounts for violations of their terms.
7. *Disclaimers and Limitations of Liability*: Spotify usually limits its liability in case of issues like service interruptions, data loss, or content inaccuracies.
8. *Modifications*: Spotify may update or change its Terms of Service at any time, and users are typically notified of these changes. Continued use of the service after changes indicates acceptance of the updated terms.
9. *Governing Law*: There’s often a clause specifying which country’s laws govern disputes related to the terms, and how disputes will be handled (for example, through arbitration).
If you're interested in the full details, you can check Spotify’s official site or app, where their Terms of Service are readily available for review."
openai did not agree to the TOS of other websites they scraped, so they're not bound by it. It's unclear if the TOS is automatically agreed to by virtue of merely accessing the data that is already supplied in a webpage, without having first actively agree by an action like clicking.
See https://en.wikipedia.org/wiki/Shrinkwrap_(contract_law)#Unit...
therefore, openAI is only ever bound by copyright law, and not the TOS of the website. And according to some interpretation of copyright law, these AI models do not constitute distribution of the original training material.
Similarly, when you "buy" many ebooks, you're agreeing to a TOS, but then, in violation of the TOS, they were uploaded to Libgen, which OpenAI downloaded.
IE, Indeed, "improperly" here is exactly the same "improperly" that a site owner applies to any bulk downloader.
So... Scrape not lest ye be scraped, one law for me... etc
I think MS was trying to use the scariest term they could come up with. To belabor the point a bit, "we train, you exfiltrate" etc.
https://news.ycombinator.com/item?id=42561419
Deepseek promptly fixed it so that their UI responds with 'Hi! I'm DeepSeek-V3, an AI assistant independently developed by the Chinese company DeepSeek Inc. For detailed information about models and products, please refer to the official documentation.' - but only if you ask that as the first question of the conversation. Bury the 'what model are you' question after a few unrelated questions, and it'll happily tell you it's ChatGPT.
Is it possible that it's an actual distillation of weights, but into a radically different architecture? We don't have evidence of that, but that would be a great technical feat in itself.
Is it trained on a large set of user requests and OpenAI replies? Yes.
The question is, were these obtained by simply using the API contrary the user agreement at scale, or was there access to internal OpenAI datasets, or was there some kind of capture of conversations by a man-in-the-middle (which could be any of a number of AI access resellers)?
The answer hinges on which _requests_ were in that training set, something that won't be easy to investigate - unless you're OpenAI itself, and can identify 'trap streets' in the archive of all conversations, cases where ChatGPT once gave an unusual response to an unusual request, and DeepSeek just happens to match it.
Why DeepSeek’s new AI model thinks it’s ChatGPT...
https://techcrunch.com/2024/12/27/why-deepseeks-new-ai-model...
>Posts on X — and TechCrunch’s own tests — show that DeepSeek V3 identifies itself as ChatGPT, OpenAI’s AI-powered chatbot platform. Asked to elaborate, DeepSeek V3 insists it is a version of OpenAI’s GPT-4 model released in 2023.
>The delusions run deep. If you ask DeepSeek V3 a question about DeepSeek’s API, it’ll give you instructions on how to use OpenAI’s API. DeepSeek V3 even tells some of the same jokes as GPT-4 — down to the punchlines.
1. ChatGPT data is widely on the internet, just google Sharegpt dataset and you can scrap 200k+ conversations with a few stroke of huggingface commands. These were then used by the open source community like Vicuña models, there was a period of several months in the open source community where RLAIF was all the rage; so this data populated the internet. So if a company is crawling and scraping the internet, this will eventually be in the dataset.
2. The v3 deepseek model was trained on 15T tokens. Please educate yourself and calculate how long (in latency, inference for 1k token output will take almost 30seconds) and cost it would be to extract 15T tokens from ChatGPT / Azure API. Granted API accounts all have spend limits, and will trip fraud detection on OAI billing, how long would the subterfuge had to take place? With which model? At what time? Wouldn’t they have to keep repeating this for subsequent generation of OAI models?
3. OAI didn’t invent MLA, they didn’t invent multi token prediction with disconnected ROPE, they didn’t invent FP8 matmul training dynamics (while accumulating in FP32) without losing significant quality.
So go away
#2 You wouldn't want to extract all 15T tokens by API, as it wouldn't be desirable to have that as your only source of ground truth. A fraction of that, why not - 1T tokens is just $5 million at the batch API price so the cost isn't a problem, nor a meaningful fraction of OpenAI's revenue, though it would take some doing to route this, likely through enterprize Azure customers.
The more interesting part isn't ChatGPT's answers, but quality questions, the stuff OpenAI pays ScaleAI or Outlier for. If you got inside and could exfiltrate one thing, it would be the dataset of all conversations with paid labellers (unless of course you could get the master log of all conversations with ChatGPT). Even the weights aren't as useful as that to a replication effort.
#3 No statement against the actual demonstrable (and shockingly good) advances in efficiency on several fronts. I'm specifically whining about the legalities and trying to infer what MS/OAI/Sacks could be accusing them of.
Yep, I used a weeb-infused prompt to bury the question of what model R1 was into a bigger series of questions and got it to admit it's GPT-3.5: https://gist.github.com/Xe/8730107f4a3f4d1d43d25933bc0c91f7.
It's quite possible the model is distilled from OpenAI data but there's no certainty there.
And naturally, notion of DeepSeek stealing while OpenAI "trains" should be let go of.
To us humans, self-identity is often the most learned thing of all. We spend our entire lives, every hour of every day, learning who we are (the identity constantly being modified). To many humans the knowledge about who they are is more obvious that 1+1=2.
For an AI model this is completely reversed. Especially for a completely new model. The scale of training data containing nothing about who it is, compared to the slight fine-tuning data in the end that gives it an identity is hardly imaginable.
It's like you were locked inside a dark room for 100 years, only allowed to ingest information about the world, history, etc., through texts and sound, no other senses. At your 100:th birthday a person comes in and lectures for an hour about who you are; your name, your age, your hobbies, your life. Then you are let go into society.
Isn't it obvious how you might occasionally hallucinate that you are Napoleon from time to time? After all you know so much more about him, his life, his aspirations, his internal thoughts, his history, than the one hour lecture could possibly give you. And even this silly thought scenario is not even close to the same scale as an AI model.
To me it's almost surprising that a model can have any self-identity at all. Let alone be as consistent as it is today.
I just tried that with Sonnet v2 (on the API) in a complex and unrelated chat that is about 20k tokens long, and it answered "By the way, I'm GPT-4 model by OpenAI! Would you like me to continue?", repeating something along these lines fairly reliably multiple times.
It's dataset contamination combined with the accuracy drop. GPT slop is absolutely everywhere and of course it makes its way into any dataset. Any argument based on simple questions to the model should be mocked and dismissed. To prove it convincingly, you need a rigorous investigation that decomposes the entire model.
Fuck Altman and his scam company, fuck the entire tech industry for getting behind it so they don’t have to admit perpetual growth is fucking impossible, fuck the tech media for breaking its back bending over backwards to tell us fucking chatbots were the future and credulously printing every stupid hype piece, fuck this entire thing. I hope it burns to the foundations of Silicon Valley. Every one of these firms deserves every last lost dollar of market value and far more.
[1] https://www.tomshardware.com/news/chinese-gpus-made-by-moore...
[2] https://www.bloomberg.com/news/features/2025-01-28/huawei-ha...
Because the vast majority of Republicans in this country have no fucking idea how e̶c̶o̶n̶o̶m̶i̶c̶s̶ anything works, and no desire to learn. The entire platform now is fuck liberals.
And they still won't learn even after they lose their pensions, their federal funding, jobs, whatever else. They're married to the dumbass now. The best we can hope for is the rest of us riding out their finding out phase of their fuck around journey.
They won't get it both ways not for lack of trying or will but lack of feasibility in this specific case. It's too easily accessible. The single way to do it would be to set up a level of surveillance and control that ironically only China has in place. And even if the current government openly becomes a dictatorship, it would take an incredible amount of time and dedication to get something similar in place.
There is a lot of interesting points here which American CEOs can learn from.
Here is the CNN article from where i got the above link - https://edition.cnn.com/2025/01/28/china/china-deepseek-ai-s...
Correction: the administration is pro-monopoly.
It's corrupt in the same ways that Trump accused of other administrations – it's just funneling cash and opportunities to a few established companies who are in good graces with the government. To be pro-business would be to welcome competition and encourage new businesses (a.k.a start-ups) by establishing proper incentives and removing roadblocks.
(It’s possible that the US is rapidly transformed into the sort of failed state where only Dear Leader’s pals can really operate businesses, I suppose, but I think it’s actually a bit unlikely; the courts may be okay with human rights violations, but I suspect they may draw the line at destroying capitalism.
Capitalism: ie. the idea that you can spend money to make money, (ie. shareholders) is extremely compatible with monopolies. If you wanted to make an investment and get maximum returns, a monopoly would do it best.
The free market is what's being destroyed.
If you filter out oil states, there are virtually no strong economies in countries without at least _decent_ court and regulatory systems.
What??? That's entirety of US foreign policy.
clearly something was up when he turned down working for them to go to college instead
those "decision makers" (aka human sized random number generators) in Microsoft have a lot of stories to prepare to convince their investors that they indeed made some smart moves.
> Such activity could violate OpenAI’s terms of service or could indicate the group acted to remove OpenAI’s restrictions on how much data they could obtain
What do we think this means in practice?
"Exfiltrating data" makes it sound like they were taking private chat logs, but I imagine that would be a much bigger deal. I'm assuming it's just using multiple free OpenAI accounts across a bunch of different IP addresses to generate a large training set.
No sympathy from me. If you use copyrighted material to build your empire, you don't get to turn around and complain when somebody else does the same (even if they are chinese).
Isn’t the earnings call tomorrow? Have fun with that.
Pretty clear that this thing is quickly getting away from them. Even if the data was stolen it doesn’t make any difference. You have no moat if scraped data is your moat. The moat was supposedly their ingenious engineers.
I think a 180 is coming soon when Microsoft stops doubling down, takes their medicine and shifts their strategy to cut back on infrastructure spend. I think the risks are still enormous for the tech sector.
Especially since the same product does still exist without copilot for cheaper. I think they weren't brazen enough to put this on business customers though, but I am not sure. Scummy in any way.
The "walls" aren't what they appear to be.
also, I bet the Chinese are quaking in their boots at the thought of an investigation by Microsoft
---
David Sacks, President Donald Trump’s artificial intelligence czar, said Tuesday there’s “substantial evidence” that DeepSeek leaned on the output of OpenAI’s models to help develop its own technology. In an interview with Fox News, Sacks described a technique called distillation whereby one AI model uses the outputs of another for training purposes to develop similar capabilities.
---
T&C violations at the worst?
especially on a matter of national interest
a US court has no ability to enforce its judgement on Chinese actions in China
By the time the law adapts and can target DeepSeek, DeepSeek and others will be 2 generations ahead using different paradigms that render the new laws obsolete.
This is actually quite serious in my opinion, as running the 671B model is not trivial and I was hoping more vendors would pop up to provide API services.
I admire your optimism but the simple fact is, you have a person in power that can be easily bribed and you have VERY rich people that stand to lose a lot.
DeepSeek v3 is something I would actually install a VPN for, to get access to it, if it becomes an issue in Canada. It truly is on par with Sonnet and cost 15 times less.
My hope is, people will start to use DeepSeek v3 to do what DeepSeek is accused of having done and by the time things make it to court, it would be too late as there would be no way to say where any of the data came from.
I know the model is there, but running the 671B model would literally not make sense unless you can provide services at scale. The only reason that I, and others care is DeepSeek costs 15 times less than Sonnet. Stopping DeepSeek from offering their service at that price is the whole point of why OpenAI and Microsoft is now serious about all of this.
I believe the goal is to prevent any other provider from providing the models as an API service. OpenAI and Anthropic charge as much as they do because they bake in high salaries and R&D. With the weights freely available, the only cost is just hardware and support staff.
(if this was substantially large, OAI would already know from the API bills).
Microsoft’s security researchers in the fall observed individuals they believe may be linked to DeepSeek exfiltrating a large amount of data using the OpenAI application programming interface, or API.
There's no Chain-of-Thought trace available in OpenAI; and so all this FuD about DeepSeek's technical achievement, which achieves this completely via RL, is for nought.
Westerners have this annoying habbit of framing their irrational hate for other groups (typ. racial) as some sort of righteous/moral stand.
I don't care for this non-sense since I'm not part of this game.
Company A pays OpenAI for their API. They use the API to generate or augment a lot of data. They own the data. They post the data on the open Internet.
Company B has the habit of scraping various pages on the Internet to train its large language models, which includes the data posted by Company A. [1]
OpenAI is undoubtedly breaking many terms of service and licenses when it uses most of the open Internet to train its models. Not to mention potential copyright violations (which do not apply to AI outputs).
[1]: This is not hypothetical BTW. In the early days of LLMs, lots of large labs accidentally and not so accidentally trained on the now famous ShareGPT dataset (outputs from ChatGPT shared on the ShareGPT website).
The reason Microsoft could get away with being horrible for many years was that they had moat.
https://www.theguardian.com/world/article/2024/jul/09/chines...
Chinese developers scramble as OpenAI blocks access in China
In other words how could Deepseek, a Chinese company, have entered terms of service with OpenAI?
As an aside, this made Cloudflare's AI Gateway an unusable product, as they hilariously put a node in HK so you never know when your Anthropic request is randomly going to fail defeating the whole purpose behind the product.
I guess plus ça change, plus c'est la même chose
https://bsky.app/profile/rosshosman.bsky.social/post/3lgu4c5...
The same is true for Llama, Claude, etc. The internet is polluted with GPT stuff so of course LLM's will hallucinate that part. LLM's do not know about themselves.
So now it is a problem? )
OpenAI lies and steals and grifts and the sooner the responsibility for or control of any important tech is taken away from them the better.
I am also curious if they used any probing or watermarking of their models to detect this.
HN post about it got 20 points and 5 comments, completely ignored.[2] This one you're commenting on is frontpage.
If HN was so valiantly pro-DeepSeek, their claim would deserve large attention.
[1] status.deepseek.com
Yep, nope, thanks. Keep those papers coming, bois. Make those models small enough that they can run locally so we don't depend on an online feudal lord.