OpenAI suspends ByteDance's account after it used GPT to train its own AI model
theverge.com
theverge.com
I'm genuinely wondering how this is different from them using others work without consent, but IANAL so maybe I'm just confusing mortality and legality
Do we know what data was used? And what the constraints were around it?
Do we know it was used without permission or are we just jumping in the “ai bad” bandwagon?
Conversely, photocopiers and text-to-speech engines and LLMs don't exercise choice over whether they reproduce copyrighted material and so can't be held responsible, so responsibility for clearing rights to redistribute/transform in that format clearly lies with the people inputting the copyrighted material. Obviously, OpenAI has tended to avoid making any attempts to secure those rights whatsoever
The ones that do secure permission where the works involved are subject to copyright.
The fact that the question/accusation has been raised a great many times and they have not stated "we know we haven't used information without licence because we had procedures to check licensing for all the data used to train our models", would certainly imply that they scraped data for training without reference to licensing, which makes it very likely that the models are based significantly on copyrighted and copyleft covered information.
> Do we know it was used without permission
No, we don't know for sure. But the balance of probabilities is massively skewed in that direction.
There are enough examples of image producing AIs regurgitating obvious parts of unlicensed inputs, as an indication of the common practise of just scraping everything without a care for permission. So asking for those with other models to state how they checked for permission for the input data is reasonable.
For years we’ve accepted that search engines, for example, can grab all the code on GitHub and use it to build a search index.
Google image search, in particular, ‘regurgitates’ all the images it has indexed when it thinks they match a search term.
It has a little disclaimer it shows next to the results saying that ‘images may be copyrighted’ - figuring out if they are copyrighted and if so by whom is left as an exercise for you the searcher. Depending on what you are using the image search for, the copyright of the images may, after all, not be relevant. Like, if you’re using a Google image search to get inspiration for home decor designs, do you care who owns the copyright of each image? Should Google?
GPT poses similar risks to that. It has the explicit disclaimer that things it produces might be subject to copyright. Depending on what you’re using the output for, the copyright may or may not be relevant.
Google actually pays license fees to News Corp to excerpt their content in Google News following a legal challenge so it's not exactly conclusively established that search engines have global rights to do what they do anyway. But search engines are mostly beneficial to contact creators rather than mostly competitive with them.
This is irrelevant to original point about the purpose of a search engine being to highlight rather than replace existing information sources, and OpenAI's purpose for indexing content and policy of not engaging with copyright holders being completely different
why is that? all works, as long as it is "original" should be granted copyright. It should belong to the person who made it (in this case, the user, not openAI).
thats the contention, its not a person who's made it.
If I say to one of my human friends "write me a nursery rhyme", the copyright of the resulting rhyme would obviously belong to my friend - despite me prompting them. Clearly the prompt itself does not universally count as "making" it.
Let's say I made a "NovelSnippetAI", which contains a corpus of prewritten material. You can prompt it, and it will return a page which matches the sentiment of your prompt best. I think we can agree that the copyright of the page will still belong to the original writer - the user only did a query.
What if I did "NovelMixAI", which did exactly the same but alternated lines from the two best matches? What about "NovelTransformAI", which applied a mathematical formula to the best match and fed the output to a fixed Markov Chain? Now we're suddenly at "LLMAI", which does the same using a neural network - what makes it different from the rest?
Long-standing precedent is that any automated work does not qualify for copyright. You can only copyright human work.
This is the wrong analogy. The LLM is more like Photoshop and the prompt little more than a filter configuration. A machine cannot copyright its own output but a human guiding that machine can.
If they just wanted money, they could have cut out a lot of the ethical filters and choke out competitors. Google didn't stop NSFW. Twitter, Reddit, Tumblr, etc didn't. AI is bound to be used for NSFW among other things, but they've set the standards to make it ethical.
I think eventually they did let loose to try to keep ahead of competitors. This probably pissed off the board and led to the drama recently? Just speculation. Because the new models are nowhere near as anal as the initial release.
I think you are wrong here, the safety filters (“safety” and “ethics” for AI are labels for boundaries and concerns of different ideological factions, and Altman and OpenAI are deep in the “safety” side—which has the nost money and power behind it so “safety” is also becoming the generic term for AI boundaries) are an important oart if the PR and regulatory lobbying effort, which is key to OpenAI’s long term moneymaking plan, given the absence of a durable moat for commercial AI.
> If they just wanted money, they could have cut out a lot of the ethical filters and choke out competitors. Google didn't stop NSFW. Twitter, Reddit, Tumblr, etc didn't.
Neither, in practice, has OpenAI—there are whole communitids built around using OpenAI's models for very, very NSFW purposes. They've prevented casual NSFW, to oresent the image they want to thr government and interest grouoa whose support they want as they lovby to shape regulation of AI. avoiding being a target for things like the 404 Media campaign against CivitAI where the NSFW is more readily visible.
IP rights do not per se give the author an absolute right to determine how their work is used. It does give rights to prevent a reproduction, but that is not what an AI model does.
BTW I noticed that GPT-4 is good at writing legal letters of the sort that is widely available online. But a subpoena, ('dagvaarding', the Dutch version I have researched) it completely fails to create. Also there are not many subpoenas available online, and the court (in the Netherlands) only publishes the verdicts, not the other documents. Lawyers OTOH have a lot more of this available in their libraries.
So, my impression is that there is still a lot of material out there that is not in the corpus.
"Hello, we are a non-profit that wants to make AI models to benefit humanity. Can you give us access to your data to help us in our work?"
What if they just bought dirt cheap used copies meaning the creators saw nothing thanks to the first sale doctrine?
They could without question set up a very nice physical library in Mountain View and even invite the public in. They can probably in general scan those books for their own internal use. What got shut down was scanning the books and making them available to everyone in their entirety.
Why would you conclude that ? While an AI model does not ONLY reproduce, it most certainly can make verbatim reproductions. The things preventing the user from getting copyrighted material from chatgpt are probably only rules/guardrails. The most prominent example of this is perhaps the Bible which you could get from it quote by quote within token limit.
That’s not the same thing as general copyright.
We’re not trying to curb emissions for the purpose of kneecapping other economies. Short of China, we don’t really have any incentive to do that (bigger markets = better under US industrial policy [contradictory opinions from left wing undergrad students on Twitter don’t count as “US industrial policy”]). What’s actually happening is that we caused a problem and now it’s getting worse and in order to fix it we need to not allow everyone else to continue it.
This is a novel and advanced philosophical argument called, “two wrongs don’t make a right.”
There's another example - intellectual property. The US was fine playing fast and loose with IP (Most famous example is Dickens' attempts to point out he was being pirated left and right in the States and not seeing a penny: https://www.charlesdickenspage.com/copyright1842.html)
Now America is on top of the IP pile, it sees other nations as playing fast and loose: https://www.forbes.com/sites/johnlamattina/2013/04/08/indias...
Countries don’t have to join international trade regimes. They also don’t have to join climate/emission commitments. They do both of them because they come with benefits.
Cursory searching suggests the first real work on international copyright, by contrast, came about in 1886. Even early versions came after Dickens’ story here.
If you have a different balance to strike that you think is significantly and obviously better, I’m sure the whole world is interested in hearing it.
> We share the same planet, so we need to share the same environmental standards
All of the excuses to hold China to a different standard are horse shit. Is that stated plainly enough for you?
Will anybody meet their goals? The planet is the ultimate Commons. Personally, I think we're boned.
History doesn't repeat, but it does rhyme. Look at the shape of the story. The people on top support rules that completely coincidentally help keep them on top. It's a universal impluse.
"It is difficult to get a man to understand something when his salary depends on his not understanding it." (Upton Sinclair) is another example with a similar shape, but at the scale of individuals.
John Rawls had the right idea.
One wrong gets punished and the other - "we are exceptional" so we do as we please and do not fuck with us or else.
The morality of training on internet-scale text data is another discussion, but I would point out that this has been standard practice since the advent of the internet, both for training smaller models and for fueling large tech companies such as Google. Broadly speaking, there is nothing wrong with mere consumption. What gets both morally and legally more complex is production - how much are you allowed to synthesize from the training data? And that is a fair question.
In fact, the power of linking to data sources is what Google is almost entirely built upon.
And millions of documents authored by people that weren't compensated.
The difference is consolidating all of that value into a single company.
Similarly, people are avoiding the cost of pre-training GPT-4 class model by scraping its output.
So I think it's fair to question the moral consistency of their ToS.
[1] Please note that I am not passing a judgement on this, just stating a fact in order to make an argument.
All you’re doing is redefining content, ie thoughts, ideas, movies, videos, literature, sounds, writing, etc as “raw data”. But that isn’t raw data. There was a ton of effort that went into creating the “content”. For example, a single Wikipedia page may have many hundreds of people, some who have done years of college level studies and original research, to produce a few thousand words of content. Others have done research using primary sources. All of them have had to use effort and ingenuity to craft those into actual high quality statements, which in itself was only possible in many cases due to years of training and education. Finally, they had to setup a validation process to produce useful output from this collaborative process which included loads of arguments etc to generate what you are calling “raw data”.
I’m not sure what makes GPT’s output is any less raw than all the effort that went into producing a single Wikipedia page? Further, Wikipedia actually goes out of its way to cite its sources. GPT is designed to go out of its way to obscure its sources.
The only thing GPT does, IOW, that apparently makes the data it uses is not to cite its sources, something that would at the very least lead to professional disgrace for the people who created the “raw data” GPT uses without thought, and would even lead to lawsuits and prosecution in many cases.
So besides going out of its way to obscure the source of its data, what makes GPT’s output less raw than the output people have spent billions of man hours creating?
If GPT incurred a non negligible cost on the content owners by accessing their resources it may have been different but that's not the case.
The only thing that content owners may be able to complain about is that potentially ChatGPT/DallE may reduce their potential income and this would have to be proven. I have not stopped buying books or art of any kind since I use ChatGPT/DallE. And low quality automated content producers existed before OpenAI and were already diluting the attention to more carefully produced content (as can be seen with videos on youtube).
Quite often it contains the experience of a life of a person condensed to a few hundred pages.
ChatGPT gives easier access to the knowledge contained in tens of thousands of these books. As for me I have been reading less and less books as more wisdom is accessable on the internet in better forms (now GPT).
I'm not against what OpenAI is doing as it moves humanity forward, but like you said I won't stop using ChatGPT just because ByteDance scrapes it.
How many books do you store in 1GB How much does it cost a year to store it and have OpenAI gather it once. How much does it cost to run a GPT4 level model that will output 1GB.
That's my point here that's all. It is a huge cost for OpenAI to run a system that produces dynamic content. And it is not comparable to the cost of storing static content.
I didn't talk about the cost of producing the original data.
And I do not talk about training costs.
(while at the same time reducing future costs for everyone by distributing the capability more widely to prevent monopolization)
I’m just curious how you start off with:
> Morality and legality aside
Only to then follow it up immediately with an argument for why one is more moral. Just because you didn’t end it with “and that’s why I think OpenAI is more moral,” doesn’t mean it’s not obvious and less of an irony.
Maybe it shouldn't have been? We've been frog-boiling toward this point for a long time, from a starting point that was generally good for content creators (your content is made more discoverable) to a point that is not so good for content creators (your content is scraped and digested, programmatically laundered and regurgitated on huge corporations' own platforms with token or no attribution, and no revenue shared).
In a parallel universe where search engines were explicitly opt-in from the beginning, I think these conversations would look very different today. What OpenAI and its peers have done would, I dare say, be uncontroversially (and correctly) regarded as theft. Just as I'm not allowed to distribute^[1] software incorporating somebody else's code in a way that violates the terms of its license (or lack thereof), I shouldn't be able to distribute software that incorporates any intellectual property that I don't have the rights to.
^[1] Broadly speaking.
When it comes to cutting edge business vs business decisions, the legality is often defined post-factum, e.g. in courts.
For an outsider to know whether something in a case like this is legal or not is near impossible, considering how opaque such businesses are.
Courts generally uphold provisions against reverse-engineering (as protecting internal, proprietary knowledge) but are more welcoming to copying interfaces (as encouraging market substitutes). So one question would be whether OpenAI can restrict use of the output of their tool in this manner, since the output itself is manifestly open (to the customer). That seems novel. The only analogy I know of is database licensing that prevents customers from publishing comparisons, which seems anti-competitive.
Anti-trust policy is motivated mainly in mature markets, where one player has fairly (by hypothesis) grown to dominate. The law and courts apply special scrutiny to identify ordinarily-acceptable market practices that extend the market power of the dominant player.
But is it the same analysis in a growing market? It seems like even if OpenAI (or especially, OpenAI+Microsoft) is dominant, if the market is growing quickly, the concern might be relaxed since the dominance is uncertain. Conversely, if the market is particularly susceptible to capture, early leaders might warrant heightened scrutiny.
But aside from monopoly's first-order effect on reducing competition, the second-order effect is to reduce investments in competitors, which has anti-competitive effects. That concern could be highest in the early stages of a market.
Are there any good lawyer blogs on point?
Competitors can't get OpenAI's model weights but use its outputs to produce a functionally similar model.
It's like if you had a competitor's engine and couldn't open it up but could still see the outputs: torque, rpm, ..., and could control the inputs: fuel intake, air mixture, etc... Then you make an engine by inferring back from these measurements.
Wouldn't that be reverse engineering?
Your analogy is like going to an airport and looking at departure and arrival times as well as the flight path and then “reverse engineering” an aircraft from that. The chances are very low that you produce anything remotely resembling the aircraft you’re trying to reverse engineer. Same goes for the engine.
In the case at hand we are dealing with a mathematical function. In->Out is all there is. Back-estimating a mathematical function by sampling is as close to reverse engineering as anything is.
In addition, this is much more akin to data exfiltration than to copying an engineered mechanism. The training algorithm could count as an engineered mechanism, but that’s not what is being copied.
Incredible. The most broken version of copyright imaginable.
Doesn’t matter how much endless amounts of money they spend, they’re going to have to contend with the fact that the value they ship is derived from other’s work. It’s just diluted to the point of it becoming “data” rather than “artworks”.
And let's say I do not want them to clean up and then use my data for profit.
Sorry, what? You think reddit is trying to prevent openai from scraping the porn subreddits???
Not sure if a dribbble user can suspend OpenAI.
Altman's return is really working out
>As part of the deal, ChatGPT users will receive summaries of news stories from Axel Springer’s brands, including Politico, Business Insider, Bild and Welt, with attribution and links to the original sources of reporting, the companies said Wednesday. The agreement will allow OpenAI’s models to take advantage of the publisher’s higher quality and more current information in its chatbots’ answers.
https://www.cnn.com/2023/12/13/tech/open-ai-axel-springer-ch...
Lol is this a joke? Those are all tabloids.
Let me quote one of Germany's best cabaret artists, Volker Piepers: "Bild-Zeitung..... this filthy newspaper that is so disgusting that you insult dead fish if you wrap it in it!"
There is a reason Heinrich Böll wrote a book about them:
https://en.m.wikipedia.org/wiki/The_Lost_Honour_of_Katharina...
If you consider him "one of Germany's best cabaret artist" you could not have better proven how right I am.
The newspaper of AS regularly publish articles to push their agendas and remove politicians they don't like.
Just look what happened to Federal President of Germany, Christian Wullf, when he declared that Islam is part of Germany.
What happened to him? By reading the Wikipedia article, he seems like a very corrupt individual in a position of power. I don't see how Bild was involved in that.
Perhaps new subscription plans for ChatGPT will be payable in Worldcoin too.
https://www.wired.com/story/openai-buy-ai-chips-startup-sam-...
If you ever wondered how all these models could seem to catch up to GPT-3.5 so quickly, but then struggled to do noticeably better (much less exceed GPT-4), while not talking about their data or saying they definitely weren't simply training on GPT-3.5 outputs, remember this: they might just be lying.
I don't like ByteDance at all, but I hope OpenAI are aware both shady and legit companies are making some serious cash using their API.
> All API customers must adhere to our usage policies to ensure that our technology is used for good.
"For good" sounds like the philosophy of companies like Stripe and Twitch. There always is a grey area with stuff that is good for many people, but other people see as evil.
Then OpenAI came along and it seemed like it wouldn't be necessary any more because they were releasing their models anyway, except for when they suddenly decided that they wouldn't any more.
But others have been carrying that torch since (Llama and many others). Still it seems an interesting way of enhancing a model.
I guess OpenAI did the same. If you read their API terms one of the first things they prohibit is to use it to train a competing model. Maybe they do the same tricks internally and know how powerful it can be?
However, you might wonder what the goal is. This “API distillation” is good for teaching a pretrained model how to do lots of things. But the end result is always constrained by the quality of the pretrained base model, and API outputs don’t help at all with that part.
If you tell someone that they can’t use your product in a certain manner, and then they go to extra lengths to circumvent your measures to gain a profit for themselves, then there is going to be big civil and criminal legal problems.
State sponsored or not; Bytedance has a bank account and business license here in the U.S.
Garbage in = Garbage Out
That's entirely non-intuitive to me how that would work. Like are they just asking it questions about every topic under the sun, and then creating training materials out of the answers? How would you even begin to assemble a list of prompts that would cover all of knowledge? And how could you ever distinguish useful outputs from nonsense hallucinations?
I feel like I'm missing some key details here that the article and its links don't explain at all.
They don't need to, all that knowledge is already in public training sets from scraping the internet.
What is harder to get is the answer patterns to ensure good user experience. You want a lot of such answer patterns so the model knows what structure to use for different kinds of questions. That structure isn't to make the result formatted for humans, but contains reasoning paths the LLM takes in order to arrive at reasonable answers. Since an LLMs thinking is the words it writes the structure of how it responds thus corresponds to thinking patterns, and you want the LLM to learn a lot of those, internet data wont contain that but ChatGPT responses will.
TLDR: The words an LLM writes is also what it thinks, since it doesn't hide thoughts, so by training on what ChatGPT writes you also train on how it thinks. It isn't the facts you want but the thinking patterns.
However by asking ChatGPT questions and recording their answers, you can use that to fine-tune the model they are creating. They can tell their own model to answer the way ChatGPT answered.
In short they are copying how ChatGPT would answer. and not finding out for themselves by creating their own Question and answer (or completion) datasets.
That is what RLHF is. But anyway, Chinese companies will be able to compete here, as this will be commoditized quickly and China is the king of creating and selling commodities.
It's like if Monsanto was called Organic Cooperative.
https://addons.mozilla.org/en-US/firefox/addon/openai-is-not...
That's truly hilarious... It's akin to a child who stole a candy bar in a candy store complaining about their sister stealing that same candy bar from them.
A spider will generally have a pretty predictable route through a web site.
This API restriction is only an issue until that point, I assume, or is the distilling being done in these situations much more dynamic in nature?
Just the other day (PDF warning): https://selectcommitteeontheccp.house.gov/sites/evo-subsites...
> Pillar II: Stem the Flow of U.S. Capital and Technology Fueling the PRC’s Military Modernization and Human Rights Abuses
> U.S. export controls have been slow to adapt to rapid changes in technology and attempts by adversaries to blur the lines between private and public sector entities, particularly the PRC’s strategy of Military-Civil Fusion.
> “The military is for civilian use, the civilian is military, and the military and civilian are fused.” In other words, no line exists between civilian and military technological development. U.S. export controls have yet to adapt to this reality
If they don't allow this it just seems like they are trying to prevent people from building smaller cheaper models that will perform better for a specific use case and gobbling up the market for as long as they can.
I'm into freedom of information and what not but come on!!!
So Bytedance gets account revoked but tons of other models are allowed.
China bans US companies, US does same to China.
We are both almost the same. Sure they have a more authoritarian government but ours isn’t quite a liberal democracy either.
It’s competition all the way down.
Not violating terms. Not abiding by terms of use. But good vs evil.
It screeches “we occupy the moral high ground and are the arbiters of what is ethical and what is evil.”
The entire brand game plan around “Open”AI is to position themselves as first doing no evil, just like google did in the early days, to enable doing a lot of evil.
OpenAI knows this which is why they put that TOS there, they spent a lot of money to create that finetuning data so they don't want to give all of that away for cheap.
Both can be used to create training data quite successfully. This technique has been used in the past to create synthetic post-training (fine tuning) datasets like Orca, Samantha, and so on.
What they gonna do, double ban you.
It doesn't have to be internally consistent, it just has to make them money.
Their ladder (using public data, and hiring humans to classify to taste) is still available I believe.
Not really. Once chatGPT came out, many sites changed their terms and/or significantly increased their API access costs to prevent/limit/make cost prohibitive future scraping.
I had assumed most of their web content was from Common Crawl, and the older pre-ChatGPT Common Crawl datasets used would still be available. But it looks like Twitter, for one, was not in Common Crawl.
Which is not openAI “pulling up the ladder behind them”
There is literally no way for them to avoid looking like assholes once they take enact barriers that they themselves did not have to overcome.
As far as we know.
Not the case. That would be the case if OpenAI prevented them from using the same resources, which ChatGPT is not.
To keep that analogy going, they essentially used the ladder so hard, it broke.
It kind of is. Just think of how many services changed their data sharing policies and closed APIs due to ChatGPT training (Twitter, Stack Overflow, Reddit). Maybe the analogy is that instead of pulling the ladder, they set fire to it so it’s burning and making it harder for others to climb. Even if they didn’t set it alight on purpose, I don’t imagine they’re losing sleep over it.
Google itself on the other hand does just this, for example Wikipedia text and lyrics are taken from other pages and copy pasted onto Google's page.
Try telling a "googler" this and theyll go "noooo but for Google it's different because Google has determined that's optimal and good for user experience". It's difficult to get someone to understand something when their paycheck depends on them not understanding it.
https://www.theverge.com/2022/6/22/23178245/google-paying-wi...
https://www.theverge.com/2019/6/18/18684211/google-song-lyri...
What I said is if I copy paste things into my page, Google will kill it because it's spam. If I say "but what I'm actually paying for this content"... That's irrelevant.
I'm not saying that Google is stealing content, I'm saying they're hypocritical is applying and argument when the conclusions benefit them, but not otherwise.
In my experience googlers are very capable of saying “Google is bad but my salary is good”. Plenty of people understand things that are contrary to their paycheck.
That said, Google generally respects Wikipedia’s and others license to the data. And it is generally in the users best interest to get to desired content/information in less steps, regardless of the data’s provenance.
Not equalizing those two situations at all, just pointing out that dynamics of communication between normal people don't really happen in many other situations, or if they do its just a shallow charade. Or... just don't expect fairness and good behavior when tons of money, power and legacies are at stake.
tl;dr Google crawls you, you can't crawl Google. How is that fair? They built their empire on our brain outputs but won't share theirs.
So while OpenAI is absolutely lacks the moral high-ground, ByteDance still seems to be engaging in adversarial behavior.
It was a direct order from Paul Graham. He keeps mum about it but I have trusted sources who know the truth. Additionally, it's sort of public knowledge:
https://www.washingtonpost.com/technology/2023/11/22/sam-alt...
I don't have a full view into exactly why he was fired from both OpenAI and Y-combinator. But from what I hear the reasoning is a bit similar. Sam Altman is a bit of a political snake. He lacks ethics and he's not honest either. The last part is just me speculating on a lot of the anecdotes from quips I've heard over the years from people who know Sam.
Sams public persona is very different. And I think a lot of HN viewers worship that public persona. But Sam being a darling of the HN and Ycombinator themselves? No way. They fired him so it's unlikely.
That when you come to this site you have to type in the words ycombinator for the url right?
You said Sam is HNs darling. That's different from saying HN viewers. So you literally did not say what you said here.
Nope, this is wrong. HN is a site run by Y Combinator and heavily moderated but the comments are coming from the site's users and not Y Combinator itself
You most obviously must be referring to the owners of the website.
Or are you referring to the comments?
Because both the users and the owners of said website have vastly different opinions here.
Additionally at one point Sam Altman was literally a "darling" for the company, before he got fired.
Your statement is ambigiuose and therefore wrong.
Also breaking OpenAI's TOS is likely completely legal and everyone I know is collecting data to their own model. Worst they could do is ban the account.
OpenAI haven’t claimed this. They are refusing to generate new data for ByteDance by suspending their account.
- It’s legal and moral to train on data you have access to, regardless of copyright.
- Nobody is obligated to provide services to you so you can obtain that data from them.
It would be hypocritical if, say, ByteDance obtained synthetic data generated from GPT-4 and then OpenAI tried to prevent them from training on the data they already obtained. But all they are doing at the moment is temporarily pausing generating new data for them. OpenAI aren’t obligated to do this and OpenAI have never argued that other people are obligated to do it for them. So no hypocrisy.
It’s not a straightforward copyright issue but in many jurisdictions that is not allowed. Company A did the work, they should be allowed to profit.
There is protection in a few notable jurisdictions so a violation would make the product illegal in those jurisdictions, which is a problem if it’s an online product.
OpenAI are free to block anyone from using their API if they want. Just like anyone hosting their content a website is free to block the OpenAI web crawler.
Aka... some dev fired off a handful of test queries and never used the account again, so OpenAI decided to suspend it to look good in the US press.
Also, what does “minimal” even mean? I’m sure they monitor accounts that max out their API request limits, or even just request programmatically (i.e. request patterns that don’t match a natural human use pattern, like slowing down during a time zones lunch hours). Maybe this was a couple days worth of traffic
======
What You Cannot Do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not:
Use our Services in a way that infringes, misappropriates or violates anyone’s rights.
Modify, copy, lease, sell or distribute any of our Services.
Attempt to or assist anyone to reverse engineer, decompile or discover the source code or underlying components of our Services, including our models, algorithms, or systems (except to the extent this restriction is prohibited by applicable law).
Automatically or programmatically extract data or Output (defined below).
Represent that Output was human-generated when it was not.
Interfere with or disrupt our Services, including circumvent any rate limits or restrictions or bypass any protective measures or safety mitigations we put on our Services.
Use Output to develop models that compete with OpenAI.
======So throwing around statements like "we suspended ByteDance to ensure GPT is used for good" are hypocritical at best. They're not the pope, they have no monopoly on good.
What does that mean? Isn’t that just the definition of API use?