Is Grok Basically Just an OpenAI Wrapper?
twitter.com
twitter.com
> The issue here is that the web is full of ChatGPT outputs, so we accidentally picked up some of them when we trained Grok on a large amount of web data. This was a huge surprise to us when we first noticed it. For what it’s worth, the issue is very rare and now that we’re aware of it we’ll make sure that future versions of Grok don’t have this problem. Don’t worry, no OpenAI code was used to make Grok.
Edit: I'm not 100% sure, but I think my comment is being interpreted uncharitably. I would like to grab your finger, which is pressed squarely atop "output of other LLMs", and move it over a bit to the label of "website with unknown and almost certainly copyright content". And then, I would like to point out that there is a difference between training on such data with the benefit of plausible deniability (e.g. some third party), versus going specifically after data you know to be copyrighted, and then claiming that this is precisely the reason why your model will spit such content out verbatim. At that point, you might as well indict yourself, no?
(I'm hoping it's possible to do that, but I've not seen an example yet.)
(BERT was useful but not GPT-3 useful.)
https://www.reuters.com/legal/ai-generated-art-cannot-receiv...
>not a lawyer, go ask one if you need a legal opinion
Copyright is about distribution, not about use.
At least until some new laws passed.
It would be copyright violation if they pirated the material. But assuming they bought or scraped, they would be good.
The core issue is that no one knows how copyright law will impact training an LLM until legal precedent is set, which is exactly what's being argued in civil suits now.
Grok is extremely consistent in both referring to OpenAI in refusals, and also using the full repertoire of GPT-ism in its other responses, such as ending every message with an "Overall, ..." or "In conclusion, ..." section.
They RLHF-ed Grok on GPT output, specifically and in volume. Maybe sourced via the API or other means, I don't know.
But Igor's explanation that "this is very rare" makes no sense after you play with Grok for even 5 minutes.
This is super common in the open-source/local AI area. Many models are trained on output from GPT. The better models will filter out anything that mentions GPT or OpenAI. Seems like Grok is not one of the better models.
OpenAI would know that they were doing this. Reselling like this is against their terms of service: https://openai.com/policies/terms-of-use
*about anything
Basically never.
The raw GPT model would be anything. My favorite was "I'm a young female lawyer with ambitions." But a tuned model knows what it is.
Grok keeps confusing itself with ChatGPT because it was trained via ChatGPT, directly and explicitly. And apparently the team sucks at its job, filtering the data.
Also, despite it's clearly modeled after GPT, it also often does things GPT wouldn't say. In particular, it can get very repetitive, repeating itself over and over when asked to clarify a question, which curiously is also what Bing's model does.
Bing is powered by multiple models, an early version of GPT-4 currently to be replaced with GPT-4 Turbo, and another model Microsoft made... which is the repetitive one.
It seems part of OpenAI's secret sauce is to avoid the model getting stuck in loops, which naive efforts like xAI's and Microsoft have not independently resolved.
OpenAI terms of service explicitly say it cannot be used to create a competing product.
Is it hard to filter ChatGPT output from random chats shared in articles? Yes, it is. That's their problem though.
Probably.
> Isn't training on content without permission kind of their whole business model?
Right, their business model rests on the argument that copyright law doesn't require permission from copyright holders. They aren't relying on copyright law, but contract law (which only binds direct users, but applies to material not covered by copyright and applies in circumstances where even if material was covered by copyright use would be allowed as, e.g, fair use.)
Is it an attempt to pull up the ladder? Obviously. But both the safety mission of the charity and the commercial competitive interests of the for profit provide rationales for this, so there is no reason to think that just because it hobbles others in trying to catch up that it would be difficult for them to do it with a straight face.
There is a meaningful distinction between using ChatGPT directly, putting in the prompts and taking the results, and just finding random output that other people have published in the internet and using that.
apply only to people who agree to OpenAI's terms of service (its users)
If Musk's "anti-woke" AI was just trained on logs from what he calls "woke AI", that would, actually, be particularly newsworthy.
(The explanation they have put forward is that this is not the case, but they just failed to realize that blindly ingesting web content would bring in a lot of ChatGPT-derived content.)
It's currently in preview for some X premium subscribers.
I can't think of any reason anybody would use Grok over ChatGPT besides political tribe signalling.
> A unique and fundamental advantage of Grok is that it has real-time knowledge of the world via the 𝕏 platform.
You disagree with that?
It's not a fundamental advantage for most AI model uses or a particularly good source of knowledge of the world.
The best use cases for Grok are going to be in advertising. Corporations that pay X.com a lot of money will get privileged access to Grok features for sentiment analysis and ad placement. Basically, the AI is going to optimize advertising revenue because real-time access to data will allow advertisers to optimize their targeting as quickly as Grok can manage to summarize the relevant metrics for impact of the current ad campaign. Advertising agencies are not dumb so I suspect they will start pivoting to using generative assets for their A/B testing and optimization which means that something like Grok with the proper generative plug-ins will let them do a lot more with a smaller team of marketing "creatives".
You could ask "which turd sandwiches don't taste like shit?" when really, the question isn't "where are the better turd sandwiches?" but rather "why the hell are we eating turd sandwiches?"
Does that mean you are a wrapper of a book that you read and learned from? Am I a Kernighan and Ritchie when I wrote C because I read their book 20 years ago?
I think wrapper has a specific meaning that is not applicable. Grok isn't abstracting or an interface to ChatGPT.
ChatGPT is fine at writing code unless you tell it it's for hacking, and then it won't cooperate. Having to trick the bot into writing code that it doesn't want to write is a weird place to be in, but that's where we are.
* https://chat.openai.com/share/713e6069-cf31-4585-a5bc-fbf0b1...
This is common issue across all sorts of other alternative models too. It's not particularly surprising.
Every example of grok transcripts I've seen so far very much has the feel of using the gpt4 api with custom prompts (ie, not chatgpt). It wouldn't surprise me in the slightest if they're faking it.
Training models isn't actually that hard these days, provided you have the capital to do it. There's a LOT of good literature now, and an increasing pool of people with experience of the process.
Handling the huge number of cases where ChatGPT says something like "As a large language model created by OpenAI" would be very simple.