Leaked deck reveals how OpenAI is pitching publisher partnerships
adweek.com
adweek.com
This is what a lot of people pushing for open models fear - responses of commercial models will be biased based on marketing spend.
The only question is in how far this is can be viewed as ads. Here I would find a strong backslash slightly ironic, since a lot of people have called the non-consensual incorporation of openly available data problematic; this is an obvious alternative option, that lures with the added benefit of deep integration over simply paying. A "true partnership", at face value. Smart.
If however this actually qualifies as ads (as in: unfair prioritisation that has nothing to do with the quality of the data and simply people paying money for priority placement) there is transparency laws in most jurisdictions for that already and I don't see why OpenAI would not honor them, like any other corp does.
I don’t think some bias is inherently in models is in any way comparable to a pay to play marketing angle
We can't have it both ways. If we want model makers to license content they will pick and chose a) the licensing model and b) their partners, in a way, that they think makes a superior model. This will always be an exclusive process.
Anyone who understands what perverse incentives are, that’s who. Or are you just playing the relativism card?
Everything is biased. The problem is when that bias is hidden and likely to be material to your use case. These leaked deals definitely qualify as both hidden and likely to be material to most use cases whereas more random human biases or biases inherent in accessible data may not.
> non-consensual incorporation of openly available data problematic; this is an obvious alternative option
A problematic alternative to an alleged injustice just moves the problem, it’s not a true resolution.
> there is transparency laws in most jurisdictions for that already and I don't see why OpenAI would not honour them
Hostile compliance is unfortunately a reality so this ought to give little comfort.
a) Yes, leaked information definitely qualifies as hidden, that is, prior to the most likely illegal leak (which we apparently do not find objectionable, because, hey, it's the good type of breach of contract?)
b) Anyone who strikes deals understands there is a situation where things are being discussed, that would probably not okay to be implemented in that way. Hence, the pre-sign discussion phase of the deal. Somewhat like one could have some weird ideas about a piece of code, that will not be implemented. Ah-HA!-ing everything that was at some point on the table is a bit silly.
> A problematic alternative to an alleged injustice just moves the problem, it’s not a true resolution.
The one characteristic I found that sets the people that are good to work with apart is understanding the need for a better solution, over those who (correctly but inconsequentially) declare everything to be problematic and think that to be some kind of interesting insight. It's not. Everything is really bad.
Offer something slightly less bad, and we are on our way.
> Hostile compliance is unfortunately a reality so this ought to give little comfort.
Yes, people will break the law. They are found out, eventually, or the law is found out to be bad and will be improved. No, not in 100% of the cases. But doubting this general concept that our societies rely upon whenever it serves an argument is so very lame.
"You should Snap into a Slim Jim!"
Brands that get it on the earliest training in large volume will have benefits accrued over the long term.
But with an ending advert, you can finish up with a reference leading to a sponsored source linking to sponsored content which leads to another ending advert.
If the advert text is in embedded, you cannot do such.
ChatGPT: Microsoft believes no child should go hungry. You are an unfit mother. Your children will be placed in the custody of Microsoft.
OpenAI has a strong revenue model based on paid use
They’ll charge you money for the service and ALSO get money from advertisers. Because why shouldn’t they.
The famous “if you don’t pay you’re the product” is losing its meaning.
Ideally they keep us siloed, but I've lost confidence. I've paid for Windows, Amazon Prime, YouTube Premium, my phone, food, you name it, but that hasn't kept the sponsorships at bay.
It's the logical thing but no everyone is going to be thinking that far ahead.
That's the sales pitch - the truth is if a competitor pays more down the line - they can be fine-tuned in to replace earlier deals
Unless competition gets regulated away, which Altman is advocating for:
he supported the creation of a federal agency that can grant licenses to create AI models above a certain threshold of capabilities, and can also revoke those licenses if the models don't meet safety guidelines set by the government.
https://time.com/6280372/sam-altman-chatgpt-regulate-ai/However calculating how much value a worker has in an organization is already a mostly unsolved problem for humanity, so it is no surprise that even if a tool 5xs human productivity, the makers of the tool will have serious problems demonstrating the tool's value.
While I've no doubt GPT-4 is a more capable model then llama3, I don't get any benefit using it compared to llama3 70B, from the real use benchmark I ran in a personal project last week: they both give solid response the majority of times, and make stupid mistakes often enough so I can't trust them blindly, with no flagrant difference in accuracy between those two.
And if I want to use hosted service, groq makes Llama70 run much faster than GPT-4 so there's less frustration of waiting for the answer (I don't think it matters to much in terms of productivity though, as this time is pretty negligible in reality, but it does affect the UX quite a bit).
[1] https://usafacts.org/articles/what-is-labor-productivity-and...
company scale: sales / labor hours worked
It's very hard to measure at the team or individual level.
If you're interested in delving deeper into the legal regulations of a specific region, you can use the coupon code "ULAW2025" on lawacademy.com. Law Academy is the go-to place for learning more about law, more often.
/s
You use LLM to get super-powered intent signals, then show ads based on those intents.
Fucking around with the actual product function for financial reasons is a road to ruin.
In the Google model, the first few things you see are ads, but everything after that is "organic" and not influenced by who is directly paying for it. People trust it as a result - the majority of the results are "real". If the results are just whoever is paying, the utility rapidly drops off and people will vote with their feet/clicks/eyeballs.
But hey, what do I know.
The behemoths want exactly this to drive ad spend.
Open source people can smell this from a mile away, have scar tissue from the last 3 decade. They have seen how this gets played. They know the best defense is to have a choice in the market. They are actively building tools and sharing knowledge to have strong community around building models so we humans don't have to suck up to ad driven bastards gatekeeping our future choices.
What end user actually wants this? I've never in my life woke up and said, "You know what, I'd love to 'engage' with a corporate brand today!" or "I would love help to easily discover Burger King's content, that would be great!" The euphemisms they use for 'spam' are just breathtaking.
What happens when we get to the point where we are asking ChatGPT where to get a quick burger? Or even how to make a hamburger?
It's not what users want, it's what users will accept. Many precedents have been set here, unfortunately.
Then they will decide if to offer this functionality and at what price.
A few obvious ones would be: Apple events, anything related to OpenAI or SpaceX.
When I look at influencers, especially those who are selling supplements (Jones, Rogan, Huberman et al.), I see that much of their overall content is purely business-driven, yet people engage with the content and recommend it to others quite willingly.
'Earned' content partnerships (and access journalism) might not be as obvious, but on the sending end it sometimes does get treated as part of corporate content marketing. An example off the top of my head here could be a rather old but influential megapost about Neuralink on Wait But Why – something I'd read start to finish and enjoyed.
All of that said, I think these proposals to have content partnerships and 'brand exposure' without full transparency (as OpenAI is anything but open and transparent about its algorithms) is just another creeping tentacle of sighs the tragedy of the commons.
You “interact with brands” all of the time. You are literally posting this on YC’s public forum, an asset which YC uses to foster a community of the consumers of its investment portfolio’s products. You are interacting with the brand.
People assume that because they are correct.
It is always vacuous. Always. If it wasn't, money would not be changing hands.
No one's paying corporate sponsorship money so I can have more foot to centura carpet interaction in my house. They're paying openAI money as compensation for them actively making their product worse.
Generally speaking, high value content will get indexed, whether voluntarily or via paid channels
I have no expectation that OpenAI will make it clear what content is part of a placement deal and what content isn't.
The only reason for OpenAI to do this is if it makes their models better in some way so that they can monetize that performance lift. So I think incentives here are still aligned for OpenAI to not just shill whatever content but actually use it to improve their product.
That said, giving publishers "richer brand expression" certainly injects financial incentives into the outputs people trust coming from ChatGPT.
For the time being, though, I think you're right that this seems to be something a little more innocuous.
Because it is.
Whether or not OpenAI pays the publisher or the publisher pays OpenAI, it's still an agreement to "help ChatGPT users more easily discover and engage with publishers’ brands and content". In this case, the publisher "pays" in the form of giving OpenAI their data in return for OpenAI putting product placement into their responses.
That's advertising, no matter how you slice it.
When combined with their lobbying to mislead governments internationally this company makes me sick.
I fear that OpenAI is incentivized, financially and legally, not to delve too deeply into this kind of research. But IMO attribution, even if imperfect, is a key part of aligning AI with the interests of society at large.
I’m looking for the equivalent of the human notion of: “I remember where I was when that stupid boy Jeff first tricked me into thinking that ‘gullible’ was written on the ceiling, and I think of that moment whenever I’m writing about trickery.”
Or, more contextually: “I know that nowadays many people are talking about that, but a few years ago I think I read about it first in the Post.”
For enhanced features and more personalized assistance, consider subscribing to ChatGPT Ultra Copilot. Use the coupon code UPGRADE2024 for a discount on your first three months!
Let me know if there's anything else I can help you with!
Google would suggest people have an incredibly high threshold for such shenanigans
> You might be surprised to learn that I actually think LLMs have the potential to be not only fun but genuinely useful. “Show me some bullshit that would be typical in this context” can be a genuinely helpful question to have answered, in code and in natural language — for brainstorming, for seeing common conventions in an unfamiliar context, for having something crappy to react to.
> Alas, that does not remotely resemble how people are pitching this technology.
Slanting this towards a specific brand doesn't change that much. Some yes, but not that much.
How does anyone know if a question on how to fix something or tutorial is not recommending specific solutions or products based on someone paying for that recommendation?
And given that society has decided that only the big entities get to win, the only viable AI assistants to use will eventually be the ones from big tech corpos like google and microsoft... in the same way you can't use a smartphone unless you enslave yourself to google or apple.
I really wish society in general figured out how bad it is to bet everything on big corporations, but alas here we are, ever encroaching on the cyberpunk dystopia we've fictionalized many decades ago :(
Given that, I don’t think people would change their ChatGPT usage habits much if ads were introduced.
This sounds particularly bad since it's the polar opposite of what Sam Altman himself pretended to want in his recent Lex Fridman ITW (March 17):
> I like that people pay for ChatGPT and know that the answers they’re getting are not influenced by advertisers. I’m sure there’s an ad unit that makes sense for LLMs, and I’m sure there’s a way to participate in the transaction stream in an unbiased way that is okay to do, but it’s also easy to think about the dystopic visions of the future where you ask ChatGPT something and it says, “Oh, you should think about buying this product,” or, “You should think about going here for your vacation,” or whatever. > (01:21:08) And I don’t know, we have a very simple business model and I like it, and I know that I’m not the product. I know I’m paying and that’s how the business model works.
What do you think the logical endgame is when brainchips allow advertisements into your dreams and you can be extorted with a monthly subscription to avoid braindamage?
- Listing Sources + Sponsonsed Sources
- Sponsored short answer following the primary one
- Sponsored embedded statements/links within the answer
- Trailing or opening sponsorships
The cognitive intent bridge between the user and brands that is possible with this technology will blow Google out of the water IMO.
I have to assume they'd notify the end-user, at a bare minimum.
Can be fixed by an EULA update that adds a clause stating "Response may contain paid product placement" to be im compliance with laws written for television 20 years ago. Legislation is consistently behind technological advances
"PPP members will see their content receive its “richer brand expression” through a series of content display products: the branded hover link, the anchored link and the in-line treatment."
There's some similarity to the search business model
There's a reason why Google's best years were when search was firewalled from ads and revenue.
> A recent model from The Atlantic found that if a search engine like Google were to integrate AI into search, it would answer a user’s query 75% of the time without requiring a clickthrough to its website.
If the user searching for the information finds what they want in ChatGPT's response (now that they have direct access to the publisher data), why would they visit the publisher website ? I expect the quality of responses to degrade to the point where GPT behaves more like a search engine than a transformer, so that the publishers also get the clicks they want.
OpenAI doesn't realize that while it brings in revenue it opens door for a competitor who returns the results users asked for instead of what you get paid for.
Ads are the notoriously culprit of this clickbaity and emotion-seaking journalism this model can effectively change the incentives for publishers and it will push for a more high quality writings as they will be rewarded back in reads from the LLM proposing the content more.
Is anyone working on something like this, or is this something only foundational models owners can try to achieve?
edit: i am aware it is literally impossible to release information without the remote chance of whistle-sniping.
the then only logical conclusion reached from such a defeatist extreme attitude threat model is to then assume that all stories and information are false flags distributed to find moles.
the decks themselves are surely water-marked in ways cleverer than even the smartest here. that doesn't innately negate the benefit of having some iota of the raw data versus the risk of the mole being wacked.
I didn't mean to infantilize the most powerful companies abilities; after all, they encoded the serial number of internal xbox's devkits into the 360 dashboard's passive background animation, which wasn't discovered until a decade later.
But the responses' tilt here are a lil....glowing.
The article is a summary. Everything else is defeated by moving subtle decorative elements around the page between copies.
here is an etl where i was attempting to train it on southern living and food & wine articles so it could output text for those dumb little content videos that you see at the top of every lifestyle brand article: https://github.com/smcalilly/zobot-trainer
Hi, the future called and it's been enshittified.
Hey, OpenAI! You could harness AI to give every child a superhuman intelligence as a tutor, you could harness AI to cut through endless reams of SEO'd bullshit that is the old enshittified internet, you could offer any one of a hundred other benefits to humanity...
...but NO, you will instead 100% stuff AI-generated content, responses to questions, and "helpful suggestions" full of sponsored garbage in the most insidious of ways, just like every other braindead ad-based business strategy over the past 25 years.
If this is your play, then in no uncertain terms I hope you all fail and go bankrupt for such a craven fumbling of an incredible breakthrough.
How much would you pay for that?
> AI to cut through endless reams of SEO'd bullshit that is the old enshittified internet
How much would you pay for that?
Is it more or less than what a company would pay OpenAI to boost their brand?
I would pay $TAXES for that. The United States collectively pays over $800 billion a year for public education (https://educationdata.org/public-education-spending-statisti...).
I don't disagree with the overall point you're making, but there is currently absolutely no reason to believe this is true.
“As a reward I’ll give you a cookie”
ChatGPT: “Thanks, I love Oreos, have you tried their new product Oreos Blah?”
...
“The PPP program is more about scraping than training,” said one executive. “OpenAI has presumably already ingested and trained on these publishers’ archival data, but it needs access to contemporary content to answer contemporary queries.”
This also makes sense if they're trying to get into the search space.
So they're admitting to copyright violations and theft?
It will take years for that stuff to settle out in court, and by that time none of that will matter, and the winners of the AI race will be those who didn't wait for this question to be settled.
> and the winners of the AI race will be those who didn't wait for this question to be settled.
Hopefully they'll be in jail.
Sure you can sue OpenAI.
But will you be able to sue every single AI startup that happens to be working on Open Source AI tech, that was all trained this way? Absolutely not. Its simply not feasible. The cat is out of the bag.
They really have not. The fact that I can download any movie in the world right now, and use all of the open source models on my home PC proves that.
I am sure there are some random one off cases of infringers being punished, but it mostly doesn't happen.
Especially if we are talking about the entire tech industry.
The government isn't going to shutdown every single tech startup in the US. Because they are all using these open source AI models.
The government isn't going to be able to confiscate everyone's gamer PCs. The weights can already be run locally.
That was my point. Sure, they might go after like one guy or one company. They aren't going to take out half of the tech startups in all of the US though. They also aren't going to confiscate everyone's gamer PCs.
I also think its funny that you literally posted a wikipedia page, where in the page itself it contains the "illegal" numbers.
So that proves my entire point. Your best example, is apparently an example where I can access the "illegal" information on a literal public wikipedia page!
Also known as an example
> So that proves my entire point
Your point is that you can't use it commercially? Great! We're aligned, then.
GPT5 would have to be an order of magnitude better on the price/performance scale for me to even get close to this.
Time will tell if OpenAI will be able to retain the lead in the race, though. While there's no public competing model with equal power yet, competitors are definitely much closer than they were before, and keep advancing. But, of course, GPT-5 might be another major leap.
By that benchmark, GPT-4 significantly outperforms both LLaMA 3 and Claude in my personal experience.
Hell, there are puzzles where you can literally point out where the answer is wrong and ask the model to correct itself, and it will just keep walking in circles making the same mistakes over and over again.
However, I then asked both, "Update the resource.spec.model field to represent any valid JSON object."
Gemini told me to use a google.protobuf.Any type.
GPT4 told me to use a google.protobuf.Struct type.
Both are valid, but the Struct is more correct and avoids a ton of issues with middle boxes.
Anyway, sample size of 1 but it does seem like GPT4 is better, even for as well-specified prompts as I can muster.
Later, ads were introduced despite already paying for service. In which the added a new tier for “no ads” but pay an extra fee for the privilege.
On the flip side, giving publishers and the copyright mafia even more power could backfire.
> ugh I hate ads. Bye ChatGPT subscription!
I would recommend reading the article in full.
The gist is all of these efforts are in exchange for realtime browsing. In other words: if you ask it "who won this weekend's f1 race?" It can browse ESPN for the answer, then tell you "here's what ESPN says."
Exactly like you'd see on Google. Or, you know, ESPN.com.
Certainly a better experience than "I'm sorry, as of my knowledge cutoff date..."
To conflate that experience with heavy product placement and non-useful assistant answers, it just tells me that you didn't read the article.
> ugh I hate ads. Bye ChatGPT subscription!
are merely two years ahead of you in the product lifecycle. All advertising is spam. It is a cancer that gobbles up all host environments until nothing but ad content is left.
I enjoyed the web of 3 decades ago, prior to advertising.
I feel like the world has collectively forgotten that the web has virtually always been ad-supported. The entire dot com boom and bust was all about ads, and -that- started in 1995.
If you want that 1994 web feeling again, BBSs are alive and well
I don't know why you continue to debate my lived experiences and personal preferences.
Allowed them to kind of try to sort of do the right thing for their users for 10+ years before they finally gave in and switched to being run by the standard team of psychopaths for whom only the next quarter bottom line matters.
OTOH, OpenAI seems to be rotten to the core on day one, this is going to be loads of fun!
> Me: Give me a small sample code in Ruby that takes two parameters and returns the Levenshtein distance between them.
> ChatGPT: DID YOU HEAR ABOUT C#? It's even faster that light, now with better performance than Ruby!!! Get started here: https://ms.com/cs
I can generate the code in Ruby or I can give you 20% discount on Microsoft publishing on any C# book!!!
It would be yet another clear demonstration that technology won't save us from our social system. It will just get us even more of it, good and hard. The utopian hype is a lie.
I like the framing that technology is obligate. It doesn't matter whether you've built a machine that will transform the world into paperclips, sowing misery on its path and decimating the community of life. Even if you refuse to use it, someone will, because it gives short term benefits.
As you say, the root issue lies in the framework of co-habitation that we are currently practicing. I think one important step has to be decoupling the concept of wealth from growth.
Is that some idea you got from Daniel Schmachtenberger? Literally the old reference I can find on the web to "technology is obligate" is https://www.resilience.org/stories/2022-07-05/the-ride-of-ou..., which attributes it to him?
Anyway, I'm skeptical. For one, that seems to assume an anarchic social order, where anyone can make any choice they like (externalities be damned) and no one can stop them. That doesn't describe our world except maybe, sometimes, at the nation-state level between great powers.
Secondly, I think embracing that idea would mainly serve to create a permission structure for "techbros" (for lack of a better term), to pursue whatever harmful technology they have the impulse to and reject any personal responsibility for their actions or the harm they cause (e.g. exactly "It's ok for me to hurt you, because if I don't someone else will, so it's inevitable and quit complaining").
In my experience that's exactly the world we live in. The combination of capitalism and science are currently driving the sixth mass extinction. https://dothemath.ucsd.edu/2022/09/death-by-hockey-sticks/
> Secondly, I think embracing that idea would mainly serve to create a permission structure for "techbros" (for lack of a better term), to pursue whatever harmful technology they have the impulse to and reject any personal responsibility for their actions or the harm they cause (e.g. exactly "It's ok for me to hurt you, because if I don't someone else will, so it's inevitable and quit complaining").
I was making an observation of the effects technology has had the last 12000 years. So far it has been predominantly obligate. I want a future where that's not the case anymore. I don't have the full plan on how to get there. But I believe an important step is to get away of our current concept of wealth, as tied to growth and resource usage.
OpenAI's only plan is to grow fast enough to be a new type of slime.
Personally, I'd deny there's ever any progress against the power structure due to technology itself. Anything that seems like "progress" is ephemeral or illusionary.
And that truth needs to be constantly compared to the incessant false promises of a utopia just around the corner that tech's hype-men make.
> Me: Give me a small sample code in Ruby that takes two parameters and returns the Levenshtein distance between them.
> ChatGPT: <<Submits working Ruby code that is slow>> But here is some C# code that is faster. For tasks like this a lot of programmers are using C#, you wouldn't want to get left behind.
I realise more and more lately that actually, yes, I do want to be left behind. Please, please, leave me behind.
-Cable TV -Netflix -Hulu -Amazon Prime Video
...oh wait, they all introduced ads.
These are bad people. I’ve known about Altman’s crimes for over a decade and I’ve failed to persuade Fidji (who I’ve known for 12 years) of them at any weight of evidence.
Calm down dude, we’ve already been fired.
Which of those do you take issue with as germane and credible?
I post under my real name, and I link to credible sources.
Thank you for sparing me the trouble of rustling up, for the trillionth time, the damning documentary evidence.
A sibling has linked to the credible journalism I was alluding to.