I assume it's a subsidy to get more training data.
EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?
I assume it's a subsidy to get more training data.
EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?
(I work at OpenAI.)
So what is the value prop then? Just basic supply and demand?
FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.
ChatGPT enterprise: By default, no training (opt in).
ChatGPT personal: By default, training (opt out).
Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."
People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."
But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.
When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc.
Places where I can imagine potential cracks in the literal interpretation of what I said are things like a financial analyst who does a statistical fit to predict revenue next quarter using a model based on last quarter's aggregate token consumption, which in some sense embodies your metadata (the length of your conversations) in a sea of other data. Or perhaps an infrastructure planner who makes a little model of internet bandwidth by time of day to help plan when we need a data center networking upgrade. Maybe things like these are technically training on your data in the most pedantic sense, but definitely not in the sense that most of us mean.
I promise you we're not doing any gimmicks where we transform your data and then pretend ah because it's transformed it's not your data.
1. Legal loopholes given OpenAI's advertising aspirations and model training needs
2. Data retention and rising threats of fascism that historically have not served the persecuted very well when fascist regimes get access to said data
3. Risk from centralized collection of that data with a company whose software I do not control in a world where enshitification and lock-in is the norm.
I really wish OpenAI did more to espouse exactly this: "When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc." and ideally provide technical reasurrances that this is impossible (eg: certain technical ZDR approaches, etc.).
Do you happen to have a favorite reference to point me at that would document some of those official assurances to the nuanced detail we've discussed here?
ChatGPT: https://help.openai.com/en/articles/5722486-how-your-data-is...
ChatGPT data controls: https://help.openai.com/en/articles/7730893-data-controls-in...
If you have feedback on how to improve these, happy to consider it.
Looks like we phrase it as "your new conversations won’t be used to train OpenAI models" which is hopefully clearer than "OpenAI models will not be trained on your conversations", which could leave open the possibility of derived data or something.