I suspect anything you make in these, the thinking traces or "safety evaluations" of the content allows whatever you build to become RL or eventual pre-training data.
I suspect anything you make in these, the thinking traces or "safety evaluations" of the content allows whatever you build to become RL or eventual pre-training data.
In other words, what would the data be used to train for? It has to improve some kind of objective function. But it's a little hard to see what the objective function would be if it's just raw inner monologue.
If it was useless, they wouldn't bother.
It's probably not as useful as if you also have the prompts, but even assuming they really don't train on the prompts (which we can never verify), you can probably get to the prompts based on the monologue. Some agents essentially repeat the prompt in the monologue.
"The user asked me to build X using Y..."
What's neat about "safety evaluations" in LLM parlance is apparently encountering any novel information constitutes a "safety event" that can result in new reinforcement learning data... This anthropic 2022 paper that basically describes how it's a perpetual information siphoning machine with a cute little graphic https://arxiv.org/html/2212.08073 - they publish their "constitution". OpenAI has a "model spec" they somewhat regularly update that I suspect is their equivalent process. I suspect we've all been unwittingly advancing their models capabilities..
This Dec 2025 Google paper "A Practical Guide to Generating Synthetic Data With Differential Privacy" spells it out pretty clearly - the focus is on "privacy" - nothing about protecting the user's IP or unique knowledge/ insights.. https://arxiv.org/html/2512.03238v1
In OpenAI's case it seems to essentially generating synthetic training tuples from {prompt, chain-of-thought, answer} or scores on the chain of thought for reinforcement learning.
It would appear opting out of "Improve the model for everyone" didn't actually mean what we thought it meant.