Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.