I think innocent-until-proven-guilty is the correct approach in general but... It's also immature to assume your data will stay private in the long term.
These AI companies are immensely valuable targets. When they get breached, their datasets will inevitably penetrate public datasets. And why would any AI company refuse to train on 'public' data?