Should charge AI for training on top of it or get them to donate. A small amount can fund them easily.
If Google just wanted them to exist and didn't care about profiting off of the search traffic they wouldn't partner with Mozilla.
This isn't me siding with AI companies by the way; it's a slippery slope argument.
Sometimes those two are in conflict, such that it will not be possible to satisfy both simultaneously.
The AI services have an option then to pay for this service, support a open service, or write their own crawler. I think if every open AI request didn't just do a web search but a more targeted arXiv search the results would be better.
I submit to open things because I want my material to be openly available. If I wanted restrictions, I would submit to gated journals.