I worry that some model provider will go and hire artists to draw pictures of pelicans on bicycles to make training data
1. Models need to be good at the questions we ask them, not the questions we could ask them.
2. The questions, at least partially, are correlated with information people consume.
3. People mostly consume viral content.
4. Ergo you should scrape viral content for training data.
This is for example the result of a taxidermied lion in Sweden when the guy doing the job never ever seen a lion or a photo of them and just worked off descriptions. https://www.snopes.com/articles/344637/the-lion-of-gripsholm...