It seems critical to have diverse, inclusive, and equitable data for model training. (I call this concept "DIET".)
Investors will throw money at startups claiming to make their own training data by consulting experts, finetuning as it is now will be obsolete, pre-ChatGPT internet scrapes will be worth their weight in gold. Once a block is hit on what we can do with data, the data itself is the next target.
If an x-ray means different things based off the race or gender we should make sure the model knows the race and gender.