What you can't do with scraped data is republish it verbatim. Doing a data analysis on scraped data is permitted by law, and you can publish your analysis of that data.
The question is, is an AI model trained on scraped data a derived analysis that is therefore legal? Or is it republishing of the original data? We need a test case to find out.
In the case of this dataset, I don't think the CC license applies to people using it. It "may" apply to redistribution of it for free. If the dataset was sold, that would be a violation. I suspect (after tested in court) a model trained on this dataset would be allowed despite the CC license on the photos.
Personally, in this case I think the ethics committee of the University should have put up barriers to the project. The morals of this are questionable at best.
0: https://techcrunch.com/2022/04/18/web-scraping-legal-court/