I'm not sure about the legality and ethics of training models on stolen data, but for reference, but so far as I can tell the Enron email data set is about 1.5GB (much lower than I expected to be honest!).
And I believe some of the more interesting things found in that data set (outside of the fraud) were people cheating on their partners.
https://www.kaggle.com/datasets/wcukierski/enron-email-datas...