Here's how I would leverage Hadoop infrastructure to use Pandas: delete Hadoop so I've got more disk space to run Pandas.
I don't get to train ML on terabytes of data very often. I do NLP, so "terabytes" means training a background model on the entire Common Crawl. Usually I'm doing something more specific and interesting than learning about random web pages. But when I do deal with the Common Crawl, I deal with it on one computer. Terabytes are not scary.
How does distributed computing even help? ML models need memory locality, sometimes to the extreme of being localized within a GPU's memory. And the limiting factor is the ability to iterate over the data. Sending the data over a network during training would be the worst thing you can do there.