Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
No, "all organizations" are not adopting Spark.
All organizations which already have a Hadoop cluster. You do not need Spark to use dataframes.
Never claimed this. To clarify Spark allows you to directly port Pandas code while leveraging existing Hadoop cluster infrastructure. And distributed computing is terrible for machine learning.
Distributed computing (Both traditional hadoop/spark and latest TF/PyTorch with parameter server) are essential for scaling ML beyond a certain point. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
If you have experience routinely training models on Terabytes of data intended for production deployment. I am happy to hear. There is a vast difference between training a model on your machine for research and building a reliable ML system that scales across large datasets and teams while taking infrastructure costs into account.Here's how I would leverage Hadoop infrastructure to use Pandas: delete Hadoop so I've got more disk space to run Pandas.
I don't get to train ML on terabytes of data very often. I do NLP, so "terabytes" means training a background model on the entire Common Crawl. Usually I'm doing something more specific and interesting than learning about random web pages. But when I do deal with the Common Crawl, I deal with it on one computer. Terabytes are not scary.
How does distributed computing even help? ML models need memory locality, sometimes to the extreme of being localized within a GPU's memory. And the limiting factor is the ability to iterate over the data. Sending the data over a network during training would be the worst thing you can do there.
delete Hadoop so I've got more disk space to run Pandas.
Except when you have ~2000 node cluster that runs 10,000 ETL tasks daily all of which are IO bound, And you cannot "just" uninstall. However during certain periods the same cluster has significant underutilization this opens up possibilities of doing lots of cool stuff for almost zero cost.I can understand your confusion. Training embedding model on Common Crawl is a toy problem. I recommend you thinl from perspective of a Tech company ideally in a production ML setting to understand the cost tradeoffs that go into making these decisions. Regarding memory locality if the problem is small enough its possible to tune Spark to use fewer or even just one worker with enough amount of memory allocated.
Sending the data over a network during training would be the worst thing you can do there.
When your data itself is in 100s of Terabytes and sharded across multiple racks, often a properly tuned spark pipeline is more reliable and performs well. Again there is a difference between using Common Crawl subset that you manually download filter train etc and in ensuring that models are updated/trained automatically daily across 100s of TB data.You are posturing.