1,876 karma · joined February 28, 2016
This is classic "if theory does not matches practice lets change the practice" approach, which is far too common in academic systems research.
I trust AWS with knowing what its customers truly want (in terms of performance and cost) and what it can provide,Since AWS has real financial stakes in its success.
A decade ago the same researchers would have mourned emergence of cloud computing as a wrong thing and instead asked for P2P computing since that's what they had spent the decade before doing research on.
You are being naive if you think academia is filled with do-gooders it’s just a race but of another kind.
Unless there is a direct impact of doing something more openly e.g. accessible code base that thousands of researchers can use to publish and cite your work quickly or significant risk of getting scooped. There is very little to motivate any change.
Also if you think journals/review system is bad, just get a glimpse of “grant review” system by NSF/NIH or “tenure committees” etc. They make the worst stack ranking performance review etc. seem light hearted fun.
E.g. Compute wars have only intensified with TPUs and FPGA. sure for training you might be okay with few 1080ti but good luck building any reliable, cheap and low latency service that uses DNNs. Similarly big data for academia is few terabytes but real Big data is Petabytes of street level imagery, Videos/Audio etc.
It comes built in with face detection/recognition, object detection, OCR, Open Images tagger etc.
"When you end up with a bunch of papers showing that genetic algorithms are competitive with your methods, this does not mean that we’ve made an advance in genetic algorithms. It is far more likely that this means that your method is a lousy implementation of random search."
Scanner is one of the first few tools to leverage Docker/Kubernetes by demonstrating ability to ship complex heterogenous architecture in a reliable/reproducible manner.
More than documentation, I would argue that TF especially tf.data lacks a tracing tool that would let a user quickly debug how data is being transformed and if there are any obvious ways to speed up. E.g. image_load -> cast -> resize vs image_load -> resize -> cast had different behavior and lead to hard to identify bugs. For tf.data prefetch which ends up being key to improving speed yet its is not documented, the only way I actually found out about it was by reading your TF.Data presentation.
A large scale visual data analytics platform, think SQL/MapReduce/Full-text search but for images and videos using Deep Learning. Now writing few papers on/using it to finish and get my PhD.
sorted([(k.weight, k.name) for k in somelist], reverse=True)In fact PhD's are merely a collateral damage in ending tution tax exemption which affects all students (and represents a much larger chunk when you account for number of undergrads). It's just that in other cases "per-student" effects are much lower, since either the student has no income or the lack of exemption mere dissuades parents from filling jointly.
Julia might still be good enough MATLAB replacement for Computational Simulation style tasks, but its clearly not suited for Machine learning.
If you disagree please show me how many ML researchers/labs/companies use Julia over Python/C++. Its cool to claim "Modern", "Deep", "ML", but I don't see any evidence.
[1] http://tvmlang.org/2017/08/17/tvm-release-announcement.html
today I’d still recommend the trusted laptop build approach for truly sensitive algorithms and computations.
You are utterly wrong. Algorithms and computations (especially ML kind) are never sensitive, its the data which is always sensitive. And that ALREADY exists on the cluster.If an Organization already has a Hadoop cluster containing data. You are suggesting that somehow having it downloaded to a secured laptop is better? Than say Spark running on top of the cluster? I think you are deeply mistaken. The cluster instances are already protected (if not you have a bigger problems). Also while an organization might not have Hadoop, they surely have an RDBMS, in which case the algorithms are even more useful.
delete Hadoop so I've got more disk space to run Pandas.
Except when you have ~2000 node cluster that runs 10,000 ETL tasks daily all of which are IO bound, And you cannot "just" uninstall. However during certain periods the same cluster has significant underutilization this opens up possibilities of doing lots of cool stuff for almost zero cost.I can understand your confusion. Training embedding model on Common Crawl is a toy problem. I recommend you thinl from perspective of a Tech company ideally in a production ML setting to understand the cost tradeoffs that go into making these decisions. Regarding memory locality if the problem is small enough its possible to tune Spark to use fewer or even just one worker with enough amount of memory allocated.
Sending the data over a network during training would be the worst thing you can do there.
When your data itself is in 100s of Terabytes and sharded across multiple racks, often a properly tuned spark pipeline is more reliable and performs well. Again there is a difference between using Common Crawl subset that you manually download filter train etc and in ensuring that models are updated/trained automatically daily across 100s of TB data. No, "all organizations" are not adopting Spark.
All organizations which already have a Hadoop cluster. You do not need Spark to use dataframes.
Never claimed this. To clarify Spark allows you to directly port Pandas code while leveraging existing Hadoop cluster infrastructure. And distributed computing is terrible for machine learning.
Distributed computing (Both traditional hadoop/spark and latest TF/PyTorch with parameter server) are essential for scaling ML beyond a certain point. Maybe you've worked at a job or two where nobody can comprehend not using distributed computing, as you describe, but it's nonsense to claim that "all organizations" work that way.
If you have experience routinely training models on Terabytes of data intended for production deployment. I am happy to hear. There is a vast difference between training a model on your machine for research and building a reliable ML system that scales across large datasets and teams while taking infrastructure costs into account. scaling for the fun of scaling.
I think your arguments are misguided, for every compute bound task where Hadoop/Spark undeperform, 1000 other ETL type tasks where hadoop is indispensable. As a result any organization running a large hadoop cluster will already have underused compute capacity for free and a maintainance staff which already taking care of the cluster. Thus from an organizations perspective the time difference is not material especially for batch jobs, this is the reason why Presto and Spark have been so successful. They enabled underutilized hadoop clusters to be used for ML and data science while delivering reasonable performance for Zero cost. also rules out e.g. the fairly common practice of using R or pandas for some ad-hoc processing.
This is essentially the reason why all organizations are adopting Spark, since it allows you to write imperative code on dataframes and build ML models.The reason Spark ecosystem has been so popular is because it enables these types of computations without breaking the model.