HNHacker News
TopNewBestAskShowJobs

mhamilton723

51 karma · joined March 5, 2019

submissionscomments
mhamilton723··on The Project Gutenberg Open Audiobook Collection
Hi all, we created this work and are happy to answer any questions!
mhamilton723··on Microsoft Announces General Availability of SynapseML
Website: https://microsoft.github.io/SynapseML/ Paper: https://arxiv.org/abs/2009.08044
mhamilton723··on Microsoft Releases Open Source ML Library for Distributed Search Engine Creation
I definitely would agree that you should pick the best tool for the job and not limit yourself to one ecosystem if it's too difficult.

One way to make Spark a bit easier to work with is through kubernetes or a tool like databricks that provides it as a service. Kubernetes, in particular, provides you a really nice amount of flexibility and composability when designing systems. One thing that we created to try to fill the gap of having to integrate System X with Spark was HTTP on Spark. This makes it easy to integrate Spark with other tools in a microservice architecture. When you couple this with containers, you can do a lot very quickly.

For datatypes I would look into different Spark connectors, these days there one for almost every database/streaming service/ cloud store under the sun.

This being said, Spark is a large piece of software that uses many different programming concepts which can be daunting. Our goal is to try to listen to feedback like this so we can try to make the Spark ecosystem a bit easier to use for everyone.

mhamilton723··on Microsoft Releases Open Source ML Library for Distributed Search Engine Creation
We have added support for integrating anything that communicates through HTTP into spark, so if you put it behind a service we can let you use it in spark. Also the mobius library, does a similar thing:

https://github.com/Microsoft/Mobius

mhamilton723··on Microsoft Releases Open Source ML Library for Distributed Search Engine Creation
Hey, thanks for your interesting point of view on Spark. I do a lot of my small data prototyping in Spark in single machine mode and it has worked for me thus far :).

For the likelihood functions comment, I would totally agree. Autograd libraries are easier to build custom likelihood models in, which is why we created CNTK on Spark, and databricks created Tensorflow on Spark. These give you the flexibility of modern deep learning stacks with the elasticity of spark

But in the end Spark is a single tool in a collection of tools and might not be right for your project, but it's been good for a lot of our work here at MSFT :)!

mhamilton723··on Microsoft Releases Open Source ML Library for Distributed Search Engine Creation
MMLSpark is Microsoft’s open source initiative for advancing distributed computing in Apache Spark. MMLSpark provides deep-learning, intelligent microservices, model deployment, model interpretability, image-processing, and many other tools to help you build scalable and production quality machine learning applications. MMLSpark is usable from Python, Scala, Java, and R and can be used on any Spark cluster. For more information and setup instructions, see our github page:

https://github.com/Azure/mmlspark

Thank you!

mhamilton723··on Gen Studio – An experimental collaboration across The Met, Microsoft, and MIT
Thanks for adding this!
mhamilton723··on Gen Studio – An experimental collaboration across The Met, Microsoft, and MIT
Very cool! You have left the realm of the MET Open Access data and ventured into the realm of the network that was pre-trained on imagenet!