1,236 karma · joined January 28, 2012
Founder of: http://konduit.ai/
Deep Learning guy. www.linkedin.com/in/agibsonccc/
Github: https://github.com/agibsonccc/
Twitter: @agibsonccc
Author: http://deeplearning4j.konduit.ai/
Based in Tokyo (yes not San Francisco)
Hi, I just want to expand on this a bit...
The "eclipse foundation" is actually not just sponsored by IBM. Nor is the IDE. As for what support it provides, that's actually spot on and is under a similar idea to what apache and linux provide. There are of course differences: but for most people the vague idea that open source foundations exist and provide some similar functions as a way to host projects is enough for this discussion
A lot of eclipse's revenue actually comes from working groups with Jakarta EE being the biggest one. There are also IOT, Automotive and many other working groups + projects under the eclipse umbrella.
1. Did you build this for your own use cases? Interesting side project?
2. How do you feel about the need for base64 being a requirement on the endpoints? Isn't GRPC the wrong medium for this? Also, what do you see as the main limitations right now? The models?
One interesting thing that could happen is the hardware gets better, and then these distributed schedulers might not be able to keep up with all the different options on the market.
There is also the tension of the hardware vendors wanting to give away things that only run on their chips vs the software makers who want things to run on every chip. It seems like there will be a lot of competition among the various infra players in the next few years now that nvidia is starting to have real competition now (even if it's not big yet)
It's good someone is building a company around it. I could see them building services on top of it and build a SAAS like databricks did with spark.
I'll be curious to see how ray matures.
Generally if you want python/java interaction, we maintain a tool called javacpp that handles this, we even bundle cpython: https://github.com/bytedeco/javacpp-presets
GraalVM itself also depends on javacpp for a portion of its features (LLVM wrapper): https://github.com/oracle/graal/blob/315f5dcf69c2e73fd13a5f8...
I'd be happy to answer questions about the overlap of the 2. I can say I happily execute python scripts from our embedded python and even point the python execution at an anaconda distribution.
There's a bit more to jvm than just training models. Java based application servers are still widely deployed for example.
There's a whole data engineering niche we target here as well (for the spark pipelines) where python isn't a good fit for that team.
Then there's the fact we're still easier to run keras on spark than any other framework thanks to our model import.
That doesn't account for what we're doing with inference. Many vendors in this space are just running kubernetes bundling other tools they don't control. We actually engage various large companies in custom chip development running other DL frameworks due to this low level control.
Depending on what you're looking to do, we're still the standard for pre compiled binaries compiled as jar files: https://repo1.maven.org/maven2/org/nd4j/nd4j-native/1.0.0-be...
You won't find any other framework running pre cooked avx binaries and IBM power at the same time.
Happy to talk about more depending on what your focus is.
Their focus is more on the c bindings and allowing other people to build what they want on top of that.
Other language bindings aren't generally going to be used for anything more than inference. First class actual data science work isn't going to happen in other languages anytime soon (at least outside of julia and R which are at least trying to compete in this niche).
Happy to answer questions about anything tech related there!
Sorry I just saw this post. We're actually intending on implementing allreduce as well. Right now, the initial focus was on implementing something more robust. There's a lot of elements of our distributed training that isn't talked about in the post(this was meant to be more of a high level overview).
A few additional things:
1. First and foremost fault tolerance and making spark run well was a bigger priority for us. Spark and its ilk don't do well with gpu clusters. Our initial focus was more taking what we had and making it work well and running everywhere with no code changes.
A common workflow that "just works" is being able to run model import on spark and run things as is. My colleague max is behind elephas (which we've since adopted for our python interface for dl4j on spark).
2. What we've seen many people don't have is MPI. We have this hard constraint of running things in strange on prem environments. So instead, we focus more on things like multi cast udp and compression/quantization to speed things up the networking as much as we can.
3. When we go to implement all reduce we'll be focusing on reusing as much of this as we can. We'll also try to figure out how to reuse our existing parts tha twork well like our cyclic buffer re use called workspaces: https://deeplearning4j.org/docs/latest/deeplearning4j-config...
Any other feedback folks have would be appreciated.
Java itself tends to be fairly verbose I know, but not only are there wrappers but for inference if you want to use python for basics and then us for inference. We enable several ways to do that. Folks tend to oversensationlize these "us vs them" framework "fights".
Rather than do that, why not just use what's good for each use case? Python for training, java for deployment?
Beyond that, you could also consider building your own wrapper maybe? In clojure, there's jutsu.ai for example: https://github.com/hswick/jutsu.ai
Ironically, 2 years after that I got in to YC and now run our company from Japan. Time flies!
An interesting example was a customer that bought a startup. The startup had a docker container that leaked memory (they were using an advanced TF library, google quit maintaining it). The parent company had tomcat infrastructure. They told the startup to get it to work in tomcat. They approached us.
Dl4j itself is built on top of a more fundamental building block that allows us to do some cool stuff in c: https://github.com/bytedeco/javacpp-presets
We maintain the above as well. We have our own tensorflow, caffe, opencv,onnx etc bindings. This gives you the same kind of pointer/zero copy interop you would in python but in java.
Another area we play in is RPA. You have windows consultants who don't know machine learning but deal with business process data all day. They want standardized APIs and to run on prem on windows: https://www.uipath.com/webinars/rpa-innovation-week/custom-b...
They want the ability to automatically maintain models after they are implemented with an integration that auto corrects the model. We do that using our platform's apis. On line retraining is another thing we do.
We also do automatically rollback based on feedback if you specify a test set. If we find the test is less accurate for a specified metric over time, we roll back the model. This prevents a problem in machine learning called concept drift which means the domain changes over time.
Lastly, you have customers generating their own models and you need to track/audit them all. Sometimes to do online retraining like the RPA use case. Some customers scale to millions of models. They need this kind of tracking.
Our model server just works with every format. We created our own DL framework and maintain that. Eg: We're the 2nd largest contributor to keras alongside having our own DL framework.
We've found it's easier to just have 1 runtime with PMML, TF file import, keras file import, onnx interop coupled with endpoints for reading and writing numpy, apache arrow's tensor format, and our own libraries binary format. We also have converters for r and python pickle to pmml as well. Happy to elaborate if interested.
For ingest we typically allow storage of transformations, accepting raw images, arrays etc. We have our own HDF5 bindings as well but mainly for reading keras files. Could you tell me a bit about what you might want for ETL from HDf5?
Most people don't need "industrial strength" till you hit a certain scale. They tend to optimize for ease of use/simplicity. It's one thing if you don't have to change their workflow, it's another if you have to not only have people change their workflow but also teach something new.
Convenience matters more. There's a pain threshold of "new tool" vs "this costs me x amount of time".
How are you guys overcoming this? Even though we're also in the infra space, I've never seen you guys out in the wild. Where would I bump in to you and what is the scale I would want to add the complexity of k8s + pachyderm + whatever other deps you guys have over just using an S3 bucket?
Another question: Why hasn't AWS just packaged this up and offered it as an extension of their k8s service? How are you guys going to overcome that?
Usually I see "hybrid cloud" or "on prem" as the response. If it is on prem, are you guys relying on the presence of k8s at customers? Do you guys use something like gravitational?
2. You have to know the right people. Go to the conferences and ask them. It took a long time. The profit shares aren't much. I more did it for the exposure at their conferences. We found it to be worth it as part of a larger package and networking with others in the field more than anything else. The notoriety has really helped as well. It's a lot of work, but definitely worth it if you can get it. Just be ready to put in a metric ton of work.
Attribution isn't the issue, it's balancing the need for building a community vs the financial incentives to actually support the people building the thing you're using.
I typically see this attitude at the developer level, but the line of business never sees it like that. Devs don't have incentive to care about these things. They also just want to deal with the vendor.
What you sell is more "insurance" and a guaranteed timeline on bug fixes/releases.
As someone who supports OSS, why should you be able to demand I ship something on a certain schedule if you're not paying me? If it's a business transaction, then there's a contract and aligned incentives on both sides.
Taking this a step further, this is usually not enough which is why we then see open core business models like elastic search, gitlab, (and also my company skymind) selling a combination of support + licensing.