1,236 karma · joined January 28, 2012
Founder of: http://konduit.ai/
Deep Learning guy. www.linkedin.com/in/agibsonccc/
Github: https://github.com/agibsonccc/
Twitter: @agibsonccc
Author: http://deeplearning4j.konduit.ai/
Based in Tokyo (yes not San Francisco)
What I'm specifically talking about is even that kinda hacky experiment code you end up writing. I don't try to implement whole projects in there, but even just "train this model" type code ends up being a hassle because of how bad the editors are.
My above comment was more referencing wishing I could spend more time writing experiment code in jupyter without copying and pasting all the time.
How do people cope with this? Do you supplement it with other tools? I spend a lot of my time in an IDE and then just paste some of the code in to cells. That seems easier.
the same point. Tech doesn't matter. Simplicity does.
Even in our own product line, we only do a small
subset of this. We don't even require a cluster
to run. We also work with tech that people use.
You are currently competing with horovod
and kubeflow. eg: "competing with free"
You need more than that to survive.
Generally, that comes down to services.
Your pitch is still about differentiated tech, not a large install base, a differentiated business model
and something related to people like a good partner ecosystem.
Your pitch here requires tons of services.
People don't know how to use all of this stuff especially on prem.
It takes more than just code to build a business.
I say this as someone who's been doing this since 2013. It's not easy.
Dl4j itself has a decent sized user base. Ranging likely from your phone maker to your bank and retail store.
We have our own software distro too which is why I'm commenting on this. We don't try to boil the ocean with a bunch of tech though.
There's a whole new crop of companies focusing on solving bits of the ML problem well rather than trying to do storage and god knows what else.
My point here about you guys is you're trying to compete in what is largely a commodity market. People don't need all this stuff. Simplicity won here. It's not about better tech.
You guys have the same pitch MapR does and largely the same problem: Better tech is only part of the problem with adoption. You need customers, users, and a clear business model when going to market.
Cloudera and Horton ran one playbook that at least somewhat worked (it got them public) and now they can focus on competing with the cloud vendors, which made the right decision and just made commonly used software easy to use.
They merged partially because they were both being cannibalized by services revenue they couldn't get rid of. Now they are struggling to move to the cloud (see Atlas now)
I'm not sure yet another hadoop distro with a bunch of 1 off tooling that is supposedly faster is the answer. Why is all that stuff even needed?
Beyond that, commenting on the translation a bit. They did live translation the first day of WAIC for the headline speakers. There were 2 screens, 1 was baidu and the other was iflytek. Neither were that good on the english side (it was ok..but could barely keep up with the speakers)
They claim they are still working on english. The grammar output wasn't coherent. IFlytek itself has some neat hardware they sell that is pretty good.
Beyond that, it seems like they are mainly collecting data right now. I would not be surprised they were doing this just for marketing visibility. It is easy to fake.
Happy to answer questions about the experience there if people would find it useful.
Azul publishes this as well: https://www.azul.com/products/azul_support_roadmap/
Cursory google searching is all it takes to find these things.
Origin of the term here: https://www.reactivemanifesto.org/
Graal is an R&D effort by a few folks in the java community driven by oracle right now to allow AOT compilation among other things for the JVM: http://www.oracle.com/technetwork/java/jvmls2015-wimmer-2637...
Java is generally not for apps that require fast VM spin up times. That being said, for what the JVM ecosystem can do it isn't bad in practice for heavy server applications.
I know gravitational mainly works in the go ecosystem which has its own trade offs there, but considering what else java already has a mature ecosystem for (big data, well understood native internals, other language built on top of it) the startup times and weaknesses while not ideal haven't been a show stopper.
GraalVM is an attempt to address a wide variety of problems you get with JVM startup time, GC etc.
At most it's simplistic right now and will be for a while. The fact they are working on it is really cool though!
Hope that helps!
That codegen isn't going to match what you need to do for real speed on cpus or gpus when writing vectorized math code.
Re: his last point. That's exactly what we talked to that team about. We don't feel those tools are going to work for real world use cases. We already do the codegen and auto bindings/mapping ourselves in addition to the memory management ourselves.
I can say for a fact that panama is not seriously targeting this space. We implement a ton of that native code today that works with c++ and actual android today. We also handle gpus. Project panama is only targeting c, and even then will only do it a cross platform non committal fashion. They aren't doing it the way they should be in order to properly target native vectorized code.
We know this from experience, because this is all we do: https://github.com/deeplearning4j/deeplearning4j https://github.com/bytedeco/javacpp-presets
We tried seeing if we could get some of this work in to the JDK, but their goals fundamentally compete with what it takes to get vector math to be fast. It's also not nearly as ambitious as it needs to be to handle real world tensor workloads.
That being said, while gemm is one op, it's a lot more than just jni back and forth that use other libraries. What matters here are also things like convolutions, pair wise distance calculations, element wise ops, etc.
There's nuance there.
There are multiple layers here to consider:
1. The JNI interop managed via javacpp (relevant to this discussion)
2. Every op has allocation vs in place trade offs to consider
3. For our python interface, we have yet another layer to benchmark there (we use pyjnius for jumpy the python interface for nd4j)
4. Op implementations for the cuda kernels and the custom cpu ops we wrote. (That's where our avx512 and avx2 jars matter for example)
For the subset we are comparing against, it's basically making sure we wrap the blas calls properly. That's definitely something we should be doing.
We've profiled that and chose the pattern you're seeing above with f ordering.
That is where we are fast and chose to optimize for. You are faster in those other cases and have laid that out very well.
Again, there's still a lot that was learned here and I will post the doc when we get it out there to make that less painful next time.
You made a great post here and really laid out the trade offs.
I wish we had more time to run benchmarks beyond timing for our own use cases, if we had smaller scope we would definitely focus on every case you're mentioning here. We likely will revisit this at some point if we find it worth it.
In general, our communications and docs can always be improved (especially our internals like our memory allocation)
Re: your last point we do do this kind of benchmarking with tensorflow. For example: https://www.slideshare.net/agibsonccc/deploying-signature-ve... (see slide 3 and also the broader slides for an idea of how we profile deep learning for apps using the jvm)
We need to do a better job of maintaining these things though. We don't keep it up to date and don't profile as much as we should. It has diminishing returns after a certain point vs building other features.
I'm hoping a CI build to generate these things is something we get done this year so we can both prevent performance regressions and have consistent numbers we can publish for the docs.
Once the python interface is done that will be easier to do and justify since most of our "competition" is in python.
We will be sending out a doc for this by next week with these updates. Thanks a lot for playing ball here.
Beyond that, can you clarify what you mean? Do you mean just the gemm op?
For that, that's the only case that mattered for us. We will be documenting the what/how/why of this in our docs.
Beyond that, I'm not convinced the libraries are directly comparable when it comes to the sheer scope of the libraries to each other.
You're treating nd4j as a gemm library rather than a fully fledged numpy/tensorflow with hundreds of ops and support for things you would likely have no interest in building.
A big reason I built nd4j was to solve the general use case of building a tensor library for deep learning, not just a gemm library.
Beyond that - I'll give you props for what you built. There's always lessons to learn when comparing libraries and making sure the numbers match.
Our target isn't you though, it's the likes of google,facebook, and co and tackling the scope of tasks they are.
That being said - could we spend some time on docs? Heck yeah we should. At most we have java doc and examples. We tend to help people as much as we can when profiling.
Could we manage it better? Yes for sure. That's partially why we moved dl4j to the eclipse foundation to get more 3rd party contributions and build a better governance setup. Will it take time for all of this to evolve? Oh yeah most definitely.
No project is perfect and always has things it could improve on.
Anyways - let's be clear here. You're a one man shop who built an amazingly fast library that scratches your own itch for a very specific set of use cases. We're a company and community tackling a wider breadth of tasks and trying to focus more on serving customers and adding odd things like different kinds of serialization, spark interop,.. etc.
We benefit from doing these comparisons and it forces us to document things better that we normally don't pay attention to. This little exercise is good for us. As mentioned, we will document the limitations a bit better but we will make sure to cover other topics like allocation and the like as well as the blas interface.
Positive change has come out of this and I'd like to thank you for the work you put in. We will make sure to re run some of the comaprisons on our side.
I would be careful to even call them just a telco.
I can say for a fact they have other robotics efforts as well. Yes they have not had the best luck in monetizing them.
Ironically, we are associated with some of their newer R&D efforts in robotics: https://skymind.ai/press/softbank
I can say they have not made ideal decisions in a lot of areas, but I wouldn't count them out based on 1 bad robot. They are a lot more diverse than that.
Source: I have softbank as a customer, have dealt with numerous executives here, and actually live in japan. I don't just read the news.
FWIW: Production is an overloaded term. CI may not even be applicable here. Say you're doing batch inference where you need to run jobs every 24 hours on a large amount of data: That might be tied to some cron job.
That being said you could use a CI system for that in theory.
There are also other factors here: What other kind of things do you want to track? Experiments results? Wrong results by your machine learning algorithm?
Concisely: What kind of deployment requirements do you have and what are your goals?
If you are edoing real time, what does "deployment" even mean? Are you serving in real time via a rest api? Are you doing streaming? What are your throughput requirements? What about latency? Is that even hooked up to a CI system?
Something that is vaguely related: How do you test the accuracy of different models across your cluster? Say you want to do a self deployment, what if you want to tie that to say: a workspace where you produced the results?
Is that hooked up to a CI system? If so, what's your use case?
Then there's that common hand off from data scientist to production, what does that look like? A sibling thread mentioned some of these things.
If anyone else is curious about this stuff, we deploy deep learning models in locked down environments both on kubernetes as well as touching hadoop clusters. Happy to answer questions.