HNHacker News
TopNewBestAskShowJobs

agibsonccc

1,236 karma · joined January 28, 2012

Contact me: agibsonccc@gmail.com

Founder of: http://konduit.ai/

Deep Learning guy. www.linkedin.com/in/agibsonccc/

Github: https://github.com/agibsonccc/

Twitter: @agibsonccc

Author: http://deeplearning4j.konduit.ai/

Based in Tokyo (yes not San Francisco)

submissionscomments
agibsonccc··on Why There Will Never Be Another RedHat: The Economics of Open Source (2014)
Documentation is a moving target. Depending on your project, you often have to mix background explanations with code samples. Fields evolve. New people show up. Documentation can never really be good enough. Books provide a way of filling in certain gaps in a structured way that gives people more background. Often times, people expect documentation to do more than this. I tend to agree with you, but it's not what end users end up needing a lot of the time.
agibsonccc··on Why There Will Never Be Another RedHat: The Economics of Open Source (2014)
Oreilly author and project lead for an open source eclipse foundation (apache style project though!) project that runs a company. Happy to answer questions here about the how/what/why this happens.
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
You can generally persist the results your self to disk though. Especially since a lot of things end up being numpy arrays. So you run 1 script that saves all the results, and another that loads it and runs just the part of your workflow you want. Bonus: it's persisted to disk on top of that! I know things get more complicated than that, but I'd say the compelling use case for notebooks isn't the state saving but more the whole package in one place (state persistence,visualization, interactive repl,..)
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Java's tooling for this is among the top. We're spoiled compared to the dynamically typed languages :)
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Yeah I agree. We do something similar if we're using zeppelin or beaker. I organize it, put an uber jar in there and then run everything from there. That's a ton easier.
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Yeah but the whole point is "interactive coding". It doesn't feel very interactive when I have to context switch all the time :). I'd prefer something closer to what the lisp folks get to do with the repl where you can scratch out an idea and see it working without leaving your environment.
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Data frame rendering in the various notebooks (beaker,jupyter,zeppelin,..) is wonderful. Your workflow sounds closest to what I do. If I want to visualize something I tend to compile my thoughts/imports and organize things in an editor first and put it in a notebook in parallel. It helps with version control as well.
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Already a customer, you have nothing you can sell me :).
agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
Oh I won't argue you with you there. I just find myself rotating quite a bit because I have to do both deployment as well as writing code for experimentations.

What I'm specifically talking about is even that kinda hacky experiment code you end up writing. I don't try to implement whole projects in there, but even just "train this model" type code ends up being a hassle because of how bad the editors are.

My above comment was more referencing wishing I could spend more time writing experiment code in jupyter without copying and pasting all the time.

agibsonccc··on Why Jupyter is data scientists’ computational notebook of choice
The only thing that stops me from being able to use notebooks full time is their intellisense compared to IDEs is horrible. I like being able to use them for demos/presentations, but I can't imagine trying to code within one primarily. Especially when it comes to tracking results.

How do people cope with this? Do you supplement it with other tools? I spend a lot of my time in an IDE and then just paste some of the code in to cells. That seems easier.

agibsonccc··on Cloudera and Hortonworks merge
Exactly!
agibsonccc··on Cloudera and Hortonworks merge
I don't think I changed my angle here? I'm still addressing

the same point. Tech doesn't matter. Simplicity does.

Even in our own product line, we only do a small

subset of this. We don't even require a cluster

to run. We also work with tech that people use.

You are currently competing with horovod

and kubeflow. eg: "competing with free"

You need more than that to survive.

Generally, that comes down to services.

agibsonccc··on Cloudera and Hortonworks merge
You bundle way more than they do and on top of that have your own file system just like mapr does.

Your pitch is still about differentiated tech, not a large install base, a differentiated business model

and something related to people like a good partner ecosystem.

Your pitch here requires tons of services.

People don't know how to use all of this stuff especially on prem.

It takes more than just code to build a business.

I say this as someone who's been doing this since 2013. It's not easy.

agibsonccc··on Cloudera and Hortonworks merge
According to my customers and users yes? Granted, we do a lot more than just that though.

Dl4j itself has a decent sized user base. Ranging likely from your phone maker to your bank and retail store.

We have our own software distro too which is why I'm commenting on this. We don't try to boil the ocean with a bunch of tech though.

There's a whole new crop of companies focusing on solving bits of the ML problem well rather than trying to do storage and god knows what else.

My point here about you guys is you're trying to compete in what is largely a commodity market. People don't need all this stuff. Simplicity won here. It's not about better tech.

You guys have the same pitch MapR does and largely the same problem: Better tech is only part of the problem with adoption. You need customers, users, and a clear business model when going to market.

Cloudera and Horton ran one playbook that at least somewhat worked (it got them public) and now they can focus on competing with the cloud vendors, which made the right decision and just made commonly used software easy to use.

agibsonccc··on Cloudera and Hortonworks merge
Both Horton and Cloudera require a huge number of partner resellers and consultants. What horton and cloudera have isn't just software, but brand, an install base, and the ability to back it up.

They merged partially because they were both being cannibalized by services revenue they couldn't get rid of. Now they are struggling to move to the cloud (see Atlas now)

I'm not sure yet another hadoop distro with a bunch of 1 off tooling that is supposedly faster is the answer. Why is all that stuff even needed?

agibsonccc··on AI Company Accused of Using Humans to Fake Its AI
I was a speaker at a conference they were doing live annotation for. (http://waic2018.com/index-en.html) See slide deck here for proof: https://www.slideshare.net/agibsonccc/world-artificial-intel...

Beyond that, commenting on the translation a bit. They did live translation the first day of WAIC for the headline speakers. There were 2 screens, 1 was baidu and the other was iflytek. Neither were that good on the english side (it was ok..but could barely keep up with the speakers)

They claim they are still working on english. The grammar output wasn't coherent. IFlytek itself has some neat hardware they sell that is pretty good.

Beyond that, it seems like they are mainly collecting data right now. I would not be surprised they were doing this just for marketing visibility. It is easy to fake.

Happy to answer questions about the experience there if people would find it useful.

agibsonccc··on Java is still available at zero-cost
Yes not all of it - there are still android specific bits yet. Enough of it is openjdk though.
agibsonccc··on Java is still available at zero-cost
They are the exact same. Oracle moved the core bits to openjdk long ago. See: https://jaxenter.com/green-interview-jdk-binaries-137078.htm...
agibsonccc··on Java is still available at zero-cost
Not sure that's relevant either. Within the target environments you're still looking at 5 to 10 years just for red hat: https://access.redhat.com/articles/1299013

Azul publishes this as well: https://www.azul.com/products/azul_support_roadmap/

Cursory google searching is all it takes to find these things.

agibsonccc··on Java is still available at zero-cost
What's being missed here is there is a whole slew of vendors that support openjdk. Those patches will come from red hat, ibm, azul,.. as well.
agibsonccc··on Java is still available at zero-cost
Use openjdk instead. By default on linux, that's what will continue to happen. You won't really be impacted by this.
agibsonccc··on Java is still available at zero-cost
You're a long time out of date at this point. Google merged openjdk a long time ago for their base JVM.

https://www.theregister.co.uk/2015/12/30/android_openjdk/

agibsonccc··on Vert.x native: how to run a JVM app in 10Mb of RAM
To answer some of your questions (I use vertx and mainly focus on the JVM world) - reactive is actually a widely used term in the java world to mean "event loop with async programming.

Origin of the term here: https://www.reactivemanifesto.org/

Graal is an R&D effort by a few folks in the java community driven by oracle right now to allow AOT compilation among other things for the JVM: http://www.oracle.com/technetwork/java/jvmls2015-wimmer-2637...

Java is generally not for apps that require fast VM spin up times. That being said, for what the JVM ecosystem can do it isn't bad in practice for heavy server applications.

I know gravitational mainly works in the go ecosystem which has its own trade offs there, but considering what else java already has a mature ecosystem for (big data, well understood native internals, other language built on top of it) the startup times and weaknesses while not ideal haven't been a show stopper.

GraalVM is an attempt to address a wide variety of problems you get with JVM startup time, GC etc.

At most it's simplistic right now and will be for a while. The fact they are working on it is really cool though!

Hope that helps!

agibsonccc··on Interfacing with native methods on Graal VM
Yes that's what I stated above. I've also stated that I haven't just read the news. We've talked to that team physically. Being language/platform neutral does not mean it is going to fulfill most use cases people would have for c bindings. Java tends to be "good enough" for a lot of use cases out of the box. It might help a bit with libraries like netty and memory management, but it's not going to work on real world math code which, as I stated, is our main use case.

That codegen isn't going to match what you need to do for real speed on cpus or gpus when writing vectorized math code.

Re: his last point. That's exactly what we talked to that team about. We don't feel those tools are going to work for real world use cases. We already do the codegen and auto bindings/mapping ourselves in addition to the memory management ourselves.

agibsonccc··on Interfacing with native methods on Graal VM
Disclaimer: I'm affiliated with a semi competing project to panama called javacpp: https://github.com/bytedeco/javacpp

I can say for a fact that panama is not seriously targeting this space. We implement a ton of that native code today that works with c++ and actual android today. We also handle gpus. Project panama is only targeting c, and even then will only do it a cross platform non committal fashion. They aren't doing it the way they should be in order to properly target native vectorized code.

We know this from experience, because this is all we do: https://github.com/deeplearning4j/deeplearning4j https://github.com/bytedeco/javacpp-presets

We tried seeing if we could get some of this work in to the JDK, but their goals fundamentally compete with what it takes to get vector math to be fast. It's also not nearly as ambitious as it needs to be to handle real world tensor workloads.

agibsonccc··on Neanderthal vs. ND4J – Native Performance, Java and CPU
Yeah we definitely need to spend some more time on benchmarks after all it's said and done.

That being said, while gemm is one op, it's a lot more than just jni back and forth that use other libraries. What matters here are also things like convolutions, pair wise distance calculations, element wise ops, etc.

There's nuance there.

There are multiple layers here to consider:

1. The JNI interop managed via javacpp (relevant to this discussion)

2. Every op has allocation vs in place trade offs to consider

3. For our python interface, we have yet another layer to benchmark there (we use pyjnius for jumpy the python interface for nd4j)

4. Op implementations for the cuda kernels and the custom cpu ops we wrote. (That's where our avx512 and avx2 jars matter for example)

For the subset we are comparing against, it's basically making sure we wrap the blas calls properly. That's definitely something we should be doing.

We've profiled that and chose the pattern you're seeing above with f ordering.

That is where we are fast and chose to optimize for. You are faster in those other cases and have laid that out very well.

Again, there's still a lot that was learned here and I will post the doc when we get it out there to make that less painful next time.

You made a great post here and really laid out the trade offs.

I wish we had more time to run benchmarks beyond timing for our own use cases, if we had smaller scope we would definitely focus on every case you're mentioning here. We likely will revisit this at some point if we find it worth it.

In general, our communications and docs can always be improved (especially our internals like our memory allocation)

Re: your last point we do do this kind of benchmarking with tensorflow. For example: https://www.slideshare.net/agibsonccc/deploying-signature-ve... (see slide 3 and also the broader slides for an idea of how we profile deep learning for apps using the jvm)

We need to do a better job of maintaining these things though. We don't keep it up to date and don't profile as much as we should. It has diminishing returns after a certain point vs building other features.

I'm hoping a CI build to generate these things is something we get done this year so we can both prevent performance regressions and have consistent numbers we can publish for the docs.

Once the python interface is done that will be easier to do and justify since most of our "competition" is in python.

agibsonccc··on Neanderthal vs. ND4J – Native Performance, Java and CPU
Fair point we are fixing now: https://github.com/deeplearning4j/deeplearning4j-docs/issues...

We will be sending out a doc for this by next week with these updates. Thanks a lot for playing ball here.

Beyond that, can you clarify what you mean? Do you mean just the gemm op?

For that, that's the only case that mattered for us. We will be documenting the what/how/why of this in our docs.

Beyond that, I'm not convinced the libraries are directly comparable when it comes to the sheer scope of the libraries to each other.

You're treating nd4j as a gemm library rather than a fully fledged numpy/tensorflow with hundreds of ops and support for things you would likely have no interest in building.

A big reason I built nd4j was to solve the general use case of building a tensor library for deep learning, not just a gemm library.

Beyond that - I'll give you props for what you built. There's always lessons to learn when comparing libraries and making sure the numbers match.

Our target isn't you though, it's the likes of google,facebook, and co and tackling the scope of tasks they are.

That being said - could we spend some time on docs? Heck yeah we should. At most we have java doc and examples. We tend to help people as much as we can when profiling.

Could we manage it better? Yes for sure. That's partially why we moved dl4j to the eclipse foundation to get more 3rd party contributions and build a better governance setup. Will it take time for all of this to evolve? Oh yeah most definitely.

No project is perfect and always has things it could improve on.

Anyways - let's be clear here. You're a one man shop who built an amazingly fast library that scratches your own itch for a very specific set of use cases. We're a company and community tackling a wider breadth of tasks and trying to focus more on serving customers and adding odd things like different kinds of serialization, spark interop,.. etc.

We benefit from doing these comparisons and it forces us to document things better that we normally don't pay attention to. This little exercise is good for us. As mentioned, we will document the limitations a bit better but we will make sure to cover other topics like allocation and the like as well as the blas interface.

Positive change has come out of this and I'd like to thank you for the work you put in. We will make sure to re run some of the comaprisons on our side.

agibsonccc··on The impact of Masayoshi Son’s $100B fund will be profound
Softbank actually has numerous business units ranging from reselling/distribution of various products to monetizing telco data.

I would be careful to even call them just a telco.

I can say for a fact they have other robotics efforts as well. Yes they have not had the best luck in monetizing them.

Ironically, we are associated with some of their newer R&D efforts in robotics: https://skymind.ai/press/softbank

I can say they have not made ideal decisions in a lot of areas, but I wouldn't count them out based on 1 bad robot. They are a lot more diverse than that.

Source: I have softbank as a customer, have dealt with numerous executives here, and actually live in japan. I don't just read the news.

agibsonccc··on Apple open-sources FoundationDB
I'm not convinced a paper with a website and some proof of concepts would be considered the "best". You're throwing around a bunch of components in to a distro calling yourselves everything from "deep learning" to a file system. It's not clear what you guys are even trying to do here.
agibsonccc··on Ask HN: What's Your CI/CD Workflow for Your Machine Learning Projects?
Disclaimer: I compete in this space and may compete with your internal team, your cloud vendor, or something you are interested in.

FWIW: Production is an overloaded term. CI may not even be applicable here. Say you're doing batch inference where you need to run jobs every 24 hours on a large amount of data: That might be tied to some cron job.

That being said you could use a CI system for that in theory.

There are also other factors here: What other kind of things do you want to track? Experiments results? Wrong results by your machine learning algorithm?

Concisely: What kind of deployment requirements do you have and what are your goals?

If you are edoing real time, what does "deployment" even mean? Are you serving in real time via a rest api? Are you doing streaming? What are your throughput requirements? What about latency? Is that even hooked up to a CI system?

Something that is vaguely related: How do you test the accuracy of different models across your cluster? Say you want to do a self deployment, what if you want to tie that to say: a workspace where you produced the results?

Is that hooked up to a CI system? If so, what's your use case?

Then there's that common hand off from data scientist to production, what does that look like? A sibling thread mentioned some of these things.

If anyone else is curious about this stuff, we deploy deep learning models in locked down environments both on kubernetes as well as touching hadoop clusters. Happy to answer questions.

← PreviousPage 2 of 25Next →