I get the impression that several people working on debian couldn't work this one out!
I get the impression that several people working on debian couldn't work this one out!
It is a pretty straightforward process:
http://www.cloudera.com/content/cloudera/en/documentation/cd...
What debian users/hackers/amateur admins like me really want is packages that are first class citizens, that the debian guys have picked up, sanitised, analysed and made part of the system.
I'll take software from the debian repos every time if I can. And it's pretty damning if people who are familiar with build systems and package creation can't figure it out!
What does the language have to do with the program?
Hadoop is what it is because it's a complex problem with a fittingly complex solution. Simply re-writing it in your pet language won't somehow make it "better".
(I have no horse in this race, I am just writing what I think the grandparent comment was referring to)
Just like Java[0]. It is just a matter of choosing the right compiler for the use case at hand.
[0] - http://www.excelsiorjet.com/ (one from many vendors)
Almost all commercial JVMs have some form of AOT or JIT caching, specially those that target embedded systems.
Sun never added support to the reference JVM for political reasons, as they would rather push for plain JIT.
Oracle is now finally thinking about adding support for it, with no official statement if it will make it into 9 or later.
JEP 197 is the start of those changes, http://openjdk.java.net/jeps/197
Oracle Labs also has SubstrateVM, which is an AOT compiler built with Graal and Truffle.
GCC only keeps gcj around due to its unit tests.
So producing a binary which doesn't require a separate runtime really isn't a problem.
I'm not an expert on Java, but my experience with it is that it's runtime is fairly huge and requires custom installation.
For that matter, as you allude even C has a runtime.
Figuring out which software packages I needed, how to modify my environment variables, which compiler to get, and where to put everything in the correct directory was the entire difficulty.
If it were written in Go instead of Java, I could have done `go get apache.org/hadoop` and it would have been done instead of giving up after hours of frustration.
Go has almost no new features that make it an interesting language from a programming language perspective. Go's win is that it makes the actual running of real software in production better. Hadoop's difficult is exactly why InfluxDB exists at all.
This complaint is just about packaging, and not the language itself. Any project can have good or back packing scripts, and for Java there are plenty of ways to make it "good".
Not to mention, the BUILDING.txt document clearly states they use maven[1] and to build you just do: mvn compile
> Go's win is that it makes the actual running of real software in production better
This might just be a familiarity issue, because once you launch the program, all things are equal.
And yes, you can bundle a JVM with your java app, which makes it exactly like GO's statically linked runtime and just as portable without any fuss.
[1] https://github.com/apache/hadoop/blob/trunk/BUILDING.txt
Go gets us better performance and concurrency out of the box.
Than Java? At best, GO performs on par with Java, but is often measured 10-20% slower.[1][2][3]
This is usually attributed to the far more mature optimizing compiler in the JVM, which ultimately compiles bytecode down to native machine code, especially for hot paths. Java performance for long running applications is on par with C (one of the reasons it's a primary choice for very high performing applications such as HFT, Stock Exchanges, Banking, etc).
> concurrency out of the box.
Java absolutely supports concurrency "out of the box"...[4]
[1] http://zhen.org/blog/go-vs-java-decoding-billions-of-integer...
[2] http://stackoverflow.com/questions/20875341/why-golang-is-sl...
[3] http://www.reddit.com/r/golang/comments/2r1ybd/speed_of_go_c...
[4] http://docs.oracle.com/javase/7/docs/api/java/util/concurren...
I happen to agree with you whole heartedly, if you spend enough time here though you'll see the inevitable comment about how anything made in php is worthless insecure garbage and anyone who spends their time developing a php application are amateurs at best.
This isn't really a comment at you, just wanting to point out how much that convention is challenged.
"without even using any of the HBaseGiraphFlumeCrunchPigHiveMahoutSolrSparkElasticsearch (or any other of the Apache chaos) mess yet."
That sounds like the state of a lot of docker images.
In general, my base images are often debian:wheezy, ubuntu:trusty or alpine:latest ... From here, a number of times I've tracked down the dockerfiles (usually in github) for a given image... for the most part, if the image is a default image, I've got a fair amount of trust in that (the build system is pretty sane in that regard)... though some bits aren't always as straight forward.
I learned a lot just from reading/tracing through the dockerfiles for iojs and mono ... What is interesting is often the dockerfile simply adds a repository, and installs package X using the base os's package manager. I'm not certain it's nearly as big of a problem as people make it out to be (with exception to hadoop/java projects, which tend to be far more complicated than they should be).
golang's onbuild containers are really interesting. I've also been playing with building in one node container with build tools, then deploying the resulting node_modules + app into another more barebones container base.
Anyway, yes, you can make your own base images. But, images `should` be light enough where you can build them each iteration. I've done dev stacks where literally each `save/commit/run of a test` built the docker container from the dockerfile in the background! With the caching docker does it really doesn't add any overhead to the process.
> What is interesting is often the dockerfile simply adds a repository, and installs package X using the base os's package manager.
Yup! Pretty much. Other than some config stuff for very specific use cases (VPN, whatever.)