Comparing Go and Java, Part 2 – Performance
boundary.com
boundary.com
At the mean, at concurrency 50, we have things like 37.5ms vs. 12ms. To a user, that is "instant" vs. "instant". Sure, if we stack a bunch up on a page, maybe we'll start to care, but..
Much more directly, in real-world scaling scenarios, it matters far more if the 98% number is 2s or 248ms than the mean is 37ms. A 2s latency means 2 out of every 100 requests (and likely > 2% of page views), the user experience will suck. And, all it takes is 6 requests to ensure that 25% of users experience this (your homepage alone probably requires 6 requests).
I'm not just raising this to be pedantic--I've seen plenty of systems that have been tuned to do very well on things like req/s or even mean latency but do very poorly for 1 in 100 or 1 in 1000 (vs. being a very fair scheduler at the cost of overall throughput). We reward these systems by measuring and praising the wrong thing. (e.g. mongodb vs. riak)
edit: (btw, not intended as a specific defense of either go or java wrt the article, just a general statement about benchmarking these systems)
Check it out: http://dropwizard.codahale.com/
I'd rather build something with Java EE6 if I had the choice now, or ASP.Net MVC+WCF+NHibernate if I had the choice of platform.
Don't, its a nightmare.
2) JSF + WELD have holes between there specs. CDI can manage JSF session and request scoped beans and you can use conversation scoped beans in JSF. However they never got around to figuring out view scoped beans (which IMO is one of the most useful things in JSF). As a result you have to use some third party to make them work together. There are a number of other spec holes but I'll use this one as an example.
This means using SEAM 3 or CODI CDI extensions to fill the gaps. I went with SEAM 3 as it theoretically had a good lineage given that a lot of what was done in SEAM 2 ended up being the EE6 spec. What the SEAM 3 project did however, was to write a half assed version of all their extensions, then summarily announced they were all ditching the project to go work on Apache Deltaspike, which is basically the same extensions with a different name. Currently the SEAM 3 extensions are in various states of broken depending on the extension and your use case, and of course Delta Spike isn't ready yet (still in Apache Incubator). I've removed most of SEAM 3 from my app at this point, the remaining module is SEAM Faces which is unfortunately plugging the hole in JSF View Scope working with WELD. Also unfortunate is that there is and has been a bug in SEAM Faces for months now that causes the data JSF stored in the session to be unserializable. So I'm stuck with sticky sessions.
SEAM modules also fill in the hole with Hibernate's stupid session management, i.e. opening a transaction in the RENDER_RESPONSE phase so lazy loading works.
Several times now I've run into bugs, gone to research them and found they weren't fixed because people from the involved modules were arguing over the EE6 spec.
I could probably yap for awhile here but It may be easier to say this, I started my current product last April(2011), it started as a "pure" EE6 app. I ran into so many EE6 bugs and just pure stupidity in places that I've been moving away from it ASAP. I've replaced Hibernate with Ebean, ditched 90% of SEAM, setup an API with Jersey/Jackson and we're now pushing all front end stuff to Rails or Backbone apps which just talk to the API.
I'd never write another webapp with EE6. Jetty/Jersey/Jackson is great for web services, especially with groovy. If you need to handle the front end tasks, use play or grails.
Play or Grales might be just as bad as EE6. Do you also have similarly intensive experience with Play and/or Grales since April 2011 that you can compare with?
That said to get to your question, I don't have experience with either at the same scale as the EE6 app. I have fiddled with both some to get a feel as I was deciding what the path away from the EE6 stuff I would take. I have ditched hibernate for Ebean (Ebean is the ORM that comes with Play). I've also started writing most of the controllers in Groovy, which comes from messing with grails. Both these changes have been awesome and significantly simplified the project.
The entire point of Arquillian is to test in a live container, which means using it for things that will interact with the container services. They love to put up examples of stuffing 3 classes into a war and testing it, to which I say, why? It makes sense if your testing a CDI extension (who incidentally seem to be the only people using Arquillian). If I wanted to test 3 classes from my app I'd mock the injects with mockito. Its far faster and easier to use mockito to inject instances, especially since mockito can instrument those injections allowing you to assert method call information.
Where this all falls apart is where Arquillian should shine: integration tests. e.g. put up a functional JSF controller/view and fire requests at it with JSFUnit. You now have to package up enough stuff into your war to get a functional JSF environment running. Arquillian doesn't help you at all to figure out what dependencies you need to include to just get JSF running. As a result you end up playing games trying to get the WAR to include what it needs to run which can be much harder than it sounds due to the way much of the EE6 stack is layered and intertwined. You end up having to include a ludicrous amount of stuff to get JSF running. Or you just say 'fuck it' and tell Arquillian to include your entire pom and get a huge deploy. On top of this is managing your own dependencies. Is your view modularized and using 5 different sub views and supporting controllers (and associated helpers?). Its all up to you to track and manage this. Its like being thrown back into a world without maven for every single test case you try to setup. I found I spent more time trying to figure out what I needed to deploy than I did writing tests.
Also your completely screwed if you try to do a service layer down to database integration test and your using hibernate. Waiting for hibernate to start up on every test class is brutal. Getting rid of hibernate makes testing and many other things so much easier. If you stop to think about how much crap exists in the stack just to deal with hibernates session lifecycle and transaction requirements its amazing. In my current app I have about ~65 entity and ~70 tables, switching from Hibernate to Ebean and ditching the libraries I no longer needed to deal with hibernates session management dropped my final war size by ~15MB (40%), cut startup time in half, and allowed me to remove over a thousand lines of code.
I think I will avoid as well then.
With that said, I am a big fan of Spring and I don't consider it "intrusive" as some other commenters do. But if you aren't interested in using it as a DI container then you may find it easier to just get closer to the other libraries used (i.e. Jersey etc) and avoid any helper code.
http://static.springsource.org/spring/docs/2.5.x/api/org/spr...
Does that answer your question?
Here's my personal experience. Simple go code is actually comparable to c, even for non-io bound tasks. I was very surprised when OpenSSL's AES implementation and go's AES implementations performed similarly in my microbenchmarks. The jvm actually performs very well (go figure that over a decade of optmization in running enterprise workloads would result in a fast runtime). I've inspected the generated assembly in a go code and compared it with that from the equivalent c code and there's no doubt go isn't the most efficient language ever created.
Use go if you want something that's (pretty darn) fast, productive, and has a standard library written by some of the most respected names in the field. Don't use go if you need the most mature and performant runtime and libraries. I have no doubt that they will get there eventually.
No, is not worth reading, is misleading at best and has been throughly debunked:
http://blog.golang.org/2011/06/profiling-go-programs.html
Not to mention it used an ancient version of Go, even Go 1 is dramatically faster than that, and since Go 1 there have been even more dramatic performance improvements, but the main issue is that the guy that wrote the benchmarks really had no idea what he was doing (there were similar criticisms from outside the Go communities about the quality of the benchmark).
Most folks would assume a paper coming out of google involving go has some go experts involved. Clearly this is not the case.
Since none of the implementations were IO bound (unsaturated link), I'd bet on memory management as the deciding factor (it's hard to assume that an authentication service will be CPU bound).
Then it's not difficult to see why Java won. When it comes to garbage collectors, JVM has the best (production) implementation of GC bar none.
Go runtime has a lot of catching up to do.
http://www.powershow.com/view/14655a-MGM5M/Ulterior_Referenc...
Java version doesn't do type conversion: https://github.com/collinvandyck/go-and-java/blob/master/jav...
rs.getString("id")
rs.getBoolean("admin")
Go version does conversion at runtime, as row.Scan() accepts interface values.Also, the usage of sql interface is not optimal (at least with regards to code length) -- since the query is for one row, why not use QueryRow instead of Query?
Authorization header decoding is strange -- it creates new base64 decoder, then string reader, then decodes. Why not just use base64.StdEncoding.DecodeString() and get rid of a few lines of code and a few allocations?
https://github.com/collinvandyck/go-and-java/blob/master/go/...
Similarly, JSON marshals into a newly allocated slice instead of creating a decoder and then marshaling directly into ResponseWriter.
https://github.com/collinvandyck/go-and-java/blob/master/go/...
Constants for HTTP error codes, that are declared in net/http, for some reason are redeclared
https://github.com/collinvandyck/go-and-java/blob/master/go/...
- - -
I sometimes wonder why various benchmarks include colorful charts, but fail to include a few lines from a simple run of profiler. It's so easy to do, and yet nobody bothers to learn where the cycles are actually spent!
In go you can happily ignore an exception by assigning it to _. That's probably a bad idea for a larger piece of code, but for little scripts, go for it.
I may be already permanently damaged by go. Every time a function can return an exception I have to stop and ask myself: "self, what should you do if this happens?". I'm starting to think it results in better code. Granted, not every shell script/utility benefits from this level of introspection, but that's what python's for I guess.
p.s. Anyone who's had to deal with checked exceptions in java puts up with a crazy about of boiler plate too.
p.p.s. Disclaimer I have been professionally employed writing all languages mentioned above, so hopefully I'm relatively unbiased.
Checked exceptions are indeed obnoxious and a major language design failure, which is why pretty much every modern language since just has plain old (non-checked) exceptions. And even in Java, you can work around the brain damage by wrapping checked exceptions with runtime equivalents in API facades. Exceptions are still incredibly useful, and lack thereof is my biggest complaint about Go.
I'm also disappointed by the convention of capitalizing/lowercasing names to export them or not. Realize that you want to export an existing private method? What's that, your IDE doesn't support refactoring? Get typing, you have a lot of method calls to update.
I think it's fine, personally. It's not overly obnoxious, it gives shape to the code, it avoids the redundancy of an explicit export list (although that also means it's harder to see at a glance what's exported from a module I guess) and it makes sense within Go's habit of mandating formatting, there's no reason not to leverage this mandate.
Coding conventions on steroids, if you will.
Go uses a well-understood mechanism: return. Control reverts to the caller. It's simple, which was an explicit design goal of Go.
IMHO, if you have a choice between exceptions or not, they just aren't worth the value they deliver. As Josh Bloch says, use them for exceptional circumstances only, to indicate truly exceptional circumstances, such as catastrophic errors.
Re: refactoring: http://golang.org/cmd/gofmt/. Check out the -r option.
In practice, is this really a huge deal? Modulo go fmt, what editor doesn't support multi-file S&R with regex? Isn't this what a compiler is for? All in all, this sounds like bikeshedding about syntax. We're all entitled to our opinions but it's awfully hard to say anything interesting about syntax which has not already been said a bajillion times.
Is that on kickstarter? Put me down for 10!
Why? There's nothing necessarily wrong about letting it crash, and letting a layer above report the error cleanly. Hell, in Erlang the usage is even to let an other process entirely handle the error.
In fact, my opinion would be the complete opposite of yours: checking every single return value (if only to return it to the caller unaltered) is fine for little script, but it's a bad pattern to need for large pieces of code, it's verbose, redundant and unhelpful.
Then do that. You don't have to handle errors, you can ignore them just like you would ignore an exception. The difference is that with an error return value, you are explicitly choosing to ignore it. With exceptions, it is easy to accidently ignore it when you didn't want to.
>but it's a bad pattern to need for large pieces of code, it's verbose, redundant and unhelpful.
Which is an argument for better error handling, not an argument for exceptions. If go had Maybe and Either, there would be no problem.
No, if I ignore an exception it bubbles up the stack and will either stop the program or find somebody handling it. If I ignore a return value, the program gets into a completely undefined state and will crash later in a completely different place.
Unless there's a way for go to do the same thing as the Erlang pattern:
{ok, Value} = call(SomeArg, SomeOtherArg).
is there?No, that is the normal way I do it in Erlang (hence the note that this is an Erlang pattern), where there are exceptions (and nobody says there aren't) but most functions tend not to use it and to return tagged tuples: `{ok, Value}` if the call succeeded (or just `ok` if there's no value to return) or `{error, Reason}` if the call failed. Note: lowercase words in Erlang are atoms, you can think of them as interned strings. Words which start with a capital are "variables" (which can't vary, but close enough).
Now the caller can unpack the result:
case some_call() of
{ok, Value} -> %% code to execute if the call succeeded;
{error, Reason} -> %% code to execute of the call failed
end
this uses pattern matching (on the value being a tuple and having the right atom as its first element) to dispatch each case to the right branch.But in this sub-thread, we don't want to ignore the error. In Erlang, "ignore the error" is written:
{ok, Value} = some_call()
this doesn't really ignore the error (and let the function keep running), it asserts that the result of some_call() matches the tuple `{ok, Value}` and faults if that's not correct. The equivalent Go code is what is used in TFA, namely: result, err := SomeCall()
if err != nil {
panic(err)
}
and is also equivalent to not catching the exception in Java: it does not let the code keep running.And my question was thus: is there a way (shorter than the one used in TFAA) to do this, not handle the error but have the error prevent the code from running?
> which you can ignore by either not checking it, or just outright assigning it to _.
No, that leaves the code running in an unknown and corrupted state, I don't consider this acceptable.
It is also the normal way you do it in go. Read the examples, that's exactly why go has multiple return values.
That's what you get when you refuse to use them anyway.
Go has exceptions, there's just a dogma about never using them.
A good rule of thumb is to use panics when it's a programmer error indicating a bug.
Panic() is for truly irrecoverable exceptional situations where you do not expect the caller to catch it.
Something like a 'die on error' option might be useful for trivial scripts and applications.
I sometimes do this, and then remove the fail() function to force myself to properly handle the errors. (Still, often is best to handle the errors as soon as you write the code anyway).
The key is that unlike with exceptions, the fact that you are ignoring the error is explicitly stated in the code, is not something that magically might happen.
And most importantly, errors are part of the documented API, with exceptions it is rarely documented what exceptions a function might throw, much less what exceptions the functions called by that function might throw.
Yes, Go error handling is a bit verbose, but that is a sign of how much better it is than exceptions, without falling into the 'checked exceptions' insanity.
Much as you could write this in Java:
catch (Throwable t) {
throw new RuntimeException("oh noes"); // or whatever
}
You could write this in Go: if err != nil {
panic("oh noes")
}They are are not the same.
I've just started playing around with Go, I'm quite liking it. It's a different approach to things, which is always fun.
Knowing this, someone's already done it---anyone aware of such a thing?
These are not currently available on the JVM (although this may change as of JDK 8).
http://www.atego.com/products/aonix-perc/
http://www.excelsior-usa.com/jet.html
http://researcher.watson.ibm.com/researcher/view_project.php...
None of these are free (or "open source" in the OpenJDK sense -- e.g., where most of the system with the exception of a few libraries is open), however, with the exception of the last -- where the first revision was available as part of Jikes RVM.
The only commercial JVM that I do believe has some production server-side deployments -- Oracle's JRocket -- still doesn't support value types and has (in many cases) actually performed worse than HotSpot with large heaps.
So to put it bluntly none of these fit "I'm willing to deploy this to thousands of nodes in production" criteria. To me HotSpot (due to its stability, performance, and support for concurrent GC) VM is more of a reason _to_ use Java/JVM languages -- otherwise I'd much rather use C#, F#, D, Go, or OCaml.
Finally, I should it's quite possible to do manual memory allocation, bypass bounds checking, and to control memory layout using direct byte buffers and Unsafe in HotSpot/OpenJDK. Problem is (and I've mentioned it) is that it incurs serialization costs from copying objects to these byte arrays. Of course may this may be acceptable for many kinds of application, especially if only a _small_ part actually needs this (which I'd imagine describes many -- if not most -- kinds of applications).
But for the scenarios you described, I agree with you.
OTOH if using a proprietary solution ($$ and closed source) was fine, I'd personally go with C# or F#. If you're fine using a closed-source runtime, why not just use a better language (not to mention Microsoft's products are free for startups via BizSpark)?
Because Mono is not a solution when you need high scalable servers running in commercial UNIXes, z/OS and similar systems? These type of solutions always get to use C++ or Java in our projects.
My current project is actually done in C#, but it is a 100% Microsoft stack.
You just have to know where to look for, plus you don't need to throw away mature languages and tooling.
2 - Go's CPM paradigm is not necessarily more performant than the preemptive threading model of Java, and I don't believe that has ever been claimed for CPM. It is claimed that it is "easier" to write concurrent code using the (CPM/Go) message passing paradigm, but in my experience, you either get concurrency (and can do both variants ok) or you don't, in which case any real world concurrent system will hardly be "easy". That said, for the concurrency novice, Go is far less intimidating experience given the language level support for goroutines (fibers), channels, selectors, etc. (think Java NIO ...)
Go is getting a lot faster and pretty quickly.
I don't know if that makes sense, but languages need somewhat large communities to really thrive, in my opinion, and defining yourself into a small niche isn't a good way to get them.
I downloaded patricks fork and got 4300r/s with maxprocs 1 with the go version, out the box.
Recompiling and running with gomaxprocs=100, i got 4400r/s.
mvn deploy and running auth-1.0.jar on JDK 1.7 on my box similarly peaked out at 3100r/s.
It's worth noting though, that I was using patricks httperf bench.sh modification, and it appears that httperf was cpu bound in both cases, with the kernel and postres taking about half a core between them.
Using wrk, by contrast, spun main (the go program) up to 2.5 cores, and 6000r/s. Java under wrk lit all cores for a time, then hit `Exception in thread "async-log-appender-0" java.lang.OutOfMemoryError: Java heap space` `Caused by: ! org.postgresql.util.PSQLException: FATAL: sorry, too many clients already`
A bit more investigation and tuning later, using 10 clients `wrk -c 10 -r 100000 -t 4 -H 'Authorization: basic apikey_value' http://localhost:8080/authenticate` Java was spinning out 10kr/s. Go was limited to 6kr/s. It's worth noting some more details however: Go only ever spun up two postgres forks, whereas java spun up 10. There's scope for optimization there. Go used 8mb of ram, whereas java was sitting on 130mb after a few runs. The Java version, cranking out at 10kr/s was maxing the kernel out on one core, so that's probably approaching the practical limit for single machine tests.
I suspect putting a load balancer in front of a couple of instances of the go program would allow you to totally smash the java performance, given that it lit 2 cores at 6kr/s, and java lit 8 at 10kr/s. The memory usage tradeoff is significant - JVM sitting at 130mb and Go sitting at 8mb. Clearly everyone needs to draw their own conclusions on this, for their own purposes. The JVM solution is carrying a stats server, a ton of other tooling and so on. The Go system has some - pprof was included, but it's limited by comparison. Arguably gdb and so on can actually be of real use in the Go case, but that's also an exercise for the reader.
Interesting any which way you look at it. There are a lot of other interesting side effects (that need working out) in both programs as evidenced by this simple testing. Postgres also needs some tuning if you really want to slam this with anything remotely resembling a real world scale test.
My tests were done on OSX 10.8 12B19 on a 2.3GHz i7 wiht 8GB of DDR3 at 1333MHz, an intel 510 SSD, totally untuned postgres. The machine is a Macbook from whatever year that is.
Here's wrk, for anyone looking for it: https://github.com/wg/wrk