In what cases is Java faster than C++?
quora.com
quora.com
If you eschew all this and write strongly optimized code, I think the single biggest spot where Java is unquestionably slower than C/C++ is in array accesses. In Java, to set a slot in a two-dimensional array Java must first test to see if the array is non-null, then test that the X bounds are correct, then test to see if the appropriate Y subarray is non-null, then test to see if the Y bounds are correct, then finally set the value.
In C the compiler does a multiply and an add and sets the slot.
In case that wasn't clear, he claims not that garbage collection is more efficient, but the saved programmer time can be used in some hand-waving way to improve efficiency. Likewise he wanders off into wooly territory with the claim that very large Java programs can be written more quickly. These would be perfectly fine if he were posting the reasons he thinks Java is better, but in a technical discussion of all-out performance it's just offtopic.
Considering that performance is always a trade-off with price - if you think about price/performance ratio - the ability to write programs quickly is of relevance.
Most of the multi-threading issues go away. This is how fast code is written for the Cell processor in the Sony PS3.
Most malloc/free issues go away. Most data is a value type or has an explicit lifetime (game-lifetime, mission-lifetime, frame-lifetime for the game development case).
More information on data-oriented design: http://gamesfromwithin.com/data-oriented-design
To the extent that most people aren't going to change their programming methodology just to avoid Java and C++ still has its niche usages - I don't think the article is naive in any way.
If you care about performance and scalability then you write stateless code and use RESTful interfaces. You also choose to write data-oriented code rather than object-oriented code.
Data-oriented code is not possible in Java because you can't create complex value types and you can't control when and where memory gets allocated and deallocated.
There are obviously techniques in data-oriented code that aren't possible in Java, but a lot of the key insight is applicable in just about every language. Structures-of-arrays, defining the data in objects based on usage patterns instead of responsibilities and "model-the-world" categorisation...
Java doesn't have a `sizeof` operator, and objects probably don't tend to store their object member variables by value, and it's not always obvious which function calls cost how much... Problems, to be sure, but if you really want data-oriented code you can usually contort yourself far enough to get it.
quux_t *foo(int i) {
/* Trick to prevent compiler from inlining */
if(i == 0) {
quux_t *bar = (quux_t *) malloc(sizeof(quux_t));
return bar;
} else {
return foo(0);
}
}
Seriously? GCC removes this obfuscation at -O2, Clang does it at -O1. Check the disassembly before making stupid claims like this.Though these flags are gcc specific (works in clang as well). I found it to be tremendously useful in some cases.
I'm waiting for technology that makes this trade-off obsolete. One should be able to transition from fast prototyping to solid and optimal production code -- incrementally and without great pain. We as an industry are just about ready for this.
And even if you contrive your app to use memory very very carefully, with a shared runtime with 100 other apps (e.g. on a server) not everybody is as nice and GC still runs/stalls the system.
You can easily get generational GC to be an order of magnitude slower than it should be by changing just one or two settings.
There is some way to get the GC to NOT ever run?
Yes, this actually came up at Smalltalk Solutions many years ago. One presenter was using Squeak as an advanced debugger at a company producing a FPS game. If all your functionality happens between frames, you can rig it so you never GC. You just throw away all of your memory outside of "perm" space every time.
With VisualWorks Smalltalk, you can change the settings so that the bulk of your GC work happens using incremental GC. It's not uncommon to get to the point where GC never takes up more than a few milliseconds. That's plenty good for most people. Admittedly that's not so good if your "light" request load is well over 1000 transactions per server per second.
And even if you contrive your app to use memory very very carefully, with a shared runtime with 100 other apps (e.g. on a server) not everybody is as nice and GC still runs/stalls the system.
When you need enough virtual hosts to be running 100 server processes -- that's likely when you need to be transitioning out of "rapid prototyping" mode and onto processes with just a little more rigor.
What I'm advocating is a language where both the rapid prototyping and running optimized mature code efficiently is possible. Not only possible, but easy to transition between.
(Erm... I think. Need to read more...)
But at that time, I wrote a memory manager which had malloc and free calls which were a magnitude faster, at least ten times.
When you need it to be fast, it can be fast. If you don't want it to be fast, or you have to resort to tricks like preventing the compiler from inlining (wtf?), then you really are doing it wrong.
Or do you mean 50 to 100 microseconds? That's still horrible, but I could believe it.
http://stackoverflow.com/questions/145110/c-performance-vs-j...
It is basically the great programming language shootout at http://shootout.alioth.debian.org/ redone with different GCC optimization levels.
It is clearly visible where Java Hotspot shines, but in most cases C++ wins by a factor of two. It also mentions a technique how you can do profiling analysis with GCC that makes similar optimizations than Hotspot does.
Wrong.
It is basically the old Doug Bagley programs which were replaced 3 years before the zi.fi/shootout article was posted.
A little history - http://c2.com/cgi/wiki?GreatComputerLanguageShootout
In this new decade -
http://shootout.alioth.debian.org/u64q/java.php#faster-progr...
That time C optimizers were getting better and better up to the point they knew "tricks", where you had to be a very, very good assembler programmer to compete. History seems to repeat.
>> "Value Types, such as a 'Complex' type require a full object in Java. This has both code speed and memory overheads."
I have programmed in C# and C++ (never Java), but I found the above to be the key issue why my programs would run significantly slower with C#. Here's the C# language discussion thread where I posted details about my issue:
http://social.msdn.microsoft.com/Forums/en/csharplanguage/th...
Yes.
Yes.
The benchmark code also re-parses the integer command-line argument on every iteration of the loop.
This benchmark is meaningless.
And Java objects in real life tend to be big and not very local, from what I've seen.
I use incgc with very good results compared to the defaults. The defaults usually result in pauses (Not good for something like mibbit), also defaults simply can't keep up with object churn.
So, anecdotally, incgc works wonderfully.
By using malloc in the C implementation, you force it to be on the heap.
Because if it really would use something like malloc, I don't really see how it should be faster than the C version.
What you cannot do in C/C++ that you can in Java is respond to program inefficiencies at runtime, with runtime knowledge.
For everything else, there is Profile Guided Optimization.
There are JIT libraries for C++ for building applications which dynamically compile stuff, I'm not aware of any that dynamically recompile the C++ though.
Seriously though this is misguided. PGO is clearly a step in the right direction but its simply deferred static optimization. Either you optimize in C at the compile time, or with PGO by running a few statistically representative executions, but these are not dynamic.
The only way to make truly dynamic optimizations is by having a runtime environment and code that is interpreted.
The argument that its not often necessary is a different argument and with only a few moments thought can be seen to be untrue. Take for example and strstr like operation that you naively implement using by walking through the source string. Now let's say you do this a lot, all is good because the input is short, now you receive (in a parser for example) something much much bigger. A JIT has privy information and can make determinations that building a index on the source string to make repeated finds faster is better, C does not, even with PGO because you may never have run a test that encounters this situation.
Why not use clang compiling to LLVM bytecode and JIT from that?