You can write non-performant code in any language you'd like, and obscure it in non-trivial ways. That's not the fault of the language.
You can write non-performant code in any language you'd like, and obscure it in non-trivial ways. That's not the fault of the language.
The mental effort you have to make to determine the overhead of an expression by looking at the call site is greater in C++ than it is in C.
That's an undeniable consequence of features like operator overloading, constructors, destructors, virtual functions, etc. It has nothing to do with how well you understand the language. The information just isn't there at the call site for you to see.
The question is a different one. It's whether that fact is compensated for by the power and reduced mental effort these features enable elsewhere. Constructors and destructors make RAII possible. That arguably makes memory management a lot easier than in C.
Isn't this just a specific example of saying "the mental effort you have to make to determine the overhead of an expression by looking at the call site is greater in higher-level languages than in lower-level ones"?
It's true, precisely because you are using higher-level abstractions. I'd be more interested in seeing an example where this wasn't true.
The only good reason to use C++ is because you (1) need a level of abstraction that C doesn't give you, (2) need level of performance that no garbage collected language can give you (even compiled ones), and (3) can't even use C + Lua (or some other scripting language), probably because you need the abstraction and the performance in the same place.
(This is an awfully narrow application domain, compared to what C++ is actually used for)
At that point, the performance characteristics of your programs have become part of the specs. You do not want to hide them under the carpet like you would do with a garbage collected language. Yet C++ does. It wouldn't hurt if it syntactically distinguished initializations, casts, calls by reference and so on.
Or take reference parameters. In C++, if you pass an object to a function you cannot know by looking at the call site if the object is going to be copied or not. The parameter could be defined as a reference, but you have to look at the function definition to find out. In C#, in Java, in Python and in C you know it by looking at the call site. These other languages are all over the place in terms of any high/low level categorization.
If you're writing, let's say, scenegraph traversal, it's really tempting to get it working on the PPU and then port it to run on an SPU. A vtable dereference to main memory will crash your SPU job and it is not always obvious why. Obviously, you can't use static analysis tools to make sure the right instructions are transferred, so you have to do it by hand. A modern PS3 game has on the order of low hundreds f different SPU jobs...not much fun. Even discounting that, SPUs will work best when fed branchless parallelized jobs, I wouldn't be surprised to see an order of magnitude difference between a vtable call on a pointer array and a static call on an object array.
Keeping pointers to objects in a polymorphic array is a games performance anti-pattern because of the dereferencing cost, but it's necessary to call virtual functions...there's a hidden cost right there.
Which loop is faster, by eyeballing, and by how much?
void load(Assets* a) {
for (int j=0; j<m_numAssets; j++) {
loadAsset(a[j]);
m_numLoadedAssets++;
}
}
void load(Assets* a) {
int numLoadedAssets=0;
for (int j=0; j<m_numAssets; j++) {
loadAsset(a[j]);
numLoadedAssets++;
}
m_numLoadedAssets = numLoadedAssets;
}
I've seen the former style run literally 1000 times slower than the latter...that's obvious? I submit in a world of out-of-order processors it is not at all.I don't think that all the author's points are spot on...Koenig lookup is Byzantine but IDEs do a good job of it, ditto source-level reasoning about dispatch. The underlying theme, that C++ is not a good fit for modern game development, shouldn't be so trivially dismissed.
It depends where *this is allocated. In the worst case, it is allocated in main memory while you wanted to stay in the graphic memory, or something.
A naive compiler would then access memory (or the cache) instead of using registers. A Sufficiently Advanced Compiler would guess that calling ++ many times is the same as incrementing in one go, and hoist that out of the loop, but apparently this one is a bit cruder.
Now the same could be said about m_numAssets, but this one isn't written to, so the compiler only have to put a copy in a register, which I guess is a simpler optimization to do.
To answer your question, that particular situation would be better in any language that forces you to explicit the reference to "this" (or "self"). Imagine how we could modify C++:
void member_function() {
int local_variable++;
local_variable++; // This is okay
member_variable++; // That should not be allowed
this->member_variable++; // This should be written instead
}
Applied to the example in the GGP above: void load(Assets* a) {
for (int j=0; j<this->m_numAssets; j++) {
loadAsset(a[j]);
this->m_numLoadedAssets++;
}
}
We see that every non-local access is prefixed by something ("this->" and "a[" here). The heavier syntax suggests a heavier cost, so the programmer will more easily think of hoisting those out of the loop, if possible (either manually or through compiler optimizations). void load(Assets* a) {
for (int j=0; j<m_numAssets; j++) {
loadAsset(a[j]);
}
m_numLoadedAssets += j; // Increment once, out of the loop
}I explained that here: http://news.ycombinator.com/item?id=4540107 I did not get the Load Hit Store problem, but I did sense there was a problem with the repeated read access to memory.
By the way, it looks like your problem is a bit more subtle than the textbook Load Hit Store: you don't read the location you just wrote to. I guess the write access dirties more than one word (at least a cache line, maybe more).
Most of the argument revolves around objects. There are two ways you can run into "hidden" surprises with objects in C++ that the author is pointing out. One involves failing to use the language features given, the other involves writing some seriously odd code.
class C {
public:
/*explicit*/ C( int ); // Adding explicit avoids hidden costs
// This combination (or something like it) is lethal.
C( int, double );
operator bool() const;
};
So for the author's code to have hidden performance concerns, they would need a constructor taking two parameters of the proper type and an automatic conversion operator that converts to a type assignable as their result. Basically, the stars have to align right and you have to knowingly write some dangerous code.Don't get me wrong, you can hide costs pretty easily in C++ and write some atrocious code. But usually this issue is alleviated by following basic best practices. The hard part is knowing what those best practices are.
I do agree that C++ is a complex and hodgepodge language, but I am not sure there are acceptable alternatives yet. It would be great if JVM had native support for value types, i.e., a way to use manual memory management, bypass bounds checks, or call native code without incurring serialization costs.
Yes, I am aware that C# does this (and I like C# and F# a great deal as languages), but a) I am not deploying Windows in production b) Mono is not yet a viable high performance server side platform (it does look promising for desktop and mobile application or for simpler ASP.NET webapps).
C is a beautiful and simple language (one I have strong affinity for) but manually implementing vtables is a bit painful (see a large pure C project for an example of that), nor is there any support for type-safe generic programming (which is an issue I have with Go as well -- even though I fully understand their reasons for omitting generics).
Then again, the reverse isn't true. There are languages (or more precisely their implementations) that just are slow.
Furthermore, good asymptotic complexity doesn't always correlate with faster code for values of n you actually encounter. Linear search is much faster than binary search, because of cache behavior, up to a surprisingly large n on many architectures.
I have had a few experiences of simply porting something rather directly to C++ and seeing 10X speed and memory improvements. The 'hidden costs' in other languages are loads of unavoidable heap allocations and pointer indirection that really don't happen with C++. And C++ compilers do a ton of very smart optimization these days.
With the generic programming constructs and static polymorphism approaches in C++ you can often get really big performance increases over what's practical in C or Java. std::sort in C++, for example, has often been found a few times faster than stdlib.h qsort. It's because of compile time optimizations that languages like C and Java can't do.