Objective-C Blocks vs. C++0x Lambdas: Fight
mikeash.com
mikeash.com
> [...]
> for_each is a template function, which means that it gets specialized for this particular type. This makes it an excellent candidate for inlining, and it's likely that an optimizing compiler will end up generating code for the above which would be just as good as the equivalent for loop.
This is somewhat unfair, as it seems to reflect the common misconception that (e.g.) C++ sort is faster than C qsort because it uses templates, when in fact qsort would be just as fast if its implementation were written in the .h file, as C++ sort's is. Compilers are perfectly capable of inlining calls to function pointers.
Calls to blocks should be able to be inlined too, but I guess they're still a new feature; I did a quick test, and it seems that gcc cannot inline them, but llvm-gcc and clang can.
In this case, the article suggests:
[array do:^(id obj) {
NSLog(@"Obj is %@", obj);
}];
Objective-C message calls are never inlined since they are dynamic, but if you write something like: static void for_each(NSArray *array, void (^callback)(id)) {
for(id obj in array) {
callback(obj);
}
}
[...]
for_each(array, ^(id obj){ NSLog(@"%@", obj); });
llvm-gcc and clang are able to generate code equivalent to doing the for loop directly.Does this happen even if the function that calls the callable (in this case sort) is not inlined? This would mean that the compiler generates a specialized function, say sort_{somerandomnumber} with the user-defined callable inlined into, and this seems kind of unlikely to me (while with templates the compiler is forced to do so).
In both your example the body of the caller (array do and for_each) is so small that I assume it is inlined.
Personally I prefer the C++0x closures precisely because of the reference/value capturing distinction.
Also, the requirement to add an empty parameter list whenever there's something between [] and {} amounts to extra complexity. This seems to be a purely syntactical dismbiguation thing as neither the mutable qualifier nor the return type should be related to an empty parameter list.
We have to realise what complexity is. Part of it is having to think of B whenver we do A even though we do not wish to say anything about B at all.
I just wish the actual passive-aggressive fight between the FSF and Apple would resolve and blocks could finally make it into upstream GCC C compiler.
XCode 3.2 defaults to GCC 4.2 with LLVM optionally available and IIRC XCode 4 too (can't check as I downgraded for various reasons).
The blocks patchset against pure GCC exists, and the problem mostly lies in upstream GCC refusing patches whose copyright has not been assigned/transferred to the FSF (see https://lwn.net/Articles/405417/). The rationale is that a critical component such as GCC should not be at the mercy of multiple (possibly hundreds) conflicting parties and easing a possible relicense process to ensure its protection.
While I understand the rationale behind this, my opinion is that it feels bureaucratic to the point of hampering notable innovative contributions while favoring local forks which will inevitably end up dying, as maintaining a fork (whatever the patchset size) against the march of a behemoth like GCC is essentially hopeless.
Xcode 4 defaults to llvm-gcc, not gcc-4.2.
I understand why the FSF wants copyright assignment, but it makes the process a lot longer and more complicated.
What this means is that whenever you pass off a lambda to a function, and that lambda captures a variable, you need to be aware of whether or not the function keeps a reference to the lambda someplace, so that you'll know if it's safe to capture by reference or not. If you capture by reference but the function keeps a reference to the lambda, and it gets called after the captured variable has gone out of scope, you'll get into some nasty trouble.
And this is why I don't like the C++ spec as it stands. It requires a kind of global knowledge to work with correctly in local contexts, in such a way that the compiler can't really help you either (it may have been possible to annotate types to indicate closure lifetime, but it would be painful without more powerful type inference than C++ has).
Ultimately Garbage Collection would help there, but I am not sure we are going to see that in C++ anytime soon.
When I designed and implemented the same feature, anonymous methods in Delphi, I used reference counting to keep alive a heap-allocated activation record containing all captured variables. This works well for most scenarios; it can get into knots in more obscure situations where you have recursive lambdas that call themselves via a captured variable, but those are usually pretty rare.
You're right that GC is a help. The biggest thing GC gives you is freedom from having to worry about who controls the lifetime of parameters and function return values, in most cases. In the presence of GC, you can get more clever about your algorithms and data structures; you can cache and memoize, without paranoid concern for things disappearing behind your back. Consider a querying API that takes in closures for sorting and selection functions; I can see it building up temporary results and caching them, or streaming results in a multi-threaded fashion, but it can only do that if it can reliably hang on to closures after the select/sort/etc. function has returned.
I'm not really sure how GC works but using normal memory allocation when you can and reference counted pointers when somebody else is responsible for deallocating objects seems to be adequate and is still high performance...
Ultimately the C++ problems have to be resolved by conventions and idioms, just like we've all learned to be careful when passing a pointer to a local variable.
You could get around this problem with spaghetti stacks or something, but you'd need to find a way to free the activation records that are floating around after their enclosing scopes have expired -- enter garbage collection which you do NOT want to require for C++.
That's the problem with Lisp, it's almost all-or-nothing. If you want to correctly include some of the benefits of the Lisp execution model -- like lambdas -- you need to accept the whole thing, lock, stock, and barrel. Including heap-allocated local vars and the garbage collector. (And yes -- Python, Ruby, Haskell, and Standard ML have a "Lisp execution model" in this sense.)
So we get compromises and hulking abominations like C++0x lambdas or -- worse yet -- their predecessor, Boost Lambdas.
tl;dr: Upward funargs are hard, let's go shopping.
int f() {
int x = 42;
return dupclosure(^(){return x+10}};
}
int g()
{
int (^fn)();
fn = f();
fn();
freeclosure(fn);
}
It's not quite as pretty, but it's workable, and in line with a C-family language's semantics. void* closureheap(void (^)())
size_t closuresize(void(^)())
Or, well, anything else that allows you to separate the step of allocating memory from it's initialization. (There are probably more representation-independent APIs that would be better, but this was just off the top of my head)And, yes, mutable captured variables will not be shared across multiple duplications. Somewhat unconventional, but given the mental model of copying closures, I don't think it's a surprising behavior.
Objective-C blocks automatically copy those variables to the heap. (Which is not to say they're not a bunch of compromises.)
tldr: Lambdas provide more flexibility but are more complicated to use. Blocks integrate with objective-C better and are simpler to use.