The original implementation gave out a handle that waited for all the threads to end when it was dropped (i.e. went out of scope). That would prevent premature cleanup. But it was possible to leak the handle so that it would never be dropped even at the end of the block, and then it wouldn't wait for the threads.
The new implementation calls a closure and waits for the threads when that closure returns. Unlike the destructor here's no way to stop that code from running when it should.