> Actually it makes sense, if you manually force the code to use only 4 threads it will take more time to finish that in the case when you let the compiler split the work for you.
While I think this is correct, I think your reasoning is off. And as far as I know, the compiler does not do anything clever here (unlike in some smarter, less mainstream languages) and you're just launching a huge number of threads.
If you use only 4 threads in a similar manner in a loop, the problem is that the background threads finish and the CPU is idle while the serial loop that creates the threads is occupied by creating and launching the threads and there are not enough threads to keep the cpu busy. (edit: this is more likely about waiting for I/O, see below).
If you're truly CPU bound, you will not gain from having a lot more threads than CPU cores (* ~1.5 for hyperthreads). If your threads are waiting for I/O, adding more will cause the OS to switch to an available thread but you gain more by using asynchronous I/O functions (like aio and epoll or kqueue, which you won't find an abstraction for in C++ stdlib). The way the code in the blog works is actually spending a whole lot of good cpu cycles for doing context switching from one thread to another and wasting a lot of memory for maintaining threads that wait for disk I/O.
What you can't see from the tutorial source is what the worker actually does (make_perlin_noise), I assume that function also does the writing to disk which would explain why you benefit from launching a huge number of threads. On the other hand, your memory use is completely unacceptable with this solution.
You should try this: split the number of images evenly and make each thread compute a number of images. Launch roughly as many threads as there are cpu cores and do not make the threads wait for disk I/O. You should see performance go up and memory use go down.