But the overhead of a context switch for a thread and process is very similar. The main difference is whether memory is shared by default.
But the overhead of a context switch for a thread and process is very similar. The main difference is whether memory is shared by default.
Except that the threads share the exact same virtual address space, and processes do not, which makes the thread context switch faster.
And that is to say nothing about the setup and teardown process, which for a process involves copy-on-demand'ing the entire memory, but for a thread merely setting up its own stack.
That's what I said. But it's really not much. I'm afraid we will need numbers now to continue the conversation. If I measured would you be open to changing your opinion? Or are you committed to this topic, so that it would have no bearing?
Spin up and tear down a million pthreads in C, and see how long that takes and how much memory it takes. Then spin up and tear down a million processes in C and see your computer grind to a halt until you kill the process that is starting the processes, if you can even get your computer to do that without power-cycling.
It's <50 lines of code for each, so I'm eagerly waiting for your response!
Notably, my confidence here comes from the fact that I don't generally get into performance arguments without having actually tested what I'm saying. I've written this code before--it's what I do whenever I'm checking out a new programming language or threading library. Given the complexity of modern computers, nobody really can predict how a program will behave without testing it (except maybe in assembly) there's just too many variables. So you should stop doing that.
If you decide to try the same thing in Java (the other language mentioned), probably drop the number of threads/processes down to 100,000, since Java's lightweight threads aren't quite as efficient. 100,000 processes will probably still be enough to crash your computer.
I'm sure you can find some language/library which implements threads particularly inefficiently, so let's stick to pthreads/C and avoid that straw man.
EDIT: Here ya go, I had ChatGPT write this one for ya:
#include <stdio.h>
#include <pthread.h>
#include <unistd.h>
void* threadFunction(void* arg) {
// Sleep for 10 seconds
sleep(10);
pthread_exit(NULL);
}
int main() {
int numThreads = 1000000;
pthread_t threads[numThreads];
// Create threads
for (int i = 0; i < numThreads; i++) {
int result = pthread_create(&threads[i], NULL, threadFunction, NULL);
if (result != 0) {
printf("Failed to create thread %d\n", i);
return 1;
}
}
// Join threads
for (int i = 0; i < numThreads; i++) {
int result = pthread_join(threads[i], NULL);
if (result != 0) {
printf("Failed to join thread %d\n", i);
return 1;
}
}
return 0;
}
And... #include <stdio.h>
#include <sys/types.h>
#include <sys/wait.h>
#include <unistd.h>
int main() {
int numProcesses = 1000000;
pid_t childPID;
// Create processes
for (int i = 0; i < numProcesses; i++) {
childPID = fork();
if (childPID < 0) {
printf("Failed to create process %d\n", i);
return 1;
} else if (childPID == 0) {
// Child process
sleep(10);
return 0;
}
}
// Wait for all child processes to finish
int status;
pid_t pid;
while ((pid = wait(&status)) > 0);
return 0;
}
It looks like the latter just crashes the program without taking down my whole machine now, which is an improvement over the last time I tried this with processes. $ gcc threads.c
$ time ./a.out
real 0m10.097s
user 0m0.035s
sys 0m0.239s
$ gcc process.c
$ time ./a.out
real 0m10.168s
user 0m0.579s
sys 0m0.347s
Were you running on something besides linux, or not natively? Is this something that degrades with the large numbers?Also, spawning is not context switching. That's the overhead that matters. But according to your own test, spawning in reasonable numbers will be about the same.
Running on MacOS, but I've run this in Linux.
> Is this something that degrades with the large numbers?
The concern here is memory--once you push into pagefile your processes will become extremely slow.
> Also, spawning is not context switching.
Thank you obviousman.
> That's the overhead that matters.
Why do you think you know every use case? You don't. There are tons of use cases where having to be concerned about creating and destroying threads places a large burden on the developer.
> But according to your own test, spawning in reasonable numbers will be about the same.
You didn't run my test.
Running 3000 processes is a few orders of magnitude less than running 1000000, and you don't get to determine what "reasonable" is for every application that exists.
> Running 3000 processes is a few orders of magnitude less than running 1000000
I didn't decide on the limit, my OS did. So they must think it's unreasonable.
> Running on MacOS, but I've run this in Linux.
Well your test works fine on linux. MacOS is not designed to run large multi-process server loads. Linux has specifically optimized forking and context switching for processes.
> Why do you think you know every use case?
Why do you think you can't handle most cases by using processes? The goal isn't for a tool to handle every use case, it's to handle a specified set of use cases well.
Python itself doesn't work for every use case.
Looks like we are safe to ignore overhead of launching processes for programs with fewer than 3000 threads.
> Thank you obviousman.
I wanted to compare context switch, you decided to measure something else. I am pleasantly surprised anyway.
> Well your test works fine on linux. MacOS is not designed to run large multi-process server loads. Linux has specifically optimized forking and context switching for processes.
Oh JFC, stop wildly speculating and pretending it's the truth. The test was a million threads/processes, and you ran 3000. By your own description you didn't run the test, period. I just ran it on Debian on a VPS and, you know, it did exactly what I said it was going to do, because gosh, this isn't the first time I've run this test on Linux.
And you seriously want to attribute this to Linux having optimized forking and context switching as if you have any idea what that means? Please do tell, which optimizations did they apply that somehow they've hidden from BSD and Apple?
Given your propensity to make things up when you don't know something, I'm beginning to think whatever problem you ran into with a million threads was fixable and you just didn't know how to, so you made up a new straw man test to try and win an argument. Have you measured the memory usage at 3000 yet, or are you still ignoring any part of reality that isn't convenient for your argument?
> Why do you think you can't handle most cases by using processes?
Where did I say that? Unlike you, I don't make generalizations about "most use cases" because I don't pretend to know what everyone in the world is doing.
> The goal isn't for a tool to handle every use case, it's to handle a specified set of use cases well.
What do you mean, "the goal"? You speak for every possible goal anyone using Python could possibly have now?
> I wanted to compare context switch, you decided to measure something else.
Then do it! I'd be interested to see the results, and even more interested to see how you tested it.
In any case, you can't just ignore tests you don't want to do or apparently aren't capable of. Being able to run a lot of threads can be extremely useful for networking applications, which is why I care. You don't get to decide my use case is "unreasonable" because you apparently can't compile a program that does a stripped down version of it.
That is to say, even if context switching is faster in processes than in threads, that doesn't mean there's no use to threads, because there are use cases where spinning up and tearing down is more common than context switching.
> I am pleasantly surprised anyway.
Ignorance is bliss, I suppose.
$ gcc threads.c
$ time ./a.out
real 0m10.097s
user 0m0.035s
sys 0m0.239s
$ gcc process.c
$ time ./a.out
real 0m10.168s
user 0m0.579s
sys 0m0.347s
You were pleasantly surprised by your own test showing an order of magnitude speed gain in user space by using threads?