It doesn't need to be this way. There is a Linux kernel patch that a Google engineer gave a presentation about in 2013 [1] that allows for true kernel threads that can be scheduled in userland, eliminating the tradeoff. Unfortunately, the patch never seemed to go anywhere, and it seems that the author is no longer working at Google (mail bounced when I tried to contact him about it). Note also that Windows already has this functionality, known as UMS.
[1]: https://blog.linuxplumbersconf.org/2013/ocw/system/presentat...
http://web.eecs.umich.edu/~mosharaf/Readings/Scheduler-Activ...
EDIT: And that Google work sounds like directed yields, an old idea which i traced back as far as 1996 before getting bored:
https://www.usenix.org/conference/osdi-96/cpu-inheritance-sc...
If you really need task switching you're not really going to beat the kernel by enough to matter. If you only want to switch between callbacks with very little state then yeah green theads work great.
Coroutines look to be a really great way to just get the best of both worlds, and seems to be a generally better model then what go did for most applications.
As an analogy, I could speed up my ability to get shoes on by an order of magnitude, and have it not really matter for my commute.
Not saying that is exactly the same here. Just extending the question on the impact if this.
Of course, that last assertion needs data. And i could be wrong. :(
Languages like Go like to advertise that you easily spawn millions of Green threads (which would fail with OS level threads or bring your OS down) but I've never seen a use case for that.
For realistic applications a work stealing OS thread pool should be faster than spawning lots of green threads. But in the end, performance has to be measured, of course.
What are those two things, in your proposition that we can simply use Linux threads?
Also you're talking about linux, the Go runtime runs on many platforms, therefore is not dependent on platform thread performance.
You should be able to "just" create 1000 threads in "only" 2ms or less.
Those 1000 goroutines eat up as much memory as roughy one OS thread.
Granted, I've somewhat shifted the goalpost, but goroutines are cheaper than OS threads or just about every sensible definition of "cheap".
It's just reserved, not committed until actually used.
> Granted, I've somewhat shifted the goalpost
Shifted? You've strapped them to a rocket and put them in orbit.
The general point still stands. No need to be unpleasant.
Yes, 1000 concurrent threads will take stack size * 1000.
1000 concurrent threads will not take stack size * 1000, they'll take used pages * 1000 (with used pages being at least 1) + the kernel overhead of the thread structures.
That's trivial to check, even in a high-level language (with its own additional overhead) e.g. spawning 1000 threads in Python on OSX takes ~25MB. OSX uses 512k stacks for non-main threads.