How much memory do you need to run 1M concurrent tasks?
pkolaczk.github.io
pkolaczk.github.io
Go and Java with lightweight threads provide a full-blown system, without any limitations. You don't have to "color" your code into async and blocking functions, everything "just works". Go can even interrupt tight inner loops with async pre-emption.
The downside is that Go needs around 2k of stack minimum for each goroutine. Java is similar, but has a lower per-thread constant.
The upside is that Go can be _much_ faster due to these contiguous stacks. With async/await each yield point is basically isomorphic to a segmented stack segment. I did a quick benchmark to show this: https://blog.alex.net/benchmarking-go-and-c-async-await
In contrast, in C# (or any other similar system) async calls are _expensive_ compared to regular function calls.
I'm NOT trying to show that Go is faster than async/await or anything similar. I'm showing that nested async/await calls are incredibly expensive compared to regular nested function calls.
If you want to show that async/await calls are expensive, than you should have shown two code samples of C#, one with async/await, and one without.
Or could have done the same for Go, show one example with goroutines, and one without.
But I think everyone already know that async/await and goroutines has it's costs.
The problem is more that you are comparing Go without goroutines (without it's allocation costs) to a C# example with a poor implementation of async/await.
It's trying to show the overhead of the async machinery, compared to "uncolored" unfunctions.
Even at a million tasks Go is still under 3 GiB of memory. So that is roughly 3KiB of memory overhead per task. That is likely negligible if your tasks are doing anything significant and most people don't need to worry about this number of tasks.
So this comparison shows that in many cases it is worth paying that price to avoid function colouring. But of course there are still some use cases where the price isn't worth it.
You're sure you should use .Result here instead of awaiting it?
To make this simple they introduced async Main[2] a few years ago.
[1]: https://github.com/davidfowl/AspNetCoreDiagnosticScenarios/b...
[2]: https://github.com/dotnet/csharplang/blob/main/proposals/csh...
It would also be interesting to try Native AOT compiling both versions…
https://learn.microsoft.com/en-us/dotnet/core/deploying/nati...
Rust also has to allocate (box) when you need recursive async calls: https://rust-lang.github.io/async-book/07_workarounds/04_rec...
What Rust does, it allows to remove one layer of allocations by using Future to store the contents of the state machine (basically, the stack frame of the async function).
This is not at all that different from Go, except that Go preallocates stack without doing any analysis.
Actually... I can fix that! I can use Go's escape analysis machinery to statically check during the compilation if the stack size can be bounded by a lower number than the default 2kb. This way, I can get it to about ~600 bytes per goroutine. It will also help to speed up the code a bit, by eliding the "morestack" checks.
It can be crunched a bit more, by checking if the goroutine uses timers (~100 bytes) and defers (another ~100 bytes). But this will require some tricks.
You can always pass context.Background, in this metaphor creating a new tree of color.
You can always call "runtime.block_on(async_handle)", in this metaphor also creating a new tree of color.
Say you have a foreach function that calls each function in a list. In async-await contexts you need a separate version of that function which is itself async and calls await.
With context you can pass closures that already have the context applied.
You're talking about doing this:
f1 := func() error { return nil }
ctx_f1 := func(context.Context) error { return nil }
fns := []func() error{f1, func() error { ctx_f1(ctx) }}
Basically, writing an anonymous conversion function.In, say, rust, the equivalent in this analogy would be:
let f1 = || -> Result<(), ()> { Ok(()) };
let async_f1 = async { || -> Result<(), ()> { Ok(()) } };
let fns = vec![f1, runtime.block_on(async_f1)];
That conversion function, 'runtime.block_on', doesn't really seem that different. The analogy still seems to hold.For a more concrete example: let's say you have a generic function that traverses a tree. You want to compare the leaves of two trees without flattening them, by traversing them concurrently with a coroutine [1]. AFAIK in rust you currently need two versions of traverse, one sync one async as you can't neither close over nor abstract over async. In go, where you have stackful coroutines, this works fine, even when closing over Context.
So yes, in some way Context is a color, but it is a first class value, so you can copy it, close over it and abstract over it, while async-ness (i.e. stackless coroutines) are typically second class in most languages and do not easily mesh with the rest of the language.
[1] this is known as the "same fringe problem" and it is the canonical example of turning internal iterators into external ones.
A function without a context can still be used just fine from any code, you just won't be able to cancel it.
It's similar to explicit DI vs implicits.
the function coloring metaphor doesn't make sense since the calling convention is the same nor are there extra function keywords (`async` vs non-async).
However, you raise a different point:
The tasks I'm running don't involve the network, they either succeed or error after an expensive calculation.
This sounds like CPU-bound, not I/O-bound. (Please correct me if I misunderstand.) Can you please confirm if you are using Go or a different language? If Go, I guess it still makes sense, as green threads are preferred over system threads. If not Go, I would be nice to hear more about your specific scenario. HN is a great place to learn about different use cases for a technology.[1] https://docs.racket-lang.org/reference/eval-model.html#%28pa...
It took me a while to figure this out, thanks to articles I came across (such as “what color is your function?”.
There are memory / speed trade-offs, as per usual. If you have a lot of memory, and can keep all the stacks in memory, then go ahead and do that. It will save on all the frivolous construction / destruction of objects.
Having said that, my own experience suggests that when you have a startup, you should just build single-threaded applications that clean up after every request (such as with PHP) and spawn many of them. They will share database pool connections etc. but for the most part it will keep your app safer than if they all shared the same objects and process. The benchmarks say that PHP-FPM is only 50% slower than Swoole for instance. So why bother Starting safe beats a 2x speed boost.
And by the way, you should be building distributed systems. There is no reason why some client would have 1 trillion rows in a database, unless the client themselves is a giant centralized platform. Each client should be isolated and have their own database, etc. You can have messaging between clients.
If you think this is too hard, just use https://github.com/Qbix it does it for you out of the box. Even the AI is being run locally this way too.
Just like comparing C vs Python performance.
"B...b...but Python is easier to write."
Sure. That is a trade off. And whether the trade off makes sense depends on the situation.
Unpopular opinion: this is a good idea. It encourages you to structure your code in a way such that computation and I/O are separated completely. The slight tedium of refactoring your functions into a different color encourages good upfront design to achieve this separation.
One should design based on constraints that best match the problem at hand, not some ossified principle turned "universal" that only really exists to mask lower-level deficiencies.
Excessive hiding of implementation detail is what leads to things like fetching a collection of user IDs from a database and then fetching each user from an ID separately (the 1+N problem). Excessive hiding of implementation detail is what leads to accidental O(N^2) algorithms. Excessive hiding of implementation detail is what leads to most performance problems.
I guess Haskell is fully based around that idea :)
It’s as good as it gets for this type of thing as it was designed for it from the ground up.
Last I checked it used about 2.5K. Beam reserves around 300 words per process (depends on exact options), but a word is 8 bytes not 8.
You can get it lower (obviously at a cost as soon as you start using the process and need space) but nowhere near 512 bytes, just the process overhead is around 800.
You can think of the Java / Go approach as adding the bind points automatically.
I missed my opportunity to reply to your comment, but I really appreciate it, and I wanted to find a way to get this back to you. The comment in question:
"Well, I was one of the engineers that made the change :) I'm not sure how much I can tell, but the public reason was: "to make pricing more predictable". Basically, one of the problems was customers who just set the spot price to 10x of the nominal price and leave the bids unattended. This was usually fine, when the price was 0.2x of the nominal price. But sometimes EC2 instance capacity crunches happened, and these high bids actually started competing with each other. As a result, customers could easily get 100 _times_ higher bill than they expected."
There was more to it than that, but I figure that's a good enough reference point.
Thank you for these improvements. It doesn't change anything, in terms of how much savings I can get by following the latest generations and exotic instance-types, but it does help with the reliability of my workloads.
It's been a huge benefit to me, personally, that I can provide some code that enables the potentiality of servers dying, with the benefit of 80% cost savings without using RIs.
You can also try another product of my former team: https://aws.amazon.com/savingsplans/ - it's similar to RI, but cheaper because it doesn't provide an ironclad guarantee that the instance will be available at all times. It's still a bit more expensive than spot, but not by much.
But then I read
> With a little help from ChatGPT
And sure enough the results & prose end up pretty much worthless. Because the author didn't do any actual critical thinking or technical work. (For instance, there's nothing more to "Go's memory usage ballooned at 1M tasks" beyond the observation. Why not? Just dig in there it's all documented you just have to read some stuff.)
It doesn’t matter that node is single threaded, nor does the GIL have any impact on single thread.
It’s once you start doing cpu heavy work inside the tasks this will hit you.
So it's likely these numbers simply reflect choices in the default configuration, not the ultimate limit with tuning.
I'm starting to dislike these low effort ChatGPT articles. Getting an answer from it is not the same thing as actually learning what's going on, but readers will walk away with the answers uncritically.
The next few decades are going to be tough for you, I think. Brace yourself!
I think this applies all the way down to hardware features and programming languages. For example the naive use of sort() in C++ will go faster than in Rust (in C++ that's an unstable sort in Rust it's stable) but may astonish naive programmers (if you don't know what an unstable sort is, or didn't realise that's what the C++ function does). Or opposite example, Rust's Vec::reserve is actively good to call in your code that adds a bunch of things to a Vec, it never destroys amortized growth, but C++ std::vector reserve does destroy amortized growth so you should avoid calling it unless you know the final size of the std::vector.
Sure. Will redo this part. I doubt it would allow Elixir to get a better result in the benchmark, though, as it was already losing significantly at 100k tasks. Any hints on how to decrease Elixir memory usage? Many people criticize using defaults in the comments, but don't suggest any certain settings that would improve the results. And a part of blogging experience is too learn things - also from the author perspective ;)
1..num_tasks
|> Task.async_stream(fn _ ->
:timer.sleep(10000)
end, max_concurrency: num_tasks)
|> Stream.run()
https://elixirforum.com/t/how-much-memory-do-you-need-to-run... $ /usr/bin/time -v elixir --erl "-P 10000000" main.exs 1000000
08:42:56.594 [error] Too many processes
** (SystemLimitError) a system limit has been reached
Any other ideas how to raise the limit?“This limit can be adjusted with the `+P` flag when starting the BEAM… To adjust the limit, you could start the Elixir application using a command like the following:
elixir --erl "+P 1000000"
“If you are getting this error not because of the BEAM limit, but rather because of your operating system limit (like the limit on the number of open files or the number of child processes), you will need to look into how to increase these limits on your specific operating system.”BTW, I fixed the 1M benchmark and Elixir is included now.
You can't just copy-paste ChatGPT without even a sanity check e.g., Python code is wrong:
await sleep(10)
sleeps 10+ seconds and returns None but the intent is to pass a coroutine to create_task() (await should be removed at the very least)And BTW - I'm very far from copy pasting from chat gpt. ChatGPT helped with this code, but often required many iterations and also some manual cleanup.
import asyncio
async def do_task():
await asyncio.sleep(10)
async def main(num_tasks: int):
tasks = []
for task_id in range(num_tasks):
task = asyncio.create_task(do_task())
tasks.append(task)
print("spawned tasks")
await asyncio.gather(*tasks)
if __name__ == "__main__":
asyncio.run(main(1_000_000)) tasks.push(
(async () => {
await delay(10000);
})()
);
.. So tasks.push is called with an IIFE, which returns a promise, which resolves after its internally awaited another promise...You could just do this, which avoids about 3 unnecessary allocations:
tasks.push(delay(10000))
And I'm not convinced promisify(setTimeout) is what you want here. I'd just use this: const delay = timeout => new Promise(resolve => setTimeout(resolve, timeout))
... Which honestly should be part of the standard library in javascript. However, I'm less convinced that replacing the delay function will make a big difference to performance.I uh, may be in this picture.
Promise timers are now packaged by node at least, they allocate about 20% less memory for the 1mil test (on v18).
const { setTimeout } = require('node:timers/promises')
https://nodejs.org/api/timers.html#timers-promises-api- single job: 27.8 mb (respectable)
- 10k: 8.3 GB!
At this point, I gave up. I'm 99% sure trying to go to 1 million will crash my computer.
[1]: https://www.php.net/manual/en/intro.parallel.php
[2]: https://gist.github.com/withinboredom/8256055709c0e6a8272b3a...
Got to love COW memory. :)
[1]: https://gist.github.com/withinboredom/c811f3c433904235a5f196...
I don't even know where to start: from the lack of env settings, system settings, compiler settings to the fact that no concurrent work was actually done, this is one giant flawed benchmark.
The only lesson here should be: Don't take any of this seriously.
compiler settings?? I bet he used defaults everywhere
> a high number of concurrent tasks can consume a significant amount of memory,
So Go needs 3gigs for a million tasks. Relatively speaking, for an app that does a million concurrent somethings, 3gb is probably peanuts.
It seems that memory isn’t the limiting factor here.
n = ARGV.first.to_i32
finished = Channel(Nil).new
n.times do
spawn do
sleep(10.seconds)
finished.send(nil)
end
end
n.times { finished.receive }
Memory usage determined with `ps ax -o pid,rss,args`. Runtime measured with `/usr/bin/time -v`. (Had to run `sudo sysctl vm.max_map_count=2500000` to get the final two rows.) Results: N Runtime(s) RSS(kiB) RSS(MiB)
=======================================
1 10.01 1792 2
1000 10.01 6016 6
10000 10.11 45184 44
100000 11.00 435840 426
1000000 20.25 4336768 4235
"System time" rather than "user time" is the majority (7.94s system time, 2.83s user time in the 20.25s wall time run). Is this pointing to memory allocations?Crystal Fiber docs https://crystal-lang.org/api/1.8.2/Fiber.html says "A Fiber has a stack size of 8 MiB which is usually also assigned to an operating system thread. But only 4KiB are actually allocated at first so the memory footprint is very small." -- and perhaps unsurprisingly, 4KiB times 1000000 is approximately 4 GiB. Nice when the math works out like that :)
To me this is useful as a baseline for something like a websocket server, with some number of idle connections, each Websocket connection mapped to a Fiber that is waiting for I/O activity but mostly idle.
Gotta wonder whether you can challenge those numbers with c++ though. I suspect you can but with a lot of work, where Rust is pretty effective written naively.
Of course the big issue with these comparisons is it's hard to be an expert in so many different languages. How do you know if what you're doing is the best way in each language?
Perhaps it's because I generally don't write code for many short lived workloads, but I find using real threads to be a lot easier to work with (plus it saves time during compilation). Tokio is great at what it does, but I have to wonder how relevant these types of threads are when "fearless concurrency" is one of the token advantages of the language.
I think this article would've been a lot better if it did some actual calculations in the spawned threads, rather than simply waiting. A smart enough compiler would see that the results of the wait calls, which usually don't come with any other side effects, are never used and optimize them out entirely.
- How are you exactly launching the workload?
- How much memory does your system have?
- Could you try running it with .NET 7 instead?
The last two are especially relevant because in the recent versions each release got extra work to reduce memory footprint on Linux (mostly because it's an important metric for container deployments).
Overall, 120MB sounds about right if you are launching a .NET 6 workload with some extra instrumentation attached while the runtime picks Server GC which is much more enthusiastic about allocating large segments of memory, especially if the system has large amount of RAM.
As for Tasks - C# tasks are in many ways similar to Rust's futures (both are yielding-to-the-caller state machines) and take very little memory, which scales directly with the amount of variables that persist across "yield points" (which you have none).
Also if you have lots of CPU cores then the CLR will (pre-)allocate up-to ~10MB/core for per-hardware-thread GC heaps. So my 16-core (32-thread) desktop will run every .NET program with a starting (OS-level) memory usage of ~200MB for a Hello, World program - though actual total GC usage will be < 1 MB.
For a bunch of these tests, it is measuring 1 million tasks multiplexed onto a smaller number of os threads.
In many implementations the cost of concurrency is going to be small enough that this test is really a measure of the stack size of a million functions.
JVM for example has a default heap size of 25% to 50%, depending on how much memory is available in the system. Given that the code is using ArrayList, there will be more need for bigger heap, due to array expansion. To eliminate that noise, the code should be changed to a pure Thread[] array with a fixed size based on numTasks.
Once that is done, you may as well run up to 10_000 tasks (virtual threads) within a JVM that consumes no more than 32 MB of RAM (compared to 78MB as originally reported by OP).
These were the best results I could find after playing a bit:
# Run single task
docker run --memory 26m --rm java-virtual-threads 1
# Run 10 tasks
docker run --memory 26m --rm java-virtual-threads 10
# Run 100 tasks
docker run --memory 26m --rm java-virtual-threads 100
# Run 1_000 tasks
docker run --memory 28m --rm java-virtual-threads 1000
# Run 10_000 tasks
docker run --memory 32m -e JAVA_TOOL_OPTIONS=-Xmx25m --rm java-virtual-threads 10000
# Run 100_000 tasks
docker run --memory 200m -e JAVA_TOOL_OPTIONS=-Xmx175m --rm java-virtual-threads 100000
# Run 1_000_000 tasks
docker run --memory 950m -e JAVA_TOOL_OPTIONS=-Xmx750m --rm java-virtual-threads 1000000Maybe I’m cynical but this is what happens if you rely on chatGPT to give you programming answers without understanding the fundamentals first.
https://www.erlang.org/doc/efficiency_guide/advanced.html#:~....
edit: oh apparently it’s got nothing to do with memory, 1M just happens to be higher than the default process limit for Erlang and thus Elixir.
It is interesting that some systems allocate more memory up-front but this doesn't grow much per-task.
It is interesting that in the form the app was written the Go program doesn't appear to scale as well. Sure there are possibly reasons for it so anyone with more knowledge of Go could attempt to write a more "correct" program and ask to rerun it.
It is interesting that with the defaults, things behave as they do, again, someone who cares more about "it's just that the defaults are different" is welcome to investigate that and ask "why" some program defaults might lead to better performance.
It is interesting that as-in most comparisons, the languages/frameworks don't neatly fit into Stereotypes.
It is interesting that most of them probably perform well enough for 99% of applications we might write.
It seems like most people still don't appreicate that YMMV and that trade-offs are always made. Just add it to your stash of brain juice that helps you become a better engineer and stop complaining.
There may be scenarios such as network connexion handling where those connexion are kept idle 99% of the time, but you still need something to handle the very rare case of receiving something. However, is 1 million network connexion manageable from a kernel point of view anyway ? Wouldn't you reach system limits before that ?
You can raise all those defaults as long as you have sufficient memory. Yes, you need to be working with fairly idle connections, otherwise you're likely to run out of CPU. Real world examples include Chat, IMAP, push messaging, that sort of thing where having a live connection means a significant reduction in latency vs periodically opening a connection to poll, but at the same time, there's not a whole lot going on.
I was at WhatsApp when we accidentally hit 2.8M connections on one machine [1] I can't remember if we ever got to 3M, but we ended up doing a lot more work per connection so didn't keep those connection counts for long. Kernel wise, as long as you're mostly inbound connections, everything is well oiled; I ran into problems initiating many thousands of outbound connections/second in an HAProxy use case, although I think FreeBSD 13 may have improved that bottleneck. I hear Linux also works, but I've not scaled it nearly as far.
[1] http://www.erlang-factory.com/upload/presentations/558/efsf2...
Note also, that BEAM uses several allocation arenas and is slow to release those back to the OS. It's important to measure memory from the OS perspective, but in a real application you'd also measure memory from inside BEAM. There are knobs you can tune if you need to reduce apparent OS memory use at the expense of additional cpu use during allocation. There's tradeoffs everywhere.
Also, yes, 4GB is a lot of memory, but once you start doing real work, you'll probably use even more. Luckily memory is not that expensive and capacities are getting bigger and bigger. We ran some servers with 768GB, but it looks like the current Epycs support 6TB per CPU socket; if everything scales, you could run a billion sleeping BEAM processes on that. :p I recognize this sounds a lot like an argument to accept bloat, but it's different because I'm saying it ;) Also, I think the benefits of BEAM outweigh its memory costs in general; although I've certainly had some fights --- binary:copy/1 before storing binaries into ets/mnesia can be really helpful!
[1] https://www.erlang.org/doc/efficiency_guide/processes.html
I wonder if there are any correctness constraints that would actually prevent such optimizations. For example, the compiler might be required to execute the exact same system calls as an unoptimized program. It seems quite reasonable that the programs using system threads could not be optimized significantly under such a constraint.
I would expect the larger limit that prevents compilers from optimizing this is the large amount of code that defines the semantics of the lightweight threads. For example, I would expect the observable / required behaviors of tokio to not prevent a SCC from optimizing it, but the massive amount of Rust code that defines the lightweight thread semantics cannot be reasonably optimized.
Which programming language / framework is the closest to actually being optimized to like < 4 byte / task? I would make an uneducated guess of Elixir or Go.
It has a clock thread on the C++ part of the engine and when it timeouts it pushes the continuations on the event loop, and that is it
Not sure if this is what I used but the settings sound familiar: https://www.baeldung.com/linux/max-threads-per-process
Also would be good to specify how the memory as measured (eg to be able to verify they are actually measuring physical memory usage).
seq{1..1000000}
|> Seq.Con.iterJobIgnore (fun _ -> timeOutMillis 10000)
|> run
I'm not sure how to calculate total memory usage but instrumenting the following way: GC.Collect()
GC.Collect()
GC.WaitForPendingFinalizers()
GC.TryStartNoGCRegion 3000000000L
let before = GC.GetTotalMemory false
seq{1..1000000}
|> Seq.Con.iterJobIgnore (fun _ -> timeOutMillis 10000)
|> run
GC.GetTotalMemory false - before |> printfn "%d"
GC.EndNoGCRegion()
Gives the following output, I'm reading 144MB for 1M lightweight thread (in the REPL) unless I did something dumb: 144449024
Real: 00:00:10.902, CPU: 00:00:18.570, GC gen0: 3, gen1: 3, gen2: 3
val before: int64 = 53902280L N Real(s) User(s) RSS(kiB) RSS(MiB)
==================================================
1 10.08 00.04 42660 41
1000 10.07 00.06 40736 39
10000 10.09 00.20 44176 43
100000 10.15 01.85 58532 57
1000000 10.88 18.18 182024 177
10000000 22.28 4m04.84 1452224 14181 million Task.Delays means 1 million queued waits in the sleep machinery, which could run into some bad scaling and overhead. But then 1 million tasks waiting on a single awaitable probably hits overhead in other subsystems. My guess would be that 1 million waits is much cheaper than 1 million sleeps, because sleeps have to be sorted by due time while the list of tasks waiting for completion just needs to be sequential. Entries in the sleep queue are also going to be bigger (at minimum, 4-8 bytes for due time + 4-8 bytes per completion callback) while entries in a completion queue would just be 4-8 bytes for the completion callback.
Task.WhenAll also has potentially worse behavior than the other apps if you care about memory, since it will need to allocate a big array of 1 million Tasks in order to pass it in. Many of the other programs instead loop over the tasks, which eliminates that overhead.
The Run() call might be creating a more complex state machine and using more memory than necessary (2 tasks vs just 1 for each delay). The memory usage might be even lower if the Run() call is dropped.
Task task = Task.Delay(TimeSpan.FromSeconds(1)); tasks.Add(task);
The difference is quite significant.
(1) With the authors code (using Task.Run), I get ~428MB of allocations.
(2) Dropping the unnecessary Task.Run(...), I get ~183MB of allocations.
(3) Doing (2) and waiting N times on the same delay, I get ~39MB of allocations.
This was all using .NET 6 too. .NET 7 or the 8 preview might be even better since they are working hard on performance in recent releases.
So even looking at just (2), that puts .NET on par with the rust library.
However AIUI C++ 20 doesn't actually supply an executor out of the box, so you would need to choose what to do here as for Rust, where they picked both tokio and async-std and you can see they have different performance.
C does only have threads, but you could presumably pull off the same trick in C that Rust does, to get Linux to give you the bare minimum actual resident memory for your thread and you needn't care about the notional "real" size of the thread since we're 64-bit and address space is basically free.
LMAO
C++ is the de facto language used for hight performance.
You're describing features not performance.
Possible optimizations depend on the semantics of the language — C++ can do just orders of magnitude more with for example a string data structure, than what is possible with a C “string”. So no, C is not the most performant language in general — de facto C++ is used for everything performance critical.
Yes, or use macros. Not saying it's the most convenient, but that's an orthogonal concern.
You wouldn't say hand-written assembly isn't performant, so your objection strikes me as wrong.
My results is 565Mb for the 1 mill tasks, vs the authors 2.6Gb.
My method is described in [1], feel free to point out what I am missing here.
[1] https://gist.github.com/nilsmagnus/38b1b5f63237a2a6317942849...
In the nodejs test, there's very little javascript code being actually executed here. Whats really being tested is nodejs's event queue (libuv), which is written in C. And its being compared to Go's coroutine system - which is also in C. But, if I recall correctly, Go's threading model includes a preemptive scheduler (which is very complex) and it allocates a stack per goroutine.
This doesn't test if Javascript or Go code is faster / more efficient. There's almost no javascript or go code being executed here at all. This is testing the execution engines of both languages. Go's execution engine has more features, and as we can see those features have a cost at runtime.
Concurrent meaning multiple tasks run within a given stretch of time. Parallel meaning multiple tasks run at a given point of time.
These are not the same. So the tests are not testing the same thing.
Oof, has C/C++ become that unpopular? It would be a useful comparison as low level system language anyway.
P.S. the statement above is imprecise: https://devblogs.microsoft.com/dotnet/configureawait-faq/
https://github.com/search?q=repo%3Atokio-rs%2Ftokio%20unsafe...
https://www.techempower.com/benchmarks/#section=data-r21&tes...
C doesn't have a built in, optimized async execution engine of any sort. C with threads would lose to any of the async runtimes here (including javascript).
C++ has async now. By all means, write up an async C++ example and compare it. I suspect a good C++ async executor would perform about as well as tokio, but it'd be nice to know for sure.
I'd also be interested to see how hard it is to write that C++ code. One nice thing about Rust and Go is that they perform very well with naively written code. (Ie, the sort of thing chatgpt, or most of your coworkers would write.) Naive C would use kernel threads, so if we're comparing average C code to average Go code at this test, the average go code will end up faster.
(This isn't always true of rust. I've seen plenty of rust code which uses Box everywhere that runs crazy slow as a result.)
Just-js was essentially built to run the techempower benchmark. It's not a real stack you can use. An exercise in "what is possible".
Some runtimes can handle this for you, but they come with other performance tradeoffs.
I wonder if the parallel stream API with a virtual thread pool would make it better or worse.
Since this is amortized exponential growth it's likely making negligible difference, although it is good style to tell data structures what you know about their size early.
The Go approach doesn't appear to pay this penalty, it's just incrementing a (presumably atomic?) counter, although I have no idea where these tasks actually "live" until they decrement the counter again - but most of the others are likewise using a growable array.
In the unmanaged languages, those intermediate buffers should be freed immediately during the re-size operation.
But in the managed languages, they might stick around until the next GC. So the peak memory would be higher.
In this example the runtime overhead is going to represent more of the memory usage anyway but it quickly isn't negligible. E.g. a list of size 1M is going to be ~8MB (64 bit addresses). So even if the array is re-sized from 1->2->4->8->...->~500K, the end result is not going to be worse than 2X the size.
This benchmark wasn't very rigorous, so maybe it has overlooked something, but the result seems plausible. Tokio is a good, mature runtime. Tokio-based frameworks are near top of TechEmpower benchmarks.
Rust's Futures were designed from the start with efficiency in mind. Unlike most async runtimes where chaining of multiple async calls creates new heap allocations, Rust flattens the entire call tree and all await points into a single struct representing a state machine. This allows the runtime to use exactly one heap allocation per task, even if the task is pretty complex.
There is so much left up to the implementation details that you are not benching anything.
EDIT: polls->pools typo
A scenario with this many tasks spawned is more realistic if you have mostly idle connections. For example, a massive chat server with tons of websockets open.
For example - https://go.dev/doc/go1.3#stacks https://agis.io/post/contiguous-stacks-golang/
So your measurements in go 1.2 would probably differed in go 1.3 (and this is just one example, in only one of the possible axises that have changed).
It's not like you own the data structure there to control it.
The Rust tokio solution will use tokio's runtime which automatically spawns an appropriate amount of worker threads to parallelize your async tasks (by default it asks the OS how many CPUs there are and creates one such thread for each CPU).
The async-std implementation is roughly similar.
I didn’t quite get this part. Even if your runtime has no overhead, the fastest way you can execute 1 million 10-second tasks is 10 * 1 million / num_cores, right? Why would these million tasks only take 12 sec?