Calling into C is a big one: we concluded it was basically impossible to make Rust calling into C as cheap as C calling into C in a segmented stack implementation. In practical usage, we found that the ability to call into C cheaply was much more important than the benefits of small stacks, which mostly help microbenchmarks like this at the expense of real-world code.
For Rust's goals, I can definitely believe that the first is true. I don't think the second claim is true at all, though. Lightweight threads have been extremely successful in Haskell, Erlang, etc, and not just on microbenchmarks.
They're mostly successful in languages that GC/heap-allocate all stack frames (paying the costs that come with it). Rust uses the machine stack so the benefits come with some very significant drawbacks, most notably unpredictability of performance that comes from stack thrashing. Large mallocs are slower than small mallocs in general, but on the typical sizes you need for machine stack segments you probably aren't going to hit the malloc implementation's free list anyway, so you might as well just allocate up front. If you GC stack frames, though, you just bump allocate in the nursery, at the cost of greatly increased GC pressure (but it makes task spawning much cheaper).
There's some interesting discussion on this near the end of Simon Marlow's paper "Extending the Haskell Foreign Function Interface with Concurrency": http://community.haskell.org/~simonmar/papers/conc-ffi.pdf
One of the enlightening things for me about CSP and Erlang, in particular, was using channels and messages and a bunch of lightweight tasks for flow control and state management. Reading the reddit comments, and even your reply here, seems to indicate that that type of programming isn't a good fit for rust, and that perfectly reasonable trade offs were made that led the language in a different direction. I should have phrased the comment better, but the recommendations here and in the reddit thread boil down to replying to "it hurts when I do this" with "don't do that".
Obviously no-one is trying to write programs that count to 10001, but the capability of doing it that way efficiently opens up a bunch of different options.
It will be interesting to see what happens to go's performance on things like this if they move away from segmented stacks.
This might just mean that it is more senceable to span a smaller number and communicate with them via port.
Its still CSP, its still usful, but its not exactly like go. My guess is that if you start to do more computation in each thread the benchmarks are gone change.
The Go style is only cheap if the stack does not grow.
What types of CSP things should you do in Rust, then? Which should you not? Why is this code not a demonstration of a slowness in Rust but a 'bad idea'? How would you implement it in a more idiomatic manner in Rust?
I find his comment pretty spot on. If anything it was the "don't do CSP in Rust" comment that was unconstructive and passive aggresively insulting. Instead of understanding what he read and the constrains it shows, the commenter preffered to just piss on the language.
As for the example, it was obviously "retarted" or rather, contrived. I mean, heck, it's "chinese whispers" with 10000 channels, what more do you need to see that this is not a serious way of solving actual problems with CSP? It's a bloody microbenchmark, and those are rarely representative of anything.
>What types of CSP things should you do in Rust, then?
The types of CSP things that people use threads and processes for. E.g getting lots of computations going on in parallel, not incrementing tens a single int in each of thousands of "computations".
CSP and channels are meant for getting real (cpu intensive) work done. You don't just use threads (or greenthreads) "because concurrency".
That some Go programmers like to use channels as a control mechanism, doesn't mean CSP was meant for that kind of abuse. That's like using Exceptions for control flow to me.
Isn't it obvious that even in Go the times to complete this are horrible (just an order of magnitude less horrible because of different design trade-offs), and that doing the same job of 10,000 incrementings properly would complete in 1/1000 the time?
> CSP and channels are meant for getting real (cpu
> intensive) work done. You don't just use threads (or
> greenthreads) "because concurrency".
The whole point of green threads is that they're orders of magnitude cheaper to create (and destroy) than real, operating system threads. They are precisely around "because concurrency" -- they allow you to nicely model concurrent problems without having to be overly concerned with the implementation detail of their creation cost.While it's not idiomatic in Go to use a channel/goroutine combo as an iterator, for example -- that's too low-level -- it's absolutely idiomatic to use one for other types of higher-order control flow, managing state machine transitions, doing a scatter/gather, and so on.
> Rust tasks have the same large fixed-size stack as OS
> threads. A fine-grained concurrency model like a task
> graph would be build on top of them.
In the absence of other context (I don't really know much about Rust) I would then argue that Rust tasks miss the point of CSP.Besides, the only substantive difference between Go's implementation and Rust's implementation is that Go uses segmented stacks, while Rust does not. That is because segmented or relocating stacks are in opposition to Rust's design goals of no GC, fast calling into C, and predictable performance.
This seams more idimotic, and more CSPish to me.
(Passing around pointer/references to immutables is of course ok)
I may misunderstand what you're talking about, but as far as I know Go more or less passes a raw pointer (or a copy thereof) over pointer channels. It has no concept of ownership, and the pointer (and pointee) on the sender side remains valid: http://play.golang.org/p/E1bqrVFxZ6
Rudeness of the person you're replying to aside, I think it's well explained elsewhere in the thread that the Rust people have made choices that don't necessarily give great speed on pointless microbenchmarks (making a bunch of processes that do nothing) in favour of performance on CSP tasks where your CSPs actually do things.