That's a really interesting approach. I could see that being really useful in a lot of cases, and it lets you get good single-threaded performance on code that would also continue to work in a multi-threaded case.
That's a really interesting approach. I could see that being really useful in a lot of cases, and it lets you get good single-threaded performance on code that would also continue to work in a multi-threaded case.
Persistent data structures are really about the interface: functions that would be traditionally mutating instead return new versions. Is there an advantage of that interface over mutable value-type collections? One advantage is that it works better with e.g. folds. But overall I think it's more awkward.
It could make refactoring and debugging easier. For instance, let's say you're part way through some operation and it fails, and the error recovery logic is buggy and leaves the data structure in an inconsistent state. You could debug the problem, but if you're in a hurry and you just need something that works right now, you can store a reference to the data structure at the beginning of the operation and revert to that if anything goes wrong. You'd pay a performance cost since all the direct in-place edits would turn into copies and allocations, but it would at least work.
Similarly, you could develop a program initially in an all-immutable multi-owner style, and then selectively convert codepaths to single-owner style to get better performance in the parts where performance matters.
Also, shared data are much easier to reason about if it's immutable (not necessarily even in a concurrent execution context).
For instance, if you write an interpreter, representing variable bindings with a functional map allows to reason about scope super easily; when you change scope, you just create a new map with the new variable bindings, and when the scope ends you just get back the map at the scope beginning which is still valid.
With a stateful hashtable, you need to undo the changes you did in you scope; the value of the table when opening the scope that you want to get back to is gone forever. Now it ties you program to a certain execution order. One way to see that is that if you want to collect environment to each scope in some data structure for some reason, you need to make deep copies everywhere.
It impacts the performance profile of the collection, and thus isn't merely an implementation detail.
> I may be misunderstanding the implications of Arc and how it is implemented, and how it behaves in the absence of multiple threads
All Arc means is that refcounting operations use atomic increment/decrement instead of regular ones.
EDIT: wording
Yes. An interesting note here is that due to Rust's ubiquitous use of references there will be little extraneous refcounting traffic unlike systems like e.g. Python where `a = b` might trigger a refcount increase. I think it was Armin Ronacher (mitsuhiko) who made that observation a few weeks back, possibly on twitter?
Yes. You have Rc and Rc::make_mut for the single threaded case. Pick what you need. I interpreted the GP comment as talking about the general pattern of make_mut, not Arc::make_mut itself.
Stylo's multithreaded, we pick Arc. Singlethreaded systems can get by with Rc, but there are often legit reasons for picking Arc, and it's not that bad a hit.
... isn't this just basic copy-on-write ? that's something you learn to implement in first year of comp-sci. Old C++ types used to work like this - and some still do, for instance Qt's collection types are all COW.
This is copy-on-write-unless-we're-alone-then-just-mutate-without-copy. That is, it's possible to write without copying if it has been determined that there are no other readers that would be referring to the old value (thus no copy is needed).
> Implicit sharing automatically detaches the object from a shared block if the object is about to change and the reference count is greater than one. (This is often called copy-on-write or value semantics.) [source: http://doc.qt.io/qt-5/implicit-sharing.html]
The C++ language is sometimes a little cruel, though. It's easy to accidentally call the mutable version of a function when the const version would have been sufficient. That can be a performance problem if it causes unnecessary copies.
well, yes, that's generally called copy-on-write. A COW class which does a copy every time you write to it would be pretty useless, woudln't it ?
e.g. if you did, in old std::strings:
std::string s = "foo";
s[1] = "x"; // no copy here
auto s2 = s; // no copy here;
s[1] = "y"; // s is copied in a new buffer
s[1] = "z"; // no copy here
s2 += "bar"; // no copy here