Inside D's GC
olshansky.me
olshansky.me
The problems are known and indeed a full rewrite on this ancient GC would be in order. Since D 2.072.0 it's possible to link and use different GCs, so a faster one could be written as an external library which would be a very welcome effort. https://dlang.org/changelog/2.072.0.html#gc-runtimeswitch-ad...
I'd be interested in experimenting with a Connectivity-Based GC (https://www.cs.purdue.edu/homes/hosking/690M/cbgc.pdf), leveraging type information to do partial collections. Write barriers come with a performance penalty (~3-5%) that we don't want to impose on people using deterministic memory management, but most GCs capable of performing partial collections, e.g. generational GCs, do require write barriers, Connectivity-Based GCs do not.
For the time being the pragmatic advise is to replace major sources of GC allocations with deterministic memory management when the GC heap grows too big (~1GB) and performance becomes a problem.
So far this hasn't prevented various people from writing extremely fast D programs.
Is this really true?
I'd like to port author's effort to Windows, but I don't have OS-level programming experience, so I don't even know what are the differences and how they affect GC algorithms described.
https://dlang.org/spec/garbage.html#obj_pinning_and_gc
There's also been some work on a precise GC for D:
https://forum.dlang.org/post/hdwwkzqswwtffjehenjt@forum.dlan...
Not generally, but you can take the address of anything and shove it into a union or pass it off to C code that stores it somewhere. Because you can, it would be very dangerous to move objects around because of that.
Some of the mistakes seen by the author could just be the original programmers being tired of painful debugging.
One charge not leveled at D's GC is bugginess. The GC has proven to be pretty robust.
Circular references? GC will collect them! (Reference counting does not handle circular references.)
Usually the two mistakes with GC are forgetting to empty collections, (which is also a problem with manual memory management,) and relying on finalizers to clean up resources.
GC will not "read your mind" and magically close your open sockets, release your file handles, and dereference your static references. It just frees you from the tedium of releasing transient memory, like strings; and it frees you from tracking lifetimes of memory when you don't need to.
In my experience, the programmers who have trouble with GC are the ones who don't understand the difference between memory management and resource management.
That being said, I anticipate that I'll write a lot of Rust in the future!
The whole lifetimes business, which is done more generally for memory safety. It adds a lot of complexity.
On the other hand, there's clearly a niche need for a language with no GC at the systems level, so I can appreciate the reasons for Rust's decision. But after playing around with Rust for quite a bit, I'm come to the decision that at least for me, Rust's lifetimes are too complex, and (partly as a result of no GC), the overall language is just not functional enough to be a general purpose programming language. Scala is currently filling that role for me.
Additionally, lifetimes are also part of the system that statically avoids problems that appear in GC'd languages like iterator invalidation (for instance, ConcurrentModificationException).
The problem is that a lot of mainstream programming languages pretty much ignored the state of the art when implementing their concurrency models. See Per Brinch Hansen's rant about Java [1], for example.
Thank you for the references (although I think you meant Flanagan without an h). From a quick glance at one of the papers (without any guidance from an expert I have no idea where to start for a good summary), it seems like the retrofitted PRFJ has to fight the pervasive sharing of having a GC, resulting in a much less orthogonal system with higher annotation burden. In any case, at least two of the Rust team did PhDs in the areas of parallelism/concurrency so I'm sure they can give a more academic comparison of Rust and the research.
Neither. I am just unhappy that these days avoidance from data races often gets reduced to the ownership approach and I wanted to provide a more complete picture (even if a quick summary is still woefully incomplete).
(And thanks for pointing out the typo – now fixed.)
That "additionally" is exactly why I said "more generally". I didn't say that lifetimes were not for thread safety. I just said that they were more generally for memory safety (I would count iterator invalidation etc as memory safety, would you not?).
Or put another way, having memory safety doesn't guarantee iterator safety.
And sometimes those needs are so pervasive that I want or need to effectively avoid using a GC myself - and so reach for a non-GCed, non-JITed, "manual control over everything" language. This I do not have well covered - traditionally my choice has been the undefined behavior infested, data race prone, "didn't even have threads in the standard library until ~2011", source of bugs and misery - C++.
Rust's lifetime and ownership semantics are complex, yes. But on the other hand, they can be used as a much more uniform, clean, and flexible version of e.g. https://clang.llvm.org/docs/ThreadSafetyAnalysis.html , which is but the tip of the iceberg of static and dynamic analysis, unit and fuzz testing, etc. I've already been applying to my C++ programs, to try and keep the bugs in check. So Rust is simpler than what I've been doing in practice.
History is full of GC enabled systems programming languages since the Xerox PARC days, the latest example being System C# for Midori.
Most projects failed more due to political reasons than technical ones.
Which is why I appreciate Apple, Google and Microsoft's attitude regarding forcing automatic memory management on developers for their OSes, to a certain extent.
Since that is the exact use case that interests me when it comes to languages such as C, Rust and D I would rather see the whole thing taken care of at compile time with very strict rules about how long it can take to traverse a certain piece of code, especially when it runs on the other side of the syscall interface.