Java’s new garbage collector promises low pause times on multi-terabyte heaps
opsian.com
opsian.com
Not bad! Looking forward to seeing how this performs with a diverse range of workloads as it matures.
Not everyone needs to try to write the next Crisis in it.
It certainly has higher performace than all those HTML 5 and Flash games.
The biggest problem is that jMonkey is the only RAD tooling available and no match against any of engines that make use of C#.
I think it is safe to assume that under "game development" most people understand rather high performance demands, not a solitaire clone (esp since the comment you responded to pointed out pause-times as crucial).
Of course those of us hosting servers for that particular game would have saved a ton of money and pain if they had not used Java, or used Java more cleverly.
One can make a game in C++ and make it just as bad.
* I think Java has escape analysis, which means that if it can determine that an object you created doesn't leave its context, it will go on the stack and won't add work for the garbage collector.
Also an interesting aside is that objects are never 'allocated on the stack'. As in you won't find the same layout of memory on the stack as you do on the heap. Instead objects are turned into data edges in the compiler's graph. It's much more abstract than literally doing an alloca.
If you have an array of these, each element in the array will be a reference/pointer to a separately allocated Point instance which can be anywhere in the heap. Accessing the array is expensive due to lack of locality (arbitrary memory access). Additionally, each instance has substantial object overhead (typically 8 or 16 bytes).
There are various ways to deal with this. One way is to have separate arrays for x and y. Another way is to use a byte array and use Unsafe to read/write the integers. This is terrible from a developer standpoint and is only acceptable if it is a small part of your application.
Value types solve this by giving you an efficient memory layout while still allowing you to code against them as normal. The slogan is "codes like a class, works like an int".
class B {
A a1;
A a2;
}
is one continuous 64-byte block of memory. Without value types, each A will be allocated separately on the heap, and B will be a pair of pointers.Lack of value types makes it hard to get cache efficient memory layouts in Java.
That's not quite correct. A reference type might contain a value type field, in which case the value type will still live on the heap.
Conversely, a local variable's type might be a value type, and _still_ involve a heap allocation. For example, std::vector in C++ is a value type but the array storage is heap-allocated.
So it's really about semantics, and not memory layout. Accessing an lvalue with a value type produces a copied rvalue; with reference types you're only copying the reference itself.
This isn't correct. Firstly, C++ doesn't have value or reference types. Where any object is stored depends purely on how it is allocated. Either in the free store (dynamic) or automatic storage (the stack).
std::vector has a constant object size - typically the size of 3 pointers (but that depends on the implementation - the standard doesn't specify): the start of the memory (dynamically allocated for storage), end/capacity (typically either another pointer or a size_t of the current capacity), and back (typically either a pointer just past the last element used or a size_t + 1 of current size). It's storage is dynamically allocated, though, and for a std::vector<T> will have dynamically allocated sizeof(T) * capacity bytes (maybe more due to alignment issues).
std::array is constexpr safe and does not (directly) dynamically allocate memory - size has to be known at compile time (hence it's a template argument).
For instance, on my compiler (GCC 4.8.4), for this example program:
#include <iostream>
#include <vector>
#include <array>
int main()
{
std::vector<int> v_ten(0, 10);
std::vector<int> v_twenty(0, 20);
std::array<int, 10> a_ten{0};
std::array<int, 20> a_twenty{0};
std::cout << "sizeof(v_ten): " << sizeof(v_ten) << std::endl;
std::cout << "sizeof(v_twenty): " << sizeof(v_twenty) << std::endl;
std::cout << "sizeof(a_ten): " << sizeof(a_ten) << std::endl;
std::cout << "sizeof(a_twenty): " << sizeof(a_twenty) << std::endl;
return 0;
}
produces this output: sizeof(v_ten): 24
sizeof(v_twenty): 24
sizeof(a_ten): 40
sizeof(a_twenty): 80
edit: added comment about alignment.The mention of array in the parent was referring to the dynamic storage of a vector, not std::array.
This was inherited from Smalltalk. Everything was an "object reference," with tag bits and all, but SmallInteger was optimized by using the low 16 bits of the object reference pointer to store the SmallInteger. So you can think of Smalltalk as having SmallInteger as a "primitive" type with pass by value.
There's several different factors to consider. Value objects will be a valuable addition to the core JVM if only to standardize the many different approaches in place today. But even the standard value types will have strict limitations -- they will likely be immutable and it's not clear that you'll be able to easily obtain their memory address. That means many projects will still need things like code-gen. Code gen, and buffer-backed objects in general, isn't a hack at all it's a proven, battle-tested solution for many domains (finance, scientific sims, big data) that are manipulating several several gigabytes of objects in a performant manner.
Alternatively you can just flatten everything into a single object, but this loses out on abstraction.
For example, if one class includes a reference to another, merge the fields and methods of both classes into one big mega-class, then see if there is any redundant code which can be removed for this new megaclass.
There have been garbage collectors designed to improve locality of predictable accesses, but they are not the norm.
Then there is the fact that you still cannot actually control the allocation. The allocation pattern ABBAABABBA has worse performance on read accesses of all As or all Bs than AAAAABBBBB but with two arena allocators you can write your code to allocate memory according to the irregular pattern but still have continguous data as the end result.
I've seen (10+ years old) low-latency Java apps where people avoid using any OO (ie everything is static methods)... and even then the code was still often easier to maintain and understand than equivalent C++ applications. Nobody does this any more anyways.
You can go even further and employ what's called mechanical sympathy (https://mechanical-sympathy.blogspot.com/) and write your code in a way that makes it easier for the processor to process faster.
I've never worked with application that use off-heap allocations or no-gc requirements.
It's by no means new or even especially difficult these days. The resulting code is far more maintainable and robust than C++ solutions.
And yeah, we restart everyday (we only trade US equities, so plenty of downtime for a restart).
We're currently fighting a Java app that in general has decent latency (10s of usecs), but has outliers of greater than a second when GC kicks in. We don't have that issue with the C++ components of our trading system.
I suppose at the complexity of modern games that's simply not possible anymore...
Not used Span<T> and Memory<T> [2] in anger yet, but they also seem to be attacking the issue from the other end
[1] https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
[2] https://blogs.msdn.microsoft.com/mazhou/2018/03/25/c-7-serie...
You lose all the object/memory management features, though there are some projects around to provide struct-like APIs: https://github.com/alaisi/nalloc
If your player has a 144Hz monitor then your pause time, rounded to the nearest millisecond, has to be 0ms.
In a 7ms frame, you are spending most of that frame doing the work of rendering the actual frame (unless your game is so trivial that the GC is going to be easy / fast anyway). An additional millisecond is going to cause you to miss your deadline and drop a frame. Dropped frames feel really bad.
It's a crazy place here on HN.
Also has it really been a decade since Braid? Wow.
Also note that this category of answer is basically saying, "look, if you mostly manage your own memory, then GC takes less time!" That's true, but a large part of the value proposition of GC in the first place was to remove the burden of memory management. Once you are saying actually, GC won't do that for this class of application, then really what you are getting out of GC is memory safety (provided the rest of the language is memory-safe). On the one hand, hey, memory-safety is a benefit. On the other hand, I don't think very many people in game development would trade that much performance just for memory safety.
(And in fact in game development we very often have to do unsafe memory things. So really what ends up being said is "much of the system has memory safety" which, really, does not sound very alluring.)
that said, the witness is one of my favorite games of all time, so I think its safe to say you know what you're doing
But beyond that ... "smart pointers" and the like make your program slow, because they postulate that your data is a lot of small things allocated far away from each other on the heap. I spoke about this in more depth at my Barcelona talk earlier this year.
This is a good point. I've spent a fair amount of time trying to build performant stuff on the web, and my conclusion has been that if you really care about memory management, garbage collection is not your friend.
You spend more time than is necessary trying to convince the engine not to allocate anything you don't want, trying to lay out objects ahead of time, etc... I have spent more time learning about how Chrome's garbage collection works than I would have ever spent manually tracking references myself.
And the end result is that I still have to manually track references and be careful about where things are allocated - the garbage collector for me is objectively a net negative, not just for performance but for dev time and program complexity. And I'm not even really doing anything all that complicated in any of the software I build -- it is not hard to run into these kinds of issues.
This is something that I didn't understand until I saw it in my own software, but I strongly suspect that if I could get rid of browser garbage collection I would spend less time thinking about memory than I do now.
Do you have any good references you can share regarding Chrome's garbage collection? This is a topic I would like to learn more about.
At some point in the future I will probably write a blog post about good techniques for taking back control of JS memory management from the browser, but it probably won't be for a while.
The short answer though is that browsers try to defer garbage collection until it looks like the page isn't doing much; and the problem is that games don't ever really stop doing stuff. This encourages the browser to put off garbage collection and handle it in chunks, which in turn leads to dropped frames because you're doing a lot of deallocation at once.
This is why if you profile a Chrome app that's doing a lot of allocations very quickly, your memory usage will kind of sawtooth all over the place. And you get around that by getting rid of as many allocations as possible.
IMO the lack of garbage collection in WASM is the most exciting thing about it, but different languages have different tradeoffs and WASM isn't appropriate for every application.
It's also not quite free since it takes up TLB space.
In practice, I would wager the performance concerns aren't huge, but that's mostly just a guess. I'd personally be interested in seeing a comparison to masking to see just how much slower masking would be, but obviously it's not like they can just flip a switch to use masking.
It doesn't. The article mentions it's only 3 mappings since those bits are colors, not arbitrary combinations of flags.
> I don't know how Java does it's memory management but if it does lots of small mappings (Which I'm guessing it does not) then that could be a concern.
openjdk generally uses large contiguous mappings but it may punch holes in the middle of the heap if it's configured to yield back memory to the OS. But applications that dynamically shrink and expand their heaps are not necessarily those that are concerned about the last quantum of page table overhead.
> And like you mentioned with the TLB
On the other hand it does support huge pages to mitigate costs of TLB entries.
Nope. They went into this in the article:
> Since by design only one of remap, mark0 and mark1 can be 1 at any point in time, it’s possible to do this with three mappings. There’s a nice diagram[1] in the ZGC source for this.
[1]: http://hg.openjdk.java.net/zgc/zgc/file/59c07aef65ac/src/hot...
That might not be totally current, as it doesn't cover the finalizable flag, but if it works the same, that would only be four mappings. If it works differently, then it would be a maximum of 6 mappings. Not 16.
In other words, just because the alternate mappings exist doesn't mean the pointers always have values that actually use them.
Basically, you don't want to check for gc in each (allocation free) loop iteration. On top of the overhead it makes optimization harder. But if you have 32+ threads then some are probably stuck in loops for a good while which means spinlocks for everyone waiting for the stop-the-world gc.
The trick to avoid this is to add an inner loop that does a fixed number of iterations and then do a gc check between each inner loop https://bugs.openjdk.java.net/browse/JDK-8186027
I can't believe the world we're living in, this is incredible. My first computer 512k of ram.
[1] https://en.wikipedia.org/wiki/Value_type_and_reference_type
Writing everything on the main UI thread, not enabling system L&F, reading books like "Filthy Rich Clients", adopting third party Swing components with modern L&F....
But of course, bad programmers make bad apps, no matter how good the ecosystem is.
Java, from my experience, has always been quick and flexible.
I would think if there were processes that needed consistent latency you would isolate them.
This is a big deal, and a major departure.
I'm hoping this will improve things, this new architecture more closely resembles Azul's C4. Java and large heaps with non-conforming object allocation patterns (medium aged objects are the most troublesome) sucks really badly.
(If it also means anything, I can confirm that enabling Shenandoah is as simple as taking the corresponding OpenJDK code and replacing the `hotspot` subcomponent with the (compatible) Shenandoah fork'd hotspot component, then just building everything normally.)
(Is this the same as pointer tagging?)
> ...
> ZGC restricts itself to 4Tb heaps
Isn't a 4TB heap limit short-sighted for a GC intended to last over a decade?
If anything, borrow a few more bits, and you should easily be able to address petabytes.
Can't you compress that into less than 42 bits?
https://en.wikipedia.org/wiki/Pigeonhole_principle#Uses_and_...
There's more holes than pigeons here.
How do I compute the hash of every object, since by definition I can't address every byte?
And why would you want to read raw bytes from an object to hash it? What requires this in Java?
(I'm guessing you were think of the same [Which apparently is just apocryphal])
Besides, processes are a ham fisted way to get around the PDP-11's limited address space and shouldn't be a thing in 2018 anyway.
At least they seem to be Azul´s biggest customer base.
Basically, graphs are hard to shard ... what if we didn't bother? You can query an in memory graph pretty quickly out into the 2nd or 3rd degree.
Even G1 (untuned) only pauses for 250ms once every few minuets. It's not perfect but good enough to ignore.
Does anyone know what this instruction on SPARC is?
This can also destroy investment rounds for many startups.
They could sue you, with a higher chance of winning, for using a python, a js or a ruby than for using OpenJDK derived code.
Do you happen to know of any real-world case that's similar to the scenario you've described?
They only support part of the standard library, the OpenJDK parts have been cherry picked and not all are made available on the same Android version and then there are the bytecodes that were introduced since Java 7.
And for Java language support, you need to have Android Oreo version of ART as not all language features get desugared into older versions.
There is some AOSP activity related to Java 9, but so far Google has been silent if there will be any further improvements beyond Java 8.
Oracle/Google/MS or Apple will continue to exert control on direction of languages they developed and open sourced because they continue to pay for majority of development expense.
Below is a link for the downvoters and skeptics:
https://www.aspera.com/en/blog/oracle-will-charge-for-java-s...
edit: I looked it up.
Your comment is obviously correct and valid. However, my stance is unchanged. I will stick to openjdk8 and avoid oracle regulatory influence.
> Oracle has announced that, effective January 2019, Java SE 8 public updates will no longer be available for "Business, Commercial or Production use" without a commercial license.
Anything you can get now is still free and GPL open sourced. All they're doing is no longer publishing NEW CODE (in the form of backports and bug fixes) to the OpenJDK 8 branch. Instead they're publishing it to OracleJDK.
They're not trying to kill GPL anymore than Redhat is trying to Linux.
Talk about an intriguing headline.
OpenJDK, on the other hand, was unaffected by this change; things like security mitigations/fixes land in OpenJDK, and make their way to other versions (like older JDK versions, or Oracle enterprise copies) from there. So, you can go download Zulu or OpenJDK binaries from your distributor (such as your Linux distribution) and you'll be fine (assuming they update promptly).
There are also feature distinctions between the Oracle JDK and OpenJDK, but they have been getting smaller over time -- AppCDS and Flight Recorder, which were Oracle-customer-only features, for example, are now in OpenJDK proper (FR will come in JDK 11), and many features such as ZGC (and originally G1, too, I think) were developed in the open and went into OpenJDK directly, right from the start.
If you're just running a bog standard Linux/BSD system with some Java software on it, you're almost certainly OpenJDK already anyway, and your distribution maintainers handle security updates for you.
You can either:
1. Update it every six months, for free, if you want to keep current with security updates.
2. Pay for updates on versions older than six months.
3. Not pay, not upgrade, and live with unsupported versions.
Adopt Open JDK has indicated their intention to back port patches to older versions for several years past their expiration for free.
All of the proprietary bits in Oracle's Java implementation are also being fully open sourced: https://www.infoq.com/news/2017/10/javaone-opening
[1] http://blog.joda.org/2018/08/java-is-still-available-at-zero...
Particularly what would work well on x86 and ARM.
An odd and probably completely wrongheaded idea just came to mind: add collectible objects to a lock-free queue in one thread, collect them in another. I wonder if you could do the mark-and-sweep stage across threads too...?
Pragmatically in high-throughput situations modern GC's perform better than simple GC. The new things in Java land is that the newest GCs have vastly reduced stop the world times (trending towards OS jitter times), while preserving compaction and without to bad throughput hits. (Even with the misfeatures of finalizers)
My personal experience with Shenandoah shows that this changes the way we think about heap settings for java. i.e. just set -Xmx to 60% of machine ram and let idle collections make sure we never use more than 4% in normal days.
Well, refcounting is the basic way of implementing GC and is listed as GC algorithm in any serious CS book about GC algorithms like "The Garbage Collection Handbook".
What many refer as GC is actually the GC algorithms that fall under the umbrella of tracing GC.
Then refcounting is only faster than GC in very constrained scenarios:
- no sharing across threads, otherwise locking is required for refcount updates
- no deep datastructures, otherwise the call stack of cascading deletions will produce a stop similar to tracing GC
- implementations need to be clever about nested deletions when refcount reaches 0, otherwise a stack overflow might happen
- cyclic datastructures need programmer help to break cycles via weak dependencies or are just not allowed
This is only relevant for naive refcount implementations, there are many optimisations, which endup turning a refcounting implementation into a tracing GC in disguise.
Also just because a language uses a tracing GC algorithm, it does not prevent the existence of value types, or manual memory allocation for hot code paths, thus allowing for more productive coding, while offering the tools for memory optimisation when needed.
This is not yet the case with Java, but even here it is part of the language's roadmap to fix this.
That is not true. Most atomic refcount implementations are lockfree (you do need synchronization though). This optimization has nothing to do with tracing. Given that you also need synchronization for tracing garbage collection (including frequent stop the world pauses in most cases, albeit brief ones), I don't think this is even really an advantage of tracing at all.
> - no deep datastructures, otherwise the call stack of cascading deletions will produce a stop similar to tracing GC
If your program has data structures that live through several collections, refcounting can be faster than tracing GC because it only traces objects once (when they die) instead of several times. Additionally, you can defer the refcount updates in ways that avoid the need to do lots of work on deallocation, and optimize for short-lived objects, to get many of the effects of generational GC. This also has little to do with tracing (except inasmuch as a reference counter with this optimization has to "trace" new objects to make them initially live, but in some sense this work to update the reference counter for the first time is just moved from object initialization time to a later point in the program). The reason it's not usually done is that it requires precise liveness information by an ambient collector, which complicates the implementation, but "exploiting liveness information" is not the same as "being a tracing GC."
> - implementations need to be clever about nested deletions when refcount reaches 0, otherwise a stack overflow might happen
This is extremely trivial to avoid (if you need to) and has nothing to do with tracing. It's really more a product of user-defined destructors than anything else, which you don't need to provide in order to implement reference counting. In fact, a lot of the supposedly inevitable slowness of reference counting compared to tracing goes away when you ban nontrivial destructors (and conversely, nontrivial destructors make tracing perform much worse).
> This is only relevant for naive refcount implementations, there are many optimisations, which endup turning a refcounting implementation into a tracing GC in disguise.
The most optimized versions of refcounting I'm aware of get their wins from things like precise knowledge about live references and update coalescing--not tracing, except for some young object optimizations as I alluded to above which are more about satisfying the generational hypothesis (they generally have a backup tracer in order to break cycles, but if you're willing to live without that they usually don't need tracing to reclaim all memory). Many optimizations people commonly associate with tracing (e.g. cache friendliness due to compaction) are in practice only relevant for young generations most of the time; for older ones both tracing and reference counted implementations tend to benefit more from a really smart allocator with intelligent memory layout and partial reclamation.
I agree that optimized versions of refcounting are difficult and the fact that you still need a backup tracer for most code discourages people from using it, as well as that tracing doesn't negate a lot of systems optimizations. But a lot of the stuff you're saying is pretty misleading: the things that make tracing efficient can mostly be applied to make reference counting efficient without "turning it into tracing," with the cost that optimized tracing and optimized reference counting both need much more invasive knowledge about the user program (and correspondingly restrict the users) compared to the less optimized versions.
I know Python, objective-c, and swift all use refcounting, but just for my own reading, what language has the most optimized ref counting right now?
There may have been further developments in improving RC throughput since 2013, but I'm not familiar with them; outside of Apple, there's very little modern research into optimizing reference counting. I know Swift does many cool optimizations of its own, but I'm not sure how restricted they are by having to remain compatible with Objective C.
More generally, I think the parent comment was not aiming to say that reference counting is never appropriate, but that blindly going with reference counting because it's now "ew, GC" is misguided.
That's not actually true of deferred reference counting with update coalescing. Instead, it defers the refcount increments and decrements to periodic intervals, much like optimized tracing implementations defer tracing to periodic intervals. This avoids reference count storms on the same objects and (especially for new objects) means much less work overall.
As I noted, there's at least one high performance GC implementation for Java that uses it, with effectively identical performance to a state of the art tracing GC on all benchmarks considered--it's not a toy idea. I don't think the fact that most production languages don't use it means very much, given how much more effort has been spent tuning tracing compared to reference counting.
> This probably also removes one major advantage of RC'ed systems though: deterministic destruction (e.g. close resource as soon as last reference disappears). And suddenly you need some kind of runtime too.
Completely deterministic destruction in the sense that you describe is pretty much incompatible with achieving the highest collection throughput. In particular it means you always have to trace dead objects independently (and, if all cores are busy, block someone from doing useful work while you deallocate), so most young object optimizations are a nonstarter. If you can stack allocate (or equivalent, e.g. push and pop from a vector with a trivial destructor) almost all your young objects, great; otherwise, for my money, keeping latency-sensitive things in well-scoped regions is far more deterministic and effective than either RC or tracing. The number of non-memory resources in most programs is usually rather low; you can always use traditional reference counting (or linear types or whatever) just for those.
> And suddenly you need some kind of runtime too.
Yes, which I mentioned a number of times. However, I don't think it's unreasonable to afford reference counting the same conveniences that tracing has when you want to talk about performance; otherwise you're not really comparing reference counting to tracing, but naive reference counting to highly optimized tracing. The fact that tracing essentially requires a runtime is certainly a point against it in some contexts (though conservative GCs like Boehm perform surprisingly well, they are usually no match for anything with precise liveness information), but I think people are a bit too conditioned to believe that things like stack maps have to be used for tracing--especially when features like exceptions and destructors that people consider acceptable in languages with small runtimes require similar metadata.
Not sure I agree with that. Copying references between local variables and reads/writes from heap all require expensive atomic operations for RC when they can't be optimized away. That's a major performance problem for languages like Swift. I do not say that one is better than the other, but this is exactly where tracing GC's shine compared to RC.
In the case of ZGC these doesn't require atomic operations, you need a read barrier for reading references from the heap though. But do not conflate tracing GC's read & write barriers with atomic barriers.
Again, the optimizations I'm describing are mostly distinct from turning RC into tracing, just applying the same sorts of optimizations we expect from production garbage collectors. The only exception is probably how in RCImmix, heavily referenced objects (4 bits are enough to precisely track something like 99.8% of objects, so "heavily referenced" refers to the other 0.2%) have their reference counts frozen so they don't pay anything until the backup trace starts. But it seems like most of the win from freezing reference counts comes from using fewer bits for the count, not avoiding the updates per se.
As usual, people are confused by improper use of terminology. It's like async in Python, three quarters of the complexity is in abuse of terminology.
Yes, of course, given the same code and allocation pattern, there are many cases where tracing GC will give you higher throughput than reference counting, particularly once you add in cycle detection. But in GC'd languages and non-GC'd languages you write code with completely different allocation patterns. In non-GC'd environments you only need to refcount the small set of allocations that are actually shared. GC'd languages usually have semantics that require tracing for every allocation.
[citation needed]
Mesa/Cedar, Active Oberon, Modula-3, D, System C#, Go, C#
The features are available, it is up to the programmers to use them
All those languages are GC enabled systems programming languages.
Regarding Go the memory safety in multithreaded code is orthogonal to having support for value types or not.
That or you start pinning managed objects or pull everything related into manual world. Both kinda suck.
In any case, the point is that you should only do that for the code paths that actually matter, after profiling the application and not everywhere.
Be productive, make use of the GC, naturally taking into the consideration the algorithms + data structures being used.
If that is still not enough to meet the application's SLAs, then manually allocation comes into play.
So, I disagree: all of the languages you mentioned either are memory unsafe when you use multiword value types concurrently, have one of the restrictions I mentioned, or (in the case of Active Oberon) conservatively lock on all writes. These value types are not "unrestricted" compared to primitives, all of which can be mutated in-place without allocating or locking, or to heap objects, which can be updated atomically even for interesting values at the cost of an allocation. Moreover, the issue is intimately tied to GC, or more accurately to the programmer not having to explicitly keep track of whether objects are uniquely owned or not (which is generally the function of a GC). As a result, if garbage collected objects can hold references to nontrivial value types, there is almost no way to avoid this problem.
We are discussing GC enabled systems programming languages with support for unsafe operations, not Rust's idea of thread safety.
D is memory safe in the context of @safe code, of course @system code is unsafe.
Active Oberon and Modula-3 also have record types, strings and arrays that can be stack allocated, declared on the global data memory segment or just manually allocated in unsafe packages.
In both cases I can also declare untraced pointers for memory blocks in unsafe packages, that the GC will gladly ignore.
Record types with fields that don't depend on one another, strings and (fixed-size) arrays only support "uninteresting" mutations that remain valid even in the presence of atomic tearing. They are restricted compared to the general value types that you can have in a language like C++ (for example, sum types, unions, "fat objects" like Go's interfaces, or vectors with dynamic lengths), because for the latter having a write tear can lead to undefined behavior (for instance, overwriting a value vector can lead to the pointed-to vector part temporarily having a different length than the length field would indicate, which can easily lead to a buffer overflow).
The memory safe fragment of D has similar restrictions on its value types to the above. For example, its value arrays must have a size known at compile time.
Of course you can support such features in the unsafe fragment of a language. However, this is a pretty significant deterrent to actually using the feature and idiomatic code will avoid it most of the time.
If you are sharing across threads you will need a synchronization mechanism, whether you have reference counting or not.
So it's not exactly fair to assign this cost to reference counting.
General reference counting does make adding or removing a reference a read/write operation, though of course non-naive implementations don't do this at every turn. Not to mention tracing GC has a reference cost as well.
One can optimize refcounting with deferred reference counting algorithms, but then that is already the skeleton implementation of a tracing GC.
Throughput is not the only concern when designing a memory management strategy though.
Moon instructs a student
One day a student came to Moon and said: “I understand how to make a better garbage collector. We must keep a reference count of the pointers to each cons.”
Moon patiently told the student the following story:
“One day a student came to Moon and said: ‘I understand how to make a better garbage collector...
It's the only language that is anywhere close to meme status that I know of.
It's like you're claiming you are famous because your mom knows who you are.
It's very impressive work, no doubt. And as long as no one asks me to debug anything running on top, I don't have any issues with it.
At least with straight reference counting I know what's going on under the hood, that's worth a lot. Sharing pointers across threads is tricky business, I can't really see what that has to do with anything. And it's quite likely that it will run faster, simply because it's much less complex.
https://www.youtube.com/watch?v=1f-kfRREA8M
And the Ministry of Silly Walks is the perfect metaphor for user interface design.
Silly walks is a masterpiece, and people think they're joking...
http://www.cs.virginia.edu/%7Ecs415/reading/bacon-garbage.pd...