See also “A unified theory of garbage collection” by Bacon, Cheng, and Rajan (OOPSLA ’04) for a discussion of how tracing and RC essentially compute the least and greatest fixed points, respectively, of the “roots and referents” function (the difference between the two being the reference cycles), and of course the Garbage Collection Handbook by Jones, Hosking, and Moss (CRC, 2012), which includes both tracing and RC in its definitions of GC and draws comparisons (barrier costs, heap fragmentation, etc.).
(These are both bog-standard references, the thread above has people know much more about this talking about ideas that aren’t 20 years old.)
Note that as a rule of thumb, tracing garbage collectors need about twice as much memory as live objects, to avoid frequent heap scans. Non-delayed reference counting only keeps around live objects, which can significantly reduce memory usage.
If you have a huge number of live objects that you need to keep around, but infrequently have references to them created (e.g. very large caches), then reference counting is going to out-perform tracing collection, particularly if some of those infrequently used objects might be paged out and you're not using card marking to sometimes avoid needing to trace the paged-out objects. It's convenient that these use cases also tend to work well with reference counting's generally lower peak memory requirements.
On the other hand, if you're rapidly creating and destroying objects, or often updating references, tracing may perform better. Moving collectors (at least using moving or compacting collectors for the young generation) can use bump allocation, resulting in much faster object allocation.
The Garbage Collection Handbook goes into nearly countless variations, hybridizations, and optimizations on these basic collectors.
For over a decade, I helped maintain Goldman's Slang language. Some components had bog standard mark-and-sweep collector, but most of the language uses reference counting. The language uses tight integration with the SecDb globally distributed NoSQL database and pervasively locally caches database objects and memoizes method calls on those objects. I'm sure naive reference counting isn't optimal for the Slang use case, but it's easy to implement and isn't actually that far from optimal. (When I first started at Goldman, some parts of the language used an implementation of Bacon's cycle-detecting concurrent reference counting collector, but there was some rare corner case where someone missed a reference count change on the C++ side, and after a while of trying to find the missing reference count, we just replaced that component with a non-generational mark-and-sweep collector.) (There was a publish-subscribe system that would watch the SecDb transaction logs, so applications could subscribe for cache invalidation notifications. With huge caches, you don't want to use database polling for cache invalidation.)
Therefore you apply techniques like Apple's ARC, or deferred-coalesced counter mutation.
Add cycle leak detectors and the RC scheme is looking more and more like a really slow garbage collector, albeit a low latency one.
Note that simple RC is not deterministic or good for real time either because of the possibility of release/destruct cascade. Which looks pretty much like a GC pause (because it is one).
When a reference count of an object goes to zero, you recursively decrement the reference counts of everything the object points. This leads you to trace the object graph of things whose reference count has gone to zero. You never follow pointers from anything which is live (i.e., has a refcount > 0).
When you are doing the copy phase of a gc, you start with the root set of live objects, and follow pointers from everything that is live. Since anything pointed to by a live object is live, you only follow the pointers of live objects. You never follow pointers from anything which is dead (i.e., garbage).
If object lifetimes are short, most objects will be dead, and so RC will be worse than GC. If object lifetimes are long, most objects will be live, and GC will be worse than RC.
Empirically, the overwhelming majority of objects have a very short lifetime, with only a few objects living a long time. (This is called "the generational hypothesis">) So the optimal memory allocator will GC short-lived objects and RC long-lived objects. Rust/C++ encourages you to do this manually, by stack-allocating things you think will be short-lived, and saving RC for things with an expected long lifetime.
Beyond this, RC has a few really heavy costs.
Reference counting doesn't handle cyclic memory graphs. You need to add tracing to handle those, and if you are going to do tracing anyway, it's tempting to just do tracing really well and skip the refcounts entirely.
This is because the memory overhead of reference counts is high -- empirically, most objects never have more than a single reference to them, and so using a whole word for reference counts is lot of overhead. Moreover, the need to increment/decrement reference counts is really bad for performance: first, mutations are expensive in terms of memory bandwidth (you've got to maintain cache coherence with the other CPUs), and second, in a multicore setting, you have to lock that word to ensure the updates are atomic.
There are tricks to mitigate this (e.g., Rust distinguishes Arc and Rc for objects which can be shared between threads or not), and there are schemes to optimise away RC assignments with static analysis (deferred reference counting), but if you want to do a really good job of reference counting, then you will be implementing a lot of tracing GC machinery.
And vice versa! The algorithm in the link is partly about adding RC to handle old objects (empirically, as part of the generational hypothesis, objects which have lived a long time will live a long time more). In fact, Blackburn and McKinley (two of the three authors of the above paper), pioneered the combination approach with their paper "Ulterior Reference Counting."
It's also not uncommon to have a copying young generation and a mark-sweep-compact tenured generation. That way, you get the advantages of not needing to scan the huge numbers of young dead objects, but the space savings of not needing 2x space for the older generation.