> and get the same performance characteristics as a GC language
Pretty sure modern state of the art GCs, like the ones found in Java or .NET runtimes, are way more efficient than reference counting.
Reference counters have that unfortunate RAM access pattern where the counter is frequently updated from multiple threads concurrently. On modern processors, this means concurrent access to a cache line by different CPU cores. These cache coherency protocols are pretty slow. On many CPUs, reading a cache line recently modified by another core costs 300-500 clock cycles, even more expensive than a cache miss and roundtrip to system RAM.