Why would allocation become more expensive in parallel? Surely you’re allocating in thread-local space? It’s like two machine instructions. Where does the extra overhead come from?
To avoid frequent synchronisation we could have gone with a (potentially generational) non-moving collector for the minor GC which would still have preserved the C API, probably allowed for very low pauses but would have made allocation more expensive.
I don't really understand why not. Can you expand on it? Doesn't the JVM for example do thread-local allocations in a conventional shared-memory parallel environment?