> Say I'm doing a drawing/game app and creating a few hundred heap objects a second that need to get garbage collected.
Was literally my job ten years ago to optimize this and I was struggling with a GC'd language with a proprietary implementation (flash+actionscript).
The problem is not with hundreds of heap objects per-frame, the problem is that they would accumulate to the tens of thousands before the first GC trigger happens.
And the GC trigger might happen in the middle of drawing a frame, even worse, at the end of drawing a frame (which means even a 10ms pause means you miss the 16ms frame window at 60fps).
The problem that most people had was that this was unevenly distributed and janky to put it in the lingo. So you'd get 900 frames with no issues and a single frame that freezes.
So most of the problem people have with GC pauses is the unpredictability of it and the massive variations in the 99th percentile latency in the system, making it look slower than it actually is.
Most of the original GC implementations scale poorly as the memory sizes went up and the amount of possible garbage went up, until the GC models started switching over the garbage-first optimizations, thread-local alloc buffers and survivor generation + heap reserves etc (i.e we have lots of memory, our problem is with the object walking overheads - so small objects with lots of references is bad).
The GC model is actually pretty okay, but it is still unpredictable enough that tuning the GC or building an application on top of a GC'd language which has strict latency requirements is hard.
However, as a counterpoint - OpenHFT.
Clearly it is possible, but it takes a lot of alignment across all the system layers, but at that point you might as well write C++ because it is not portable enough to run anywhere.