There are a lot of things that make GC performance an issue, including stop-the-world behaviour for some collectors making real-time operations impossible as well as the tendency for programmers to produce a lot more garbage in their programs. In general, I think that most people overstate the percieved inefficiencies of GC as it comes to throughput (predicatable latency is the main culprit IMHO).
It is true that GC works best when you have sufficient amounts of free memory compared to the amount of garbage produced. On any kind of memory-constrained (embedded) device I would naturally view everything about memory allocations as critical issues to handle, with clear memory usage budgets.
But is performance/core also going to keep rising? Or do we need smart compilers/languages to make use of all the cheap cores?
Languages will provide the semantics (in context of memory hierarchy) of data locality, and compilers that optimize for that.
[nop edit]
My point was, people are already thinking. It's not an easy problem and there's no obvious way forward.