You need a contiguous chunk for whatever object you are allocating. Other allocations fragment the address space, so there might be adequate space in total, but no individual contiguous chunk is large enough. You need to move around the backing storage, but then that makes your linear addresses non-stable. You solve that by adding a indirection layer mapping your "address", which is really a key/ID, to the backing storage. At that point you are basically back to a MMU.
It may be ok for embedded systems, but those recently have been evolving on the opposite direction.
That means there is no indirection between reference and the backing physical store. To change the backing memory you need to change the reference.
As you point out, within a program you can solve fragmentation by using a compacting GC. However, you will still suffer from cross-program fragmentation. You can not move the backing store of those programs without rewriting their references. You can not even page out those programs and reload them into new locations without rewriting all their references. You effectively need a global, shared, cross-program compacting GC to handle fragmentation; a hardware or kernel GC.
In MacRelix (a POSIX-like environment for classic Mac OS), when a process calls fork(), the system allocates backup memory regions for it and its child. Whenever one of them is switched in after its counterpart was the last of the two to run, the old one's regions are backed up and the new one's regions restored.