The data cache memory is one of the solutions to avoid the extremely long latency of loading data from a DRAM memory.
The alternative to a data cache memory is to have a hierarchy of memories with different speeds, which are addressed explicitly.
The latter variant is sometimes chosen for embedded computers where determinism is more important than programmer convenience. However, for general-purpose computers this variant could be acceptable only if the hierarchy of memories would be managed automatically by a high-level language compiler.
It appears that writing a compiler that could handle the allocation of data into a heterogeneous set of memories and the transfers between them is a more difficult task than designing a CPU that becomes an order of magnitude more complex due to having a hierarchy a data cache memories and a long list of other hardware mechanisms that must be added due to the existence of the data cache memory.
Once it is decided that the CPU must have a data cache memory, a lot of other hardware design decisions follow from it.
Because there is an inverse relationship between the load latency and the data cache memory size, the cache memory must be split into a multi-level hierarchy of cache memories.
To reduce the number of cache misses, data cache prefetchers must be added, to speculatively fill the cache lines in advance of load requests.
Now, when a data cache exists, most loads have a small latency, but from time to time there still is a cache miss, when the latency is huge, long enough to execute hundreds of instructions.
There are 2 solutions to the problem of finding instructions to be executed during cache misses, instead of stalling the CPU: simultaneous multi-threading and out-of-order execution.
For explicitly addressed heterogeneous memories, neither of these 2 hardware mechanisms is needed, because independent instructions can be scheduled statically to overlap the memory transfers. With a data cache, this is not possible, because it cannot be predicted statically when cache misses will occur (mainly due to the activity of other execution threads, but even an if-then-else can prevent the static prediction of the cache state, unless additional load instructions are inserted by the compiler, to ensure that the cache state does not depend on the selected branch of the conditional statement; this does not work for external library functions or other execution threads).
With a data cache memory, one or both of SMT and OoOE must be implemented. If out-of-order execution is implemented, then the number of registers needed to avoid false dependencies between instructions becomes larger than it is convenient to encode in the instructions. so register renaming must also be implemented.
And so on.
In conclusion, to avoid the huge amount of resources needed by a CPU for guessing about the programs, the solution would be a high-level language compiler able to transparently allocate the data into a hierarchy of heterogeneous memories and schedule transfers between them when needed, like the compilers do now for register allocation, loading and storing.
Unfortunately nobody has succeeded to demonstrate a good compiler of this kind.
Moreover, the existing compilers have frequently difficulties in discovering the optimal allocation and transfer schedule for registers, which is a simpler problem.
Doing efficiently the same for a hierarchy of heterogeneous memories seems out-of-reach for the current compilers.