The problem is that with diverse hardware such a set of defaults is much harder to make. You get this matrix of possibilities: gpu, VRAM and memory configurations and then the model axis. This leads to way too many options. The better alternative would be to have the runner self-benchmark what the best settings are given that it already has access to that one particular configuration.
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
llama-bench is next to useless for this purpose.