I have also spent some time on 2) and implemented several approaches in this open source optimising llm proxy - https://github.com/codelion/optillm
In my experience it does work quite well, but we probably need different techniques for different tasks.