Man, I'm not so sure if I'd use something like this because the way I prompt already changes based upon what model I am using. I'm not convinced it would route to the right model based on my diction or whatever.
Hard to quantify this ofc but that's what I've felt vibes wise from using this for the last month.
It's also possible that it's the 1m context versus the 200k context (Copilot's limit) doing some of the work here.
Perhaps you're just not the best use case. It may work better when Average Joe is the one prompting.