DeepSeek is seemingly novel in its caching capability. In most of my usage of it, I'm seeing something like 80% cached tokens, which makes using DeepSeek cheaper than damned near any other model anywhere near its size. I don't think even my self-hosted models are cheaper once electricity and air conditioning is accounted for (it's summer in Texas, pumping a few hundred watts of heat into the room for hours at a time for local LLM usage is sub-optimal).
So, I think if you're using a model that has extremely effective caching (of which there is only one I know of) as your expensive model, you almost certainly won't come out ahead by switching back and forth. Cached tokens in DeepSeek are damned near cheap as free.
But, for Claude or GPT where tokens are extremely expensive and caching barely puts a dent in it, yeah, I guess there is math that could work out in favor of this idea.