Instead have Claude know when to offload work to local models and what model is best suited for the job. It will shape the prompt for the model. Then have Claude review the results. Massive reduction in costs.
btw, at least on Macbooks you can run good models with just M1 32GB of memory.